A Hierarchical Federated Learning Optimization Method in a Dynamic Edge Computing Environment
By using dynamic clustering algorithms to optimize the equipment aggregation topology, predicting and recommending the optimal training frequency, dynamically adjusting the training rounds, and establishing a hierarchical time-out fault tolerance mechanism in the dynamic edge computing environment, the problems of insufficient resource utilization of edge equipment and insufficient adaptability to dynamic changes in the existing technology are solved, and efficient distributed model training is achieved.
Patent Information
- Application Number
- CN202510475312.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art is difficult to effectively deal with the problem of communication bottlenecks of central nodes caused by the large number of edge devices in a dynamic edge computing environment, and cannot fully utilize all edge device resources. The static aggregation structure and frequency optimization are insufficient to adapt to the dynamic changes of device status and network bandwidth.
Optimize the equipment aggregation topology through dynamic clustering algorithms to maintain stable device relationships while adapting to device changes; predict and recommend the optimal training frequency based on historical performance data to realize resource-aware adaptive training scheduling; the equipment dynamically adjusts the training rounds according to its own resource status; establishes a hierarchical time-out fault tolerance mechanism to handle abnormal equipment status.
The distributed model training efficiency in the edge computing environment is significantly optimized, the waste of computing resources and communication bottlenecks are reduced, and the overall efficiency of model training is improved. It is suitable for edge computing scenarios with heterogeneous computing capabilities and frequent network fluctuations.
Smart Images

Figure CN120017513B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge computing, and particularly relates to a hierarchical federated learning optimization method in a dynamic edge computing environment. Background Art
[0002] The current application of distributed machine learning technology in the edge computing environment mainly focuses on the heterogeneity problem and hierarchical structure optimization. The mainstream methods include adaptive aggregation strategies (such as DRAG), partial model averaging frameworks, client selection schemes (such as clustering sampling), and heterogeneous-aware scheduling mechanisms. These methods have achieved certain results in centralized structures such as federated learning, but there are still significant limitations: on the one hand, they are difficult to cope with the communication bottleneck problem of the central node caused by the large number of edge devices and cannot fully utilize the resources of all edge devices; on the other hand, although the existing hierarchical structure research (such as Spread, Auto-Group, HPFL) alleviates the impact of non-independent and identically distributed data through grouped aggregation, it generally ignores the characteristics of dynamic changes in device resources. Although the KMA and RAF algorithms improve the training efficiency through static aggregation structures and frequency optimization, their adaptability to dynamic changes such as device status and network bandwidth is insufficient, and it is difficult to maintain stable performance in a real edge environment. Summary of the Invention
[0003] Object of the Invention: The technical problem to be solved by the present invention is to provide a hierarchical federated learning optimization method in a dynamic edge computing environment in view of the deficiencies of the prior art. For each round of distributed model parameter training and aggregation of hierarchical collaboration, the following steps are performed:
[0004] Step 1, optimize the device aggregation topology of collaborative training composed of terminal devices and aggregation nodes through a dynamic clustering algorithm, while maintaining stable device relationships and adapting to device changes; the terminal devices include smart phones and Internet of Things terminals; the nodes include edge servers and gateways;
[0005] Step 2, predict and recommend the optimal training frequency based on historical performance data to achieve resource-aware adaptive training scheduling;
[0006] Step 3, the device dynamically adjusts the number of training rounds according to its own resource status to balance the training intensity and time efficiency;
[0007] Step 4, establish a hierarchical timeout fault tolerance mechanism to handle abnormal device conditions while ensuring the continuity of training.
[0008] Step 1 includes:
[0009] Step 1-1, optimize the device aggregation topology based on the dynamic clustering algorithm, and cluster and divide the edge devices according to the principle of minimizing pre-training accuracy and communication delay;
[0010] The dynamic clustering algorithm includes the following steps: calculating the pre-training accuracy difference and communication delay between devices; selecting central nodes with the goal of minimizing the total communication delay within the cluster; preferentially allocating new devices to the cluster with the largest optimization amount of pre-training accuracy difference;
[0011] The formula for selecting and aggregating different devices is:
[0012] ,
[0013] where is the th cluster of the , , which contains several nodes. The common aggregation node of the nodes in is , being the central node; N refers to the number of edge devices included, and are respectively the node representations of the th device in the th layer and the node representation of the th device in the th layer. The same device can be a node in different layers; H is the total number of layers of the aggregation structure, refers to and 's communication distance, refers to the pre-training accuracy of the th device, refers to the average pre-training accuracy of the th cluster of the th layer, is the weight coefficient used to balance the calculation weights between pre-training accuracy difference and communication delay;
[0014] Step 1 - 2, after each round of training ends and before the next round of training starts, dynamically adjust the aggregation structure: When the th round of training is completed and the root aggregation node sends the model parameters obtained from training (including but not limited to trainable parameters such as neural network weights and biases) to the control node of the central management unit responsible for coordinating the training process, the control node compares the th round (the devices in the th round refer to edge devices that meet the following conditions: 1. Currently online and network reachable; 2. Sufficient storage space; 3. Computational resources meet the minimum training requirements) with the device set in the th round: If the devices remain unchanged, keep the aggregation structure unchanged; otherwise, count the set of exited devices and the set of newly added devices and generate the next round of aggregation structure based on the following formula :
[0015] The tree structure of the r-th round of training is :
[0016] ,
[0017] where represents the tree structure of the r-th round of training;
[0018] Step 1-3, delete the aggregation structure of the departing device: For the departing device , obtain the highest occurrence layer of device in the aggregation structure . If = 0, directly remove the node of device at the 0-th layer ; if , then from the 0-th layer to the -th layer, mark the node corresponding to device at the -th layer as and update the relevant cluster center nodes; represents the vacancy mark generated after deleting device ;
[0019] Step 1-4, add the aggregation structure of the newly added device: For the newly added device , traverse all clusters at the 0-th layer, calculate the optimization gain of the average pre-training accuracy of the cluster after adding device , and select the cluster with the largest gain to add. If there is a cluster that is empty , then set the average pre-training accuracy of
[0020] to 0 or a negative value to ensure that new devices preferentially fill empty clusters;
[0021] Step 1-5, method for adjusting the aggregation structure of supplementary vacant aggregation nodes: Detect vacant aggregation nodes layer by layer from bottom to top . If the center node of cluster is in a vacant state, reselect the center device based on the dynamic clustering algorithm and synchronously update the associated nodes of each layer.
[0022] In Steps 1 - 3, the method for deleting the aggregation structure leaving the device includes a hierarchical node deletion mechanism and a vacancy marking mechanism, with the formula:
[0023] ,
[0024] ,
[0025] where, represents the th clustering set in Layer 0, is the training node of the device in Layer 0, represents the set difference operation, represents the state transition, that is, if the left - hand condition holds, then the right - hand operation is executed, represents assigning the right - hand value to the left - hand value.
[0026] In Steps 1 - 4, the method for adding the aggregation structure of newly - added devices includes: achieving the optimal clustering allocation of devices through pre - training accuracy difference calculation, optimal clustering selection, and special handling of empty clusters;
[0027] The formula for the pre - training accuracy difference calculation is:
[0028] ,
[0029] where, is the pre - training accuracy difference, is the average pre - training accuracy of the clustering before the device is added, is the average pre - training accuracy of the clustering after the device is added to the clustering, is the global average pre - training accuracy;
[0030] The formula for the optimal clustering selection is:
[0031] ,
[0032] where is the clustering index that makes the pre - training accuracy difference the largest;
[0033] The formula for the special handling of empty clusters is:
[0034] If , then set: ,
[0035] wherein is a preset negative constant, to ensure that empty clusters are preferentially assigned new devices.
[0036] Steps 1-5 include: dynamically supplementing vacant nodes by reselection of central nodes and marking the vacancies of departing devices;
[0037] The reselection of central nodes is performed using the following formula:
[0038] ,
[0039] wherein is the newly added device serves as the node representation of the newly selected central node at the th layer;
[0040] The vacancy of the departing device is marked using the following formula:
[0041] ,
[0042] When , cross-layer synchronization update is performed:
[0043] ,
[0044] wherein, refers to the th cluster at the th layer; and are respectively the node representations of the th device at the th layer and the node representations of the th device at the th layer.
[0045] Step 2 includes:
[0046] Step 2-1, the aggregating node collects the model parameters of each child node and simultaneously counts the time consumed by each child node in this aggregation, the model download transmission time, the model upload transmission time, the actual aggregation frequency in this time, calculates the average calculation time per training or aggregation of each child node and the actual calculation time consumed by the aggregating node in the th aggregation;
[0047] Subsequently, the exponential smoothing method is used to predict the average calculation time of each node in the th aggregation , the model distribution transmission time and the model upload transmission time , and predict the total transmission time of each child node in the th aggregation ;
[0048] Among them, the following formula is used to calculate the average calculation time of each child node per training or aggregation in this aggregation :
[0049] ;
[0050] The following formula is used to calculate the actual calculation time consumed by the aggregation node in the th aggregation :
[0051] ;
[0052] The following formula is used to predict the total transmission time of each child node in the th aggregation :
[0053] ;
[0054] Step 2-2, predict the calculation time of the th aggregation through the actual calculation time ; If this (the th) aggregation is the first aggregation or the currently predicted node is a newly emerged aggregation node due to the aggregation structure adjustment, at this time is the null value None (i.e., the state of no historical data), and the actual calculation time is used as the calculation time to predict the next time : If is not the null value None, and is less than or equal to , then the previously predicted calculation time is still used as the predicted calculation time for the next time. If is greater than , then is set to the smaller value of ; When Greater than When The calculation formula of is:
[0055] ,
[0056] Wherein Is a slow growth rate and can be set to a number slightly greater than 1.0;
[0057] Step 2-3, according to the expected aggregation calculation time Subtract the transmission time of the predicted child node To obtain the calculation time of the recommended child node , and then Divide by the average calculation time of the predicted child node And round down to obtain the aggregation frequency of the recommended child node ;
[0058] Step 2-4, based on the frequency allocation strategy of the remaining time, in the dynamic aggregation frequency adjustment, each child node first performs a local update, records the time consumption of this update , and accumulates the total calculation time And the number of updates (i.e., the current aggregation frequency); subsequently, predict the time of the next update through the exponential smoothing method , and then predict the total time after the update ; The child node decides whether to continue the update according to the relationship between the current number of updates and the recommended aggregation frequency ;
[0059] Among them, the calculation formula of the total time after the update Is:
[0060] ,
[0061] In the decision of whether to continue the update by the child node according to the relationship between the current number of updates and the recommended aggregation frequency , the child node performs the following judgment steps:
[0062] When it is detected that the current actual number of updates Is less than the recommended aggregation frequency , further verify whether the predicted total time Meets the first condition: , if the first condition is met, control the child node to continue to perform the model update operation;
[0063] When it is detected that the current actual number of updates Reaches or exceeds the recommended aggregation frequency , verify the predicted total time Whether the second condition is satisfied: Only when the second condition holds, are the child nodes allowed to perform additional model update operations;
[0064] The algorithm formula for the child nodes to determine whether to continue the update based on the relationship between the current update times and the recommended aggregation frequency is as follows:
[0065] ,
[0066] where is a boolean variable used to determine whether the child node continues to perform the next model update; is the appropriate timeout ratio, , is the preventive timeout ratio, .
[0067] In steps 2-1 and 2-4, the formula for the exponential smoothing method is:
[0068] ,
[0069] where is the time of the th prediction, is the time of the th update prediction, is the exponential smoothing coefficient.
[0070] In step 2-4, the formula for the frequency allocation strategy based on the remaining time is:
[0071] ,
[0072] ,
[0073] where is the estimated aggregation calculation time, is the predicted transmission time of the child node, is the calculation time of the recommended child node, is the average calculation time of the predicted child node, is the aggregation frequency of the recommended child node.
[0074] Step 3 includes;
[0075] Step 3-1, after initially constructing the aggregation structure, set the recommended aggregation frequency of all child nodes , the recommended calculation time , and recommend the aggregation frequency and the calculation time for the child nodes according to the method in step 2-3; During the subsequent update process of the child node, adaptively adjust its own aggregation frequency according to the method in Step 2-3 ;
[0076] Step 3-2, for the round of training, according to Step 1, adjust the aggregation structure from to After that, for the parent-child relationships deleted during the adjustment process, delete the recommended aggregation frequency of the aggregation node to the child node and the calculation time as well;
[0077] Step 3-3, for the parent-child relationships that have not been changed during the adjustment process, still retain the corresponding recommended aggregation frequency of the parent node and the child node and the calculation time , and conduct the next round of training based on the recommended aggregation frequency and the calculation time ; For the newly added parent-child relationships during the adjustment process, set the recommended aggregation frequency of all child nodes, the recommended calculation time , and before the child node has not received the recommended aggregation frequency and the calculation time from the aggregation node, upload the data only after updating (training or aggregating) once.
[0078] Step 4 includes:
[0079] When the node is in the th aggregation waiting process, use the calculation prediction time of this node as the standard to set a timeout abandonment ratio , , and set the timeout abandonment threshold to ; When the aggregation node waits for more than time and still does not receive the parameter data of the faulty child node (referring to some child nodes that cannot upload parameters due to network timeout, etc.), give up waiting and only use the trained model parameter data of each received child node for aggregation and subsequent operations;
[0080] Among them, the calculation formula of the timeout abandonment threshold is:
[0081] .
[0082] The present invention has the following beneficial effects: (1) The proposed dynamic adaptive hierarchical aggregation frequency adjustment method in the present invention effectively optimizes the distributed model training efficiency in the edge computing environment by real-time monitoring the computing time, transmission time, and aggregation frequency of each child node, and dynamically adjusting the aggregation strategy using exponential smoothing prediction technology. Compared with the traditional fixed aggregation frequency method, this method can adaptively adjust the aggregation rhythm according to the dynamic changes in device computing capabilities and network conditions, significantly reducing computational resource waste and communication bottleneck problems, and improving the overall efficiency of model training. It is particularly suitable for edge computing scenarios with heterogeneous computing capabilities and frequent network fluctuations. In addition, the present invention combines exponential smoothing prediction and dynamic threshold control mechanisms to avoid training delay problems caused by device performance differences while ensuring the model convergence speed.
[0083] (2) The proposed prediction-based aggregation frequency recommendation mechanism in the present invention dynamically determines whether to continue local updates or trigger aggregation by calculating the recommended aggregation frequency ( ) of each child node and combining timeout prevention strategies ( and ). Compared with the traditional uniform aggregation strategy, this method can intelligently adjust the aggregation rhythm according to device computing loads and communication delays, making full use of the computing capabilities of high-performance devices while avoiding low-speed devices from becoming training bottlenecks. This mechanism significantly improves resource utilization and reduces the overall time consumption of model training by dynamically balancing computational and communication overheads, and is particularly suitable for distributed learning tasks in non-independent and identically distributed (Non-IID) data environments.
[0084] (3) The proposed dynamic aggregation time prediction and adjustment method in the present invention ensures reasonable growth of aggregation time and avoids training instability problems caused by sudden computational or communication delays by introducing a slow growth ratio ( ) and exponential smoothing prediction technology. Compared with static aggregation time strategies, this method can adapt to the dynamic performance changes of edge devices, maximizing aggregation efficiency while ensuring training stability. This method is particularly suitable for large-scale edge computing clusters, effectively reducing the load pressure on the central node and improving the scalability and robustness of distributed model training.
[0085] (4) The proposed real-time monitoring-based adaptive aggregation decision mechanism in the present invention continuously tracks the computing time ( ), transmission time ( , ), and aggregation frequency ( ), and dynamically adjusts the training strategy in combination with the prediction model to ensure the efficiency and stability of the training process. Compared with the traditional method that relies on fixed thresholds or empirical values, this mechanism can make optimal decisions based on the real-time operating state, effectively reducing resource waste and communication conflicts during the training process, and providing an efficient and reliable solution for distributed machine learning in a dynamic edge computing environment. Brief Description of the Drawings
[0086] Figure 1 is the flowchart of the method of the present invention. Detailed Embodiments
[0087] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0088] As Figure 1 shown, an embodiment of the present invention provides a hierarchical federated learning optimization method in a dynamic edge computing environment, which dynamically adjusts the aggregation structure and aggregation frequency in the hierarchical model training framework. The embodiment of the present invention is verified in a simulated dynamic edge computing environment, setting 100 edge devices, with an initial online rate of 70%, the device computing power fluctuates dynamically (0.2 - 1.0 TFLOPS), the network topology is randomly generated (direct connection probability 0.2), and the channel bandwidth changes dynamically (10 - 100 Mbps). The MNIST dataset (Non-IID distribution, each device is assigned 2 classes of 200 - 300 samples) is used, and each device uses a CNN convolutional neural network model to train through an SGD optimizer (learning rate 0.1, batch size 64). The method includes the following steps:
[0089] Step 1, optimize the device aggregation topology through a dynamic clustering algorithm to adapt to device changes while maintaining stable device relationships;
[0090] The dynamic aggregation structure adjustment of hierarchical federated learning is designed to solve the problem of reduced training efficiency caused by the dynamic addition / removal of edge devices. It aims to optimize the device aggregation topology through a dynamic clustering algorithm, and adaptively optimize the aggregation topology structure while maintaining stable communication relationships between devices. During this process, the control node dynamically adjusts the clustering division according to the pre-training accuracy and communication delay of the devices to ensure the aggregation efficiency of each round of training. The specific adjustment is divided into three cases: deleting the departing device node, adding a new device node, and supplementing the vacant central node.
[0091] In a specific embodiment, the aggregation structure adjustment is divided into the following steps: First, based on the principle of minimizing the difference in pre-training accuracy and communication latency, the initial clusters are divided through a dynamic clustering algorithm, and the central nodes are selected; the departing devices are dynamically removed. If a device exits, its node is removed from the aggregation structure, and the relevant position is marked as empty; at the same time, new devices are dynamically added. The new devices select the optimal cluster to join according to the pre-training accuracy gain, and preferentially fill the empty clusters to improve the global training consistency; finally, the vacant central nodes are filled to be complete. The vacant central nodes are detected from bottom to top, re-elected according to the principle of minimizing communication latency, and the associated aggregation nodes are updated synchronously.
[0092] In a specific embodiment, based on the principle of minimizing the difference in pre-training accuracy and communication latency, the dynamic clustering algorithm is used to divide the initial clusters and determine the central nodes; the device changes are dynamically managed. If a device exits, the corresponding node is removed from the aggregation structure and its position is marked as vacant ( ); at the same time, new devices are dynamically admitted, enabling them to select the optimal cluster to join according to the pre-training accuracy gain, and preferentially filling the empty clusters to improve the global training consistency; finally, the vacant central nodes are detected from bottom to top, re-elected according to the principle of minimum communication latency, and the associated aggregation nodes are updated synchronously to ensure the structural integrity.
[0093] In a specific embodiment, the algorithm for using the dynamic clustering algorithm to divide the specific initial clusters is as follows:
[0094] ,
[0095] where , is the th cluster of the th layer, , which contains several nodes, and their common aggregation node is , is the central node (aggregation node); N refers to the number of edge devices included, and are the representations of the th device in the th layer and the th device in the th layer. The same device can be a node in different layers; H is the total number of layers of the aggregation structure, refers to the and communication distance of the node, is the pre-training accuracy of the th device, is the average pre-training accuracy of the th cluster of the th layer, Refers to the weight coefficient, which is used to balance the computational weights between the pre-training accuracy difference and the communication delay;
[0096] In a specific embodiment, the tree structure of the round of training is
[0097] ,
[0098] where ;
[0099] In a specific embodiment, the dynamic device departure deletion method includes: a hierarchical node deletion mechanism and a vacancy marking mechanism;
[0100] The algorithms of the hierarchical node deletion mechanism and the vacancy marking are:
[0101] ,
[0102] ,
[0103] In the formula, represents the th clustering set of the 0th layer, is the training node of device in the 0th layer, represents the set difference operation; is the highest layer where device appears, represents the th clustering of the th layer, is the central node of the clustering , represents the vacancy mark generated after deleting device .
[0104] In a specific embodiment, the dynamic new device admission includes: achieving the optimal clustering allocation of devices through pre-training accuracy difference calculation and optimal clustering selection;
[0105] The calculation formula of the pre-training accuracy difference calculation is:
[0106] ,
[0107] In the formula, is the pre-training accuracy difference is the average pre-training accuracy of the clustering before device joins, is the average pre-training accuracy of the clustering after joins, is the global average pre-training accuracy;
[0108] The calculation formula for the optimal clustering selection is as follows:
[0109] ,
[0110] In the formula is the clustering index that maximizes the maximum clustering; is the maximum clustering;
[0111] The calculation formula for the special processing of empty clusters is as follows:
[0112] If , then set ;
[0113] In the formula is a preset negative constant to ensure that empty clusters are preferentially assigned new devices;
[0114] In a specific embodiment, the vacant nodes are re-elected and the associated aggregation nodes are updated synchronously to ensure structural integrity, including: reselecting the central node according to the principle of minimum communication delay and marking the vacant positions of the departing devices to achieve dynamic replenishment of the vacant nodes.
[0115] The formula for reselecting the central node is as follows:
[0116] ,
[0117] where is the newly added device as the node representation of the newly selected central node at the th layer;
[0118] The formula for replacing the vacant mark is as follows:
[0119] ,
[0120] In the formula represents the vacant mark generated after the device leaves;
[0121] The formula for cross-layer synchronous update is (when ):
[0122] ,
[0123] In the formula is the total number of layers of the aggregation structure;
[0124] In step 1, 10 clusters are initially formed (10 devices per cluster, average latency of 45 ms). When the performance of device A (computing power 0.8 → 0.3) and device B (bandwidth 50 → 15 Mbps) degrades, an adjustment is triggered: the aggregation frequency of device A is increased from 1 to 2 to reduce the burden, and device B is moved to a neighboring cluster to reduce the latency from 80 ms to 35 ms. At the same time, the original aggregation frequency of 8 pairs of parent - child relationships is retained (3 times). This adjustment reduces the training time by 18% (120 s → 98 s), improves the accuracy by 3.4 percentage points (58.3% → 61.7%), and avoids the reconstruction overhead by maintaining a stable aggregation structure, significantly improving the system efficiency and stability.
[0125] In step 2, based on historical performance data, predict and recommend the optimal training frequency to achieve resource - aware adaptive training scheduling. The steps are as follows:
[0126] In step 2 - 1, the aggregation node collects the model parameters of each child node and simultaneously counts the time consumed by each child node in this aggregation , the model distribution transmission time , the model upload transmission time , and the actual aggregation frequency of this time , and calculates the average calculation time per training or aggregation of each child node in this aggregation and the actual calculation time consumed by the aggregation node in the th aggregation ;
[0127] Subsequently, using the exponential smoothing method, predict the average calculation time , the model distribution transmission time , and the model upload transmission time of each node in the th aggregation, and calculate the total transmission time of each predicted child node in the th aggregation;
[0128] Among them, the following formula is used to calculate the average calculation time per training or aggregation of each child node in this aggregation :
[0129] ,
[0130] In the formula, is the average calculation time per training or aggregation of each child node in this aggregation, is the time consumed by each child node in this aggregation, is the actual aggregation frequency for this time;
[0131] The aggregation node is calculated using the following formula At the th aggregation, the actual calculation time consumed :
[0132] ,
[0133] In the formula, is the actual calculation time consumed by the aggregation node for this aggregation, is the model distribution transmission time, is the model upload transmission time;
[0134] The total transmission time of each sub-node at the th aggregation is predicted using the following formula :
[0135] ,
[0136] Step 2-2, predict the calculation time of the th aggregation through the actual calculation time ; If this (the th) aggregation is the first aggregation or the currently predicted node is a newly emerged aggregation node due to the aggregation structure adjustment, at this time is a null value None (i.e., in the state of no historical data), and the actual calculation time is used as the calculation time to predict the next time: If is not a null value None, and is less than or equal to , then the previously predicted calculation time is still used as the predicted calculation time for the next time. If is greater than , then is set to the smaller value of and ; When is greater than , the calculation formula of is:
[0137] ,
[0138] where is the slow growth ratio, which can be set to a number slightly greater than 1.0;
[0139] Step 2-3: Based on the predicted aggregation calculation time Subtract the transmission time of the predicted child node to calculate the calculation time of the recommended child node Then divide it by the average calculation time of the predicted child node and round down to calculate the aggregation frequency of the recommended child node ;
[0140] Step 2-4: Based on the frequency allocation strategy of the remaining time, in the dynamic aggregation frequency adjustment, each child node first performs a local update, records the time consumed for this update and accumulates the total calculation time and the number of updates (i.e., the current aggregation frequency); Subsequently, predict the time of the next update through the exponential smoothing method and then predict the total time after the update ; The child node decides whether to continue the update according to the relationship between the current number of updates and the recommended aggregation frequency ;
[0141] Among them, the calculation formula for the total time after the update is:
[0142] ,
[0143] In the process where the child node decides whether to continue the update according to the relationship between the current number of updates and the recommended aggregation frequency , the child node performs the following judgment steps:
[0144] When it is detected that the current actual number of updates is less than the recommended aggregation frequency , further verify whether the predicted total time satisfies the first condition: , if it satisfies the first condition, control the child node to continue to perform the model update operation;
[0145] When it is detected that the current actual number of updates reaches or exceeds the recommended aggregation frequency , verify whether the predicted total time satisfies the second condition: , only when the second condition holds, is the child node allowed to perform an additional model update operation;
[0146] The algorithm formula for the child node to decide whether to continue the update according to the relationship between the current number of updates and the recommended aggregation frequency is:
[0147] ,
[0148] Among them, is a Boolean variable used to determine whether the child node continues to perform the next model update; is the appropriate timeout ratio, , is the preventive timeout ratio, .
[0149] In steps 2-1 and 2-4, the formula of the exponential smoothing method is:
[0150] ,
[0151] where is the time of the th prediction, is the time of the th update prediction, is the exponential smoothing coefficient.
[0152] In step 2-4, the formula of the frequency allocation strategy based on the remaining time is:
[0153] ,
[0154] ,
[0155] where is the expected aggregation calculation time, is the transmission time of the predicted child node, is the calculation time of the recommended child node, is the average calculation time of the predicted child node, is the aggregation frequency of the recommended child node.
[0156] In step 2, the system dynamically adjusts the training frequency according to the real-time computing power of the device: when the computing power of device C drops from 0.6 to 0.2, its aggregation frequency is increased from 1 to 3 through exponential smoothing prediction (β = 0.9), and the training time is reduced by 40%; while device D with stable computing power maintains the frequency of 1. This strategy shortens the global training time by 13% (98s → 85s), and effectively optimizes the resource utilization rate and avoids the problem of lagging nodes while ensuring that low-computing-power devices continue to contribute data (the accuracy is stable at 61.5%).
[0157] Step 3, the device dynamically adjusts the training rounds according to its own resource status to balance the training intensity and time efficiency. Including:
[0158] Step 3-1, after initially constructing the aggregation structure, set the recommended aggregation frequency of all child nodes, and the recommended calculation time , recommend the aggregation frequency for child nodes according to the method proposed in Step 2-3 and the calculation time . During the subsequent update process of the child nodes, adaptively adjust their own aggregation frequency according to the method proposed in Step 2-4 .
[0159] Step 3-2, for each round , after adjusting the aggregation structure from to according to Step 1, for the parent-child relationships deleted during the adjustment process, delete the recommended aggregation frequency of the aggregation node for the child nodes and the calculation time as well.
[0160] Step 3-3, for the parent-child relationships that have not been changed during the adjustment process, still retain their corresponding recommended aggregation frequency and the calculation time , and conduct the next round of training accordingly; for the newly added parent-child relationships during the adjustment process, set the recommended aggregation frequency of all child nodes, recommend the calculation time , and upload the data only after updating (training or aggregating) once when the child nodes have not received the recommended aggregation frequency and the calculation time from the aggregation node.
[0161] In Step 3, the system intelligently regulates the number of training times based on the actual computing power of the device: the efficient device E (computing power 0.7) completes 1 additional update (total 3 times) after meeting the recommended frequency 2 times, improving the accuracy by 1.2%; while the inefficient device F (computing power 0.3) terminates early when the prediction times out (110s > ρ_p × 90s) to avoid dragging down the overall situation. This strategy achieves an 8% optimization of the training time (85s → 78s), maintains the accuracy balance (global accuracy is improved to 62.7%) through the intelligent termination mechanism while ensuring sufficient training for the efficient device, demonstrating excellent flexibility and resource allocation capabilities.
[0162] Step 4: Establish a hierarchical timeout fault tolerance mechanism to handle abnormal device conditions while ensuring the continuity of training.
[0163] During the th aggregation waiting process of the node, this predicted value can be used as a standard to set a timeout abandonment ratio , and set the timeout abandonment threshold to ; when the aggregation node waits for more than time and still does not receive data from some child nodes, give up waiting and only use the received model data for aggregation and subsequent operations.
[0164] The timeout abandonment threshold has the following calculation formula:
[0165] ,
[0166] In step 4, when the computing power of device G suddenly drops to 0, the system sets an 88s timeout threshold (ρ_t = 1.1) based on the predicted aggregation time of 80s. After the timeout, the data of this device is immediately abandoned and the remaining 99 devices are aggregated. This mechanism shortens the training time by 10% (78s → 70s). Although it causes a slight decrease in accuracy of 0.3% (62.7% → 62.4%), it effectively avoids the problem of infinite waiting. This design significantly improves the training efficiency while ensuring that the loss of model accuracy is controllable (<0.5%), and at the same time demonstrates a strong fault tolerance ability, and can reliably prevent the risk of training stagnation caused by single-point failures.
[0167] The present invention provides a hierarchical federated learning optimization method in a dynamic edge computing environment. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A hierarchical federated learning optimization method in a dynamic edge computing environment, characterized in that: For each round of hierarchical collaborative distributed model parameter training and aggregation, the following steps are executed: Step 1, optimize the device aggregation topology of the collaborative training participants composed of terminal devices and aggregation nodes through a dynamic clustering algorithm, adapting to device changes while maintaining stable device relationships; the terminal devices include smart phones and Internet of Things terminals; the nodes include edge servers and gateways; Step 2, predict and recommend the optimal training frequency based on historical performance data to achieve resource-aware adaptive training scheduling; Step 3, the devices dynamically adjust the number of training rounds according to their own resource conditions, balancing training intensity and time efficiency; Step 4, establish a hierarchical timeout fault tolerance mechanism to handle abnormal device conditions while ensuring training continuity; Step 1 includes: Step 1-1, optimize the device aggregation topology based on the dynamic clustering algorithm, and cluster and divide the edge devices according to the principle of minimizing pre-training accuracy and communication delay; The dynamic clustering algorithm includes the following steps: calculate the pre-training accuracy difference and communication delay between devices; select the central node with the goal of minimizing the total communication delay within the cluster; preferentially allocate new devices to the cluster with the largest optimization amount of pre-training accuracy difference; The formula for selecting and aggregating different devices is: in is the jth cluster of the hth layer, h∈[0,H-1], j∈[1,N], Contains nodes in The common aggregation node of the nodes in is is the central node; N refers to the number of edge devices included, and are the node representation of the i-th device at the h-th layer and the node representation of the k-th device at the h-th layer respectively; H is the total number of layers of the aggregation structure, means and The communication distance, acc i refers to the pre-training accuracy of the i-th device, refers to the average pre-training accuracy of the jth cluster of the hth layer, and λ refers to the weight coefficient; Step 1-2, after each round of training and before the next round of training, dynamically adjust the aggregation structure: when the rth round of training is completed and the root aggregation node sends the model parameters obtained from the training to the control node of the central management unit responsible for coordinating the training process, the control node compares the device set of the r+1th round with the rth round: if the device has not changed, the aggregation structure S is maintained r unchanged; otherwise, statistics exit the device set {V out } and the newly added device collection {V in }, and generate the next round of aggregation structure S based on the following formula r+1 : Where S r Represents the tree structure of the rth round of training; Step 1-3, delete the aggregation structure of the leaving device: For the leaving device V i , get device V i In the aggregate structure S r+1 The highest occurrence level in h max , if h max = 0, then directly remove the device V i Nodes at level 0 If h max ≥1, then from layer 0 to layer h max layer, the device V i The corresponding node in the h layer represents Mark as empty i , and update the relevant cluster center nodes; empty i Indicates device V i The empty mark created after deletion; Step 1-4, add the aggregation structure of the newly added device: For the newly added device V q , traverse all clusters in layer 0 Computing Device V q The optimization gain delta of the average pre-training accuracy of clusters after adding each cluster j , and select the cluster with the largest gain Join if clustering exists is empty, p∈[1,N], then The average pre-training accuracy is set to 0 or a negative value; Step 1-5, the method for adjusting the aggregation structure of the vacant aggregation node: Detect empty aggregation nodes layer by layer from bottom to top i , if clustering The central node If the central device V is vacant, the dynamic clustering algorithm is used to reselect the central device V. k , and synchronously update the associated nodes of each layer.
2. The method according to claim 1, characterized in that In Step 1-3, the method for deleting the aggregation structure of the departing device includes a hierarchical node deletion mechanism and a vacant marking mechanism, and the formula is: in, represents the jth cluster set at layer 0, For device V i In the training node at level 0, \ represents the set difference operation, Indicates state transition, ← means assigning the right value to the left value.
3. The method according to claim 2, characterized in that In Step 1-4, the method for adding the aggregation structure of the newly added device includes: realizing the optimal clustering allocation of the device through pre-training accuracy difference calculation, optimal cluster selection, and special processing of the empty cluster; The formula for calculating the pre-training accuracy difference is: Among them, delta j is the pre-training accuracy difference, For device V i Clustering before joining The average pre-training accuracy of For device V i join in After The average pre-training accuracy of clustering, acc allavg is the global average pre-training accuracy; The calculation formula for the optimal cluster selection is: where j * In order to make The difference in pre-training accuracy delta j The largest cluster index; The calculation formula for the special processing of the empty cluster is: if Then set: Where α is a preset negative constant.
4. The method according to claim 3, characterized in that Step 1-5 includes: realizing the dynamic supplement of the vacant node by reselecting the central node and marking the departing device as vacant; The following formula is used to reselect the central node: in For the newly added device V q As the newly selected central node, the node representation at the hth layer; The following formula is used to mark the departing device as vacant: When h < H - 1, perform cross-layer synchronous update: in, refers to the jth cluster at the h+1th layer; and They are the node representation of the i-th device at the h+1 layer and the node representation of the k-th device at the h+1 layer respectively.
5. The method according to claim 4, characterized in that Step 2 includes: Step 2-1: Aggregation node collects each child node The model parameters of each child node are counted The time consumed in this aggregation Model transmission time Model upload transfer time The actual aggregation frequency f i,h , calculate each child node The average computation time per training or aggregation in this aggregation and aggregation nodes The actual computing time consumed in the mth aggregation Then, the exponential smoothing method is used to predict the average computing time of each node in the m+1th aggregation Model transmission time and model upload transfer time And predict each child node Total transmission time in the m+1th aggregation Among them, the following formula is used to calculate each child node The average computation time per training or aggregation in this aggregation The following formula is used to calculate the aggregation node The actual computing time consumed in the mth aggregation Use the following formula to predict each child node Total transmission time in the m+1th aggregation Step 2-2, by actually calculating the time Predict the computation time of the m+1th aggregation If this aggregation is the first aggregation or the currently predicted node is a newly appeared aggregation node due to the adjustment of the aggregation structure, If it is None, the actual calculation time will be As the prediction of the next calculation time if is not None, and Less than or equal to The calculation time of the last prediction is still used. As the next prediction calculation time if Greater than Then Set to and The smaller value of Greater than hour, The calculation formula is: where ρ g is a slow growth rate; Step 2-3, calculate the time based on the estimated aggregation Subtract the transmission time of the predicted child node Get the calculation time of the recommended child node Then Divided by the average computation time of the prediction child nodes And round down to get the aggregate frequency rf of the recommended child node i,h ; Step 2-4, based on the frequency allocation strategy of the remaining time, in the dynamic aggregation frequency adjustment, each child node first performs a local update and records the time consumed by this update. And the total calculation time and the number of updates f i,h ; Then, the time of the next update is predicted by exponential smoothing method Then predict the total time after update The child node is based on the current update number and the recommended aggregation frequency rf i,h The relationship determines whether to continue updating; Among them, the total time after update The calculation formula is: In the child node, according to the current update number and the recommended aggregation frequency rf i,h The relationship between the child node and the node determines whether to continue updating. The child node performs the following judgment steps: When the current actual update number f is detected i,h Less than the recommended aggregation frequency rf i,h To further verify the predicted total time Whether the first condition is met: If the first condition is met, the control subnode continues to perform the model update operation; When the current actual update number f is detected i,h Reach or exceed the recommended aggregation frequency rf i,h The total time for verification prediction is Whether the second condition is met: Only when the second condition is met, the child node is allowed to perform additional model update operations; The child node is based on the current update number and the recommended aggregation frequency rf i,h The relationship between determines whether to continue updating. The algorithm formula is: Among them, Continue is a Boolean variable used to determine whether the child node continues to execute the next model update; ρ a is the appropriate timeout ratio, ρ a >1.0,ρ p To prevent the timeout ratio, ρ p <1.
0.
6. The method according to claim 5, characterized in that In Step 2-1 and Step 2-4, the formula for the exponential smoothing method is: t m =βt m +(1-β)t m-1 , where t m-1 is the predicted time of the m-1th time, t m is the time of the mth update of the forecast, and β is the exponential smoothing coefficient.
7. The method according to claim 6, characterized in that In Step 2-4, the formula for the frequency allocation strategy based on the remaining time is: in, To estimate the aggregation calculation time, To predict the transmission time of a child node, is the computation time of the recommended child node, is the average computation time of the predicted child nodes, rf i,h is the aggregation frequency of the recommended child nodes.
8. The method according to claim 7, characterized in that Step 3 includes; Step 3-1: After initially building the aggregation structure, set the recommended aggregation frequency rf for all child nodes i,h =1, recommended calculation time Recommend aggregation frequency rf for child nodes according to the method of steps 2-3 i,h and calculation time In the subsequent update process of the child node, the aggregation frequency f is adaptively adjusted according to the method of steps 2-3 i,h ; Step 3-2: For the rth round of training, follow step 1 to convert the aggregate structure from S r Adjust to S r+1 After that, for the parent-child relationship that was deleted during the adjustment process, the recommended aggregation frequency rf of the aggregation node to the child node is i,h and calculation time Also deleted; Step 3-3: For the parent-child relationship that has not been changed during the adjustment process, the recommended aggregation frequency rf of the parent node and the child node is still retained. i,h and calculation time And according to the recommended aggregation frequency rf i,h and calculation time Carry out the next round of training; for the new parent-child relationship added during the adjustment process, set the recommended aggregation frequency rf of all child nodes i,h =1, recommended calculation time The child node does not receive the aggregation frequency rf recommended by the aggregation node i,h and calculation time Previously, data was uploaded only after updating once.
9. The method according to claim 8, characterized in that Step 4 includes: When the node is waiting for the m+1th aggregation, the node's calculation prediction time is used As a standard, set a timeout abandonment ratio ρ t , ρ t ≥1.0, set the timeout abandonment threshold to When the aggregation node waits for more than If the parameter data of the faulty sub-node is still not received within the time, the system will give up waiting and only use the trained model parameter data of each sub-node that has been received for aggregation and subsequent operations; Wherein, the timeout abandonment threshold The calculation formula is:
Citation Information
Patent Citations
Edge calculation method based on AI
CN119046010A