A Method for Accelerating Federated Learning Training

Through dynamic hierarchical decision-making algorithm and two-stage aggregation method, the collaborative training of edge devices and edge servers, the problem of low efficiency of federated learning training is solved, and a more efficient and accurate training process is achieved.

CN115408151BActive Publication Date: 2025-06-17HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211014211.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-23
Publication Date
2025-06-17
Estimated Expiration
2042-08-23

AI Technical Summary

Technical Problem

While ensuring accuracy, existing federated learning training methods are difficult to effectively improve training efficiency, especially when facing edge device heterogeneity and unstable network environments.

Method used

Using dynamic hierarchical decision-making algorithm and two-stage aggregation method, edge devices and edge servers build front-end models and back-end models respectively, local model parameters are obtained through collaborative training, and intermediate aggregation is performed on edge servers to reduce network transmission and computing pressure.

Benefits of technology

It improves the efficiency and accuracy of federated learning training, effectively solves the heterogeneous resource problems of edge device and the instability of network environment, and reduces the number of global model iterations and updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408151B_ABST
    Figure CN115408151B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for accelerating federated learning training. The method includes: edge devices construct local models according to initial model parameters, and calculate the training tasks of the edge devices and the edge server according to a dynamic hierarchical decision algorithm; the edge devices and the edge server respectively construct a front-end model and a back-end model according to the training tasks, and co-train the front-end model and the back-end model to obtain local model parameters and send them to the edge server; the edge server performs intermediate aggregation according to the local model parameters sent by each edge device, obtains intermediate model parameters and sends them to the central cloud; the central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and iteratively updates the global model until the global model converges. The beneficial effects of the present invention: while ensuring the accuracy of federated learning training, the efficiency of federated learning training is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the Internet of Things, and in particular, to a method for accelerating federated learning training. Background Art

[0002] With the advent of the big data era, data privacy has received increasing attention, and users are reluctant to upload sensitive data to cloud servers. In this context, federated learning has been proposed to solve data security problems. Federated learning is a distributed artificial intelligence technology capable of privacy protection, which sinks the operations to edge devices. The edge devices are responsible for data storage and the calculation of edge models. The client only needs to upload relevant information such as the gradients of the edge models, and the central server updates the global model based on the edge gradients generated on the clients. Federated learning enables edge devices to store sensitive data locally while collaborating with each other to build a global model.

[0003] However, with the increase in data, the training efficiency requirements for federated learning are gradually increasing. Currently, three main methods are adopted to improve the training efficiency of federated learning: 1. Optimize the communication overhead in federated learning. The client only needs to upload some of the local updates obtained to the central server, rather than uploading the entire model; 2. By selecting devices, mitigate the straggler problem in the synchronous aggregation algorithm. For example, adopt a greedy-based protocol mechanism to select the client with the highest score for aggregation through a reward value; 3. Reduce the waiting time and improve the training efficiency through an asynchronous aggregation algorithm. The so-called asynchronous aggregation algorithm means that after each device uploads its local model, the central node will immediately aggregate according to the parameters of the device without waiting for other devices to finish. Its advantage is that it can reduce the synchronous waiting time of the devices. However, due to the heterogeneous problems of edge devices and the instability of edge network environments existing in federated learning, using the above methods for federated learning training will lead to a decrease in accuracy and uncertainty in training efficiency, and cannot effectively improve the training efficiency of federated learning. Summary of the Invention

[0004] The problem solved by the present invention is how to improve the training efficiency while ensuring the training accuracy of federated learning.

[0005] To solve the above problems, the present invention provides a method for accelerating federated learning training, including the following steps:

[0006] Step S100: The central cloud constructs a global model and initial model parameters corresponding to the global model, and sends the initial model parameters to the edge server.

[0007] Step S200: The edge server receives the initial model parameters and sends the initial model parameters to the edge devices.

[0008] Step S300: The edge device constructs a local model according to the initial model parameters, and calculates the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm, where the training task is the hierarchical ratio of the local model.

[0009] Step S400: The edge device and the edge server respectively construct a front-end model and a back-end model according to the training tasks. The edge device and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters, and send the local model parameters to the edge server.

[0010] Step S500: The edge server performs intermediate aggregation on the local model parameters sent by each edge device to obtain intermediate model parameters, and sends the intermediate model parameters to the central cloud.

[0011] Step S600: The central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges.

[0012] Thus, the central cloud is first initialized, and a global model and the initial model parameters corresponding to the global model are constructed, and the purpose and tasks of federated learning training are set. The initial model parameters are sent to the edge server, the edge server receives the initial model parameters, and sends the initial model parameters to the edge devices to ensure the normal operation of the abnormal detection model construction process. The edge devices construct local models according to the initial model parameters, and calculate the training tasks of the edge devices and the edge server according to the dynamic hierarchical decision algorithm. Among them, the training task is the hierarchical ratio of the local model. The method of hierarchical training can effectively solve the problem of heterogeneous resources of edge devices, provide a solution to the problem of insufficient computing power of edge devices in classical federated learning. At the same time, the dynamic hierarchical decision algorithm is adopted to monitor and calculate the hierarchical training tasks of edge devices and edge servers in real time. If abnormal situations or low training efficiency are found, timely adaptive adjustments can be made, and hierarchical decisions are given for specific cloud-edge environments to adapt to the changing cloud-edge environment. The edge devices and the edge server construct the front-end model and the back-end model according to the training tasks respectively. The edge devices and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters. The front-end model and the back-end model are constructed through reasonable hierarchical training tasks and jointly trained. Since some model parameters of the back-end model are deployed on the edge server side, the network transmission can be reduced to a certain extent, ensuring less network overhead. At the same time, the computing speed and computing accuracy of the edge devices are improved. The edge server performs intermediate aggregation on the local model parameters sent by each edge device to obtain intermediate model parameters. The two-stage aggregation method is adopted, and the edge server is used to perform the first aggregation of model parameters to reduce the computing pressure of the central cloud model aggregation, increase the computing accuracy, reduce the amount of data transmitted to the central cloud node at the same time, and reduce the probability of network congestion. The central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge devices, and iteratively updates the global model until the global model converges to obtain the trained global model. The aggregated central model parameters of the edge server are used to perform the aggregation update of the global model again, effectively reducing the number of times of iterative update of the global model. The two-stage aggregation method is adopted to perform the aggregation calculation of model parameters on the edge server and the central cloud respectively, avoiding the problems of unstable data transmission and crowded network channel transmission caused by the edge devices directly communicating with the central cloud frequently using the wide area network, effectively improving the calculation accuracy of model parameters, reducing the computing pressure of the central cloud, and then improving the efficiency and accuracy of model training in federated learning.

[0013] Optionally, the calculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes:

[0014] Obtain the average training time of a single sample of the local model in the edge device, compare the average training time of the single sample with the historical average training time. If the average training time of the single sample is greater than a preset multiple of the historical average training time and lasts for a preset number of times, then according to the dynamic hierarchical decision algorithm, recalculate the training tasks of the edge device and the edge server.

[0015] Optionally, the recalculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes:

[0016] Obtain the status information of the edge device, and according to the status information, use the dynamic hierarchical policy based on reinforcement learning to calculate the training tasks of the edge device and the edge server at the next moment. Wherein, the status information includes the total training time in a single round of edge device training, the training time of the edge device, the network transmission time, the proportion of the edge server training time in the total training time, and the training hierarchical policy at this moment.

[0017] Optionally, the edge device and the edge server respectively construct a front-end model and a back-end model according to the training tasks, and the edge device and the edge server jointly train the front-end model and the back-end model according to the data set. Obtaining the local model parameters includes:

[0018] The edge device performs hierarchical processing on the local model according to the training task, constructs the front-end model, and sends the hierarchically processed local model to the edge server, where the edge server constructs the back-end model according to the hierarchically processed local model;

[0019] The edge device trains the front-end model according to the data set to obtain intermediate results, and sends the intermediate results to the edge server, where the edge server trains the back-end model according to the intermediate results to obtain gradients, and sends the gradients to the edge device;

[0020] The edge device trains the front-end model according to the gradients to obtain local model parameters.

[0021] Optionally, before the edge server trains the back-end model according to the intermediate results, it further includes:

[0022] Score each of the training tasks according to a preset scoring mechanism, sort each of the training tasks in descending order according to the scoring results, and set a priority queue;

[0023] According to the priority queue, use the GPU to train the back-end model one by one.

[0024] Optionally, the step of training the backend model one by one using the GPU according to the priority queue includes:

[0025] When the GPU memory is insufficient, schedule the remaining backend models in the priority queue to the CPU for training;

[0026] After the training of the backend model in the GPU is completed, use the Watch mechanism to notify the CPU, and schedule the backend model in the CPU to the GPU for training.

[0027] Optionally, the edge device and the edge server co-train the front-end model and the backend model according to the data set to obtain local model parameters, including:

[0028] The edge device calculates the number of device training times according to the device loss rate. The edge device and the edge server co-train the front-end model and the backend model according to the number of device training times, obtain the local model parameters, and calculate the device training time. Send the device training time and the local model parameters to the edge server, where the device loss rate is initialized and preset by the central cloud and sent to the edge server, and the edge server sends the device loss rate to the edge device.

[0029] Optionally, the edge server performs intermediate aggregation according to the local model parameters sent by each edge device to obtain intermediate model parameters, and further includes:

[0030] The edge server calculates the number of intermediate aggregation times according to the server loss rate, performs intermediate aggregation according to each local model parameter and the number of intermediate aggregation times to obtain the intermediate model parameters, calculates the intermediate aggregation time, and sends the intermediate model parameters and the intermediate aggregation time to the central cloud, where the server loss rate is initialized and preset by the central cloud.

[0031] Optionally, the central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges, including:

[0032] The central cloud updates the global model according to the intermediate model parameters sent by each edge server, calculates the total aggregation time according to the intermediate aggregation time, calculates the optimal device loss rate and the optimal server loss rate according to the total aggregation time, and sends the model parameters of the updated global model, the optimal device loss rate and the optimal server loss rate to each edge server, and returns to execute the step of the edge server sending the device loss rate to the edge device, and iteratively updates the global model until the global model converges.

[0033] Optionally, calculating the optimal device loss rate and the optimal server loss rate according to the total aggregation time includes:

[0034] Construct a combinatorial optimization problem according to the total aggregation time, the device loss rate and the server loss rate;

[0035] Use the simulated annealing algorithm to iteratively solve the combinatorial optimization problem to obtain the optimal device loss rate and the optimal server loss rate. Description of the Drawings

[0036] Figure 1 Flow chart of the federated learning training acceleration method according to the embodiment of the present invention Figure 1 ;

[0037] Figure 2 Schematic diagram of the computing resource allocation strategy according to the embodiment of the present invention;

[0038] Figure 3 Flow chart of the federated learning training acceleration method according to the embodiment of the present invention II;

[0039] Figure 4 Schematic diagram of the two-stage aggregation principle when k1 = 2 and k2 = 2 according to the embodiment of the present invention;

[0040] Figure 5 Schematic diagram of the simulated annealing algorithm process according to the embodiment of the present invention;

[0041] Figure 6 Schematic diagram of the cloud-edge new connection platform hierarchical architecture according to the embodiment of the present invention;

[0042] Figure 7 Schematic diagram of the federated learning component structure according to the embodiment of the present invention. Detailed Embodiments

[0043] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be given with reference to the accompanying drawings.

[0044] Refer to Figure 1As shown in the figure, a method for accelerating federated learning training according to an embodiment of the present invention includes:

[0045] Step S100: The central cloud constructs a global model and initial model parameters corresponding to the global model, and sends the initial model parameters to the edge server.

[0046] Specifically, the central cloud is initialized to confirm the tasks and goals of model training. The central cloud can use a neural network model to construct an initial global model, such as a Frontend model, a TCN model, etc., and obtain the corresponding global model parameters, and broadcast the corresponding initialized model parameters to all edge servers participating in the federated learning model training. The central cloud and the edge server are connected through methods such as Bluetooth and wireless local area network to achieve efficient data transmission.

[0047] Step S200: The edge server receives the initial model parameters and sends the initial model parameters to the edge device.

[0048] Specifically, in the cloud-edge scenario, edge devices are scattered as participants everywhere, so the aggregation node often needs to be deployed in a remote data center with a public network environment. The communication between the edge device and the aggregation server needs to pass through a wide area network (WAN), which brings problems such as transmission delay and network stability, and is not conducive to effective aggregation. The edge server is deployed close to the edge device to provide partial services for the edge device. The edge server and the edge device are connected through a high-speed local area network to achieve efficient task offloading and low-latency services. The high-speed local area network has a high transmission rate, high reliability and security, which can effectively improve the security and stability of data transmission and increase the federated learning training efficiency.

[0049] Step S300: The edge device constructs a local model according to the initial model parameters, and calculates the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm, where the training task is the hierarchical ratio of the local model.

[0050] Specifically, the edge device receives the initial model parameters of the global model forwarded by the edge server and constructs a local model corresponding to the global model according to the initial model parameters. In classical federated learning, due to the gradual increase in the amount of data, the problem of insufficient computing power of edge devices appears, which affects the computing efficiency of federated learning. Therefore, the embodiment of the present invention adopts a dynamic hierarchical decision algorithm based on reinforcement learning (DDPG-SL algorithm). First, a ratio is defined using the reinforcement learning algorithm indicating device c at time t kThe computational workload ratio of the deployed model (i.e., the workload ratio of the Frontend model to the entire model workload), this value is between 0 and 1, and this value cannot be 0 because the data cannot be moved. Therefore, at least one layer of neural network must be retained on the edge device and cannot be all transferred to the edge server e; if this value is 1, it means that all training is performed on device c k executed on, that is, this ratio represents the hierarchical training task of the current neural network and is trained according to this training task. Secondly, the training tasks of the edge device and the edge server are monitored and calculated in real time, that is, in a hierarchical training method, the complete model training task is split into two parts, the front end and the back end, based on the neural network layer, and are respectively deployed on the device side and the edge server side, making full use of the data parallelism characteristics of federated learning and the model parallelism characteristics in hierarchical training, so as to achieve an efficient hybrid parallel training scheme. According to the DDPG-SL algorithm, it can be monitored in real time whether the training tasks of the edge device and the edge server meet the requirements of federated learning training at this time. If there are abnormalities or low efficiency, timely adaptive adjustments can be made.

[0051] Step S400: The edge device and the edge server respectively construct a front-end model and a back-end model according to the training task, and the edge device and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters, and send the local model parameters to the edge server.

[0052] Specifically, according to the training task calculated by the DDPG-SL algorithm, a front-end model is constructed on the edge device and a back-end model is constructed on the edge device. The edge device obtains the data sent by the Internet of Things device through wireless communication and constructs a data set for federated learning training. Since the data in the edge device cannot be moved, the front-end model located in the edge device and the back-end model located in the edge server are jointly trained. The front-end model is trained according to the data set, and the obtained training result is sent to the edge server for the training of the back-end model. After obtaining the training result, it is returned to the front-end model, and the front-end model is trained according to the training result of the back-end model. Finally, the complete local model parameters are obtained, and the edge device sends the local model parameters to the edge server for the model aggregation step.

[0053] Step S500: The edge server performs intermediate aggregation according to the local model parameters sent by each edge device to obtain intermediate model parameters, and sends the intermediate model parameters to the central cloud.

[0054] Specifically, the edge server receives the local model parameters sent by each edge device after training. The edge server can use aggregation algorithms such as FedAVG aggregation to aggregate the models, obtain the intermediate model parameters, and avoid the problem of data islands. The intermediate model parameters are sent to the central cloud for global model aggregation.

[0055] Step S600: The central cloud updates the global model according to the intermediate model parameters sent by each of the edge servers, and sends the model parameters of the updated global model to each of the edge servers, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges.

[0056] Specifically, the central cloud receives the intermediate model parameters sent by multiple edge servers, performs calculation processing on the intermediate model parameters, such as average value calculation and other methods, obtains the optimal model parameters, updates the global model according to the optimal model parameters, and then sends the optimal model parameters to each of the edge servers, returns to steps S200 to S500, iteratively obtains the local model parameters and intermediate model parameters to iteratively update the global model until the global model converges, and obtains the global model after training is completed.

[0057] In this embodiment, the central cloud is first initialized, and a global model and the initial model parameters corresponding to the global model are constructed, and the purpose and tasks of federated learning training are set. The initial model parameters are sent to the edge server, and the edge server receives the initial model parameters and sends the initial model parameters to the edge devices to ensure the normal operation of the abnormal detection model construction process. The edge devices construct local models according to the initial model parameters, and calculate the training tasks of the edge devices and the edge server according to the dynamic hierarchical decision algorithm. Among them, the training task is the hierarchical ratio of the local model. The method of hierarchical training can effectively solve the problem of heterogeneous resources of edge devices, provide a solution to the problem of insufficient computing power of edge devices in classical federated learning. At the same time, the dynamic hierarchical decision algorithm is adopted to monitor and calculate the hierarchical training tasks of edge devices and edge servers in real time. If abnormal situations or low training efficiency are found, timely adaptive adjustments can be made, and hierarchical decisions are given for specific cloud-edge environments to adapt to the changing cloud-edge environment. The edge devices and the edge server construct the front-end model and the back-end model respectively according to the training tasks. The edge devices and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters. The front-end model and the back-end model are constructed through reasonable hierarchical training tasks and jointly trained. Since some model parameters of the back-end model are deployed on the edge server side, the network transmission can be reduced to a certain extent, ensuring less network overhead. At the same time, the computing speed and computing accuracy of the edge devices are improved. The edge server performs intermediate aggregation on the local model parameters sent by each edge device to obtain intermediate model parameters. The two-stage aggregation method is adopted, and the edge server is used to perform the first aggregation of the model parameters to reduce the computing pressure of the central cloud model aggregation, increase the computing accuracy, reduce the amount of data transmitted to the central cloud node at the same time, and reduce the probability of network congestion. The central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge devices, and iteratively updates the global model until the global model converges to obtain the trained global model. The aggregated central model parameters of the edge server are used to perform the aggregation update of the global model again, effectively reducing the number of iterations of the global model update. The two-stage aggregation method is adopted to perform the aggregation calculation of the model parameters on the edge server and the central cloud respectively, avoiding the problems of unstable data transmission and crowded network channel transmission caused by the edge devices directly communicating with the central cloud frequently using the wide area network, effectively improving the calculation accuracy of the model parameters, reducing the computing pressure of the central cloud, and thus improving the efficiency and accuracy of model training in federated learning.

[0058] Optionally, the calculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes:

[0059] Obtain the average training time of a single sample of the local model in the edge device, compare the average training time of the single sample with the historical average training time. If the average training time of the single sample is greater than a preset multiple of the historical average training time and lasts for a preset number of times, then according to the dynamic hierarchical decision algorithm, recalculate the training tasks of the edge device and the edge server.

[0060] Specifically, after multiple rounds of iterative training, the dynamic hierarchical decision algorithm makes a timeout judgment on the training process to determine whether the hierarchical training task being trained at this time meets the requirements of federated learning training and whether it is necessary to recalculate the hierarchical training task. Real-time obtain the training time of a single sample of the local model in the edge device (i.e., the time to train a sample data on a certain edge device) and the historical average training time, and compare the training time of the single sample with the historical average training time. If the average training time of the single sample is greater than a preset multiple of the historical average training time and lasts for a preset number of times, then it is determined that the hierarchical training task being trained at this time is abnormal and does not meet the requirements of federated learning training. At this time, recalculate a new hierarchical training task according to the dynamic hierarchical decision algorithm for the model training of federated learning. For example, set the preset multiple to 2 times and the preset number of times to 5 times. Then, if the average training time of a single sample is greater than 2 times the historical average training time and lasts for 5 times, at this time, it is necessary to recalculate a new hierarchical training task according to the dynamic hierarchical decision algorithm for the model training of federated learning. Further, it can be thought that if the average training time of a single sample is less than 2 times the historical average training time, or the average training time of a single sample is greater than 2 times the historical average training time but does not reach 5 times, then continue to monitor the training process in real time for timely adjustment.

[0061] In this embodiment, the preset multiple and the preset number of times are set, and the average training time of a single sample of the local model is used as a condition to compare the training time of the single sample with the historical training time to judge whether to perform dynamic hierarchical decision-making. When both the preset multiple and the preset number of times are satisfied, the dynamic hierarchical strategy is performed, which coarsely controls the frequency of dynamic hierarchical decision-making and avoids the hierarchical decision-making itself becoming a bottleneck due to frequent dynamic hierarchical decision-making.

[0062] Optionally, the recalculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes:

[0063] Obtain the status information of the edge device. According to the status information, use the dynamic hierarchical strategy based on reinforcement learning to calculate the training task of the edge device and the edge server at the next moment. Wherein, the status information includes the total training time in a single round of edge device training, the training time of the edge device, the network transmission time, the proportion of the edge server training time in the total training time, and the training hierarchical strategy at this moment.

[0064] Specifically, after determining that the hierarchical training task does not meet the requirements of federated learning training, obtain the status information of the corresponding edge device, where the status information is represented by formula (1):

[0065]

[0066] Where k represents the device and t represents the iteration round. Represents the training status of device k in round t. Represents the total training time of device k in local training in round t. Represents the training time of edge device k in round t (i.e., the training time of the front-end model in round t). Represents the network transmission time. Represents the proportion of the training time of the edge server in round t (i.e., the training time of the back-end model in round t) in the total training time. Represents the hierarchical training task in round t.

[0067] In federated learning training, each edge device needs to obtain the most suitable hierarchical strategy for itself. For example, the DRL algorithm needs to obtain actions with the same dimension as the number of edge devices, that is, the training task. However, the number of edge devices may be different each time federated learning is performed, and users will also select the most suitable edge server for the device for hierarchical training according to the actual scenario. This will result in the output dimension of, for example, the DRL algorithm being of variable size, and thus it is impossible to achieve accurate model layering. In the embodiments of the present invention, according to the status information of each edge device, a hierarchical calculation is performed once using the dynamic hierarchical algorithm to obtain its optimal hierarchical training task, as shown in formula (2):

[0068]

[0069] Where and Represents the hierarchical strategy at time t + 1 for the action. Each time the algorithm runs, only one action, that is, the training task, is output for one edge device, and targeted training is performed on the edge device, so that the federated learning training acceleration method in the embodiments of the present invention can be applied to different neural networks, improving the application universality.

[0070] In this embodiment, according to the status information of edge devices, a dynamic stratification strategy is used to specifically calculate the most suitable stratified training task for each edge device, effectively avoiding the reduction of the training accuracy and efficiency of federated learning caused by the inability to accurately calculate the stratified training task due to the different numbers of edge devices in each federated learning, effectively improving the training accuracy and efficiency of federated learning, and improving the applicability of the federated learning model training.

[0071] Optionally, the edge device and the edge server respectively construct a front-end model and a back-end model according to the training task, and the edge device and the edge server jointly train the front-end model and the back-end model according to the data set. Obtaining local model parameters includes:

[0072] The edge device performs stratification processing on the local model according to the training task, constructs the front-end model, and sends the locally modeled and stratified model to the edge server, where the edge server constructs the back-end model according to the locally modeled and stratified model.

[0073] Specifically, the edge device receives the training task, stratifies the local model in the edge device according to the training task, retains the front-end part, constructs the front-end model on the edge device, and sends the back-end part after the local model is stratified to the edge server to construct the back-end model.

[0074] The edge device trains the front-end model according to the data set, obtains an intermediate result, and sends the intermediate result to the edge server, where the edge server trains the back-end model according to the intermediate result, obtains a gradient, and sends the gradient to the edge device.

[0075] Specifically, the front-end model is trained according to the obtained data set, that is, forward propagation is performed on the front-end model to obtain an intermediate result, and the intermediate result is sent to the edge server. The back-end model is trained using the intermediate result, forward propagation and backward propagation are performed on the back-end model, the back-end model parameters and gradients of the back-end model are obtained, and the gradients are sent to the edge device.

[0076] The edge device trains the front-end model according to the gradient to obtain local model parameters.

[0077] Specifically, backward propagation is performed on the front-end model according to the gradient sent by the edge server to obtain the front-end model parameters for aggregation and sending to the edge device.

[0078] In this embodiment, the edge server is fully utilized to build a front-end model and a back-end model on the edge device and the edge server respectively, and collaborative training is carried out. By leveraging the data parallelism feature of federated learning and the model parallelism feature in hierarchical training, an efficient training scheme is achieved. At the same time, after the back-end model is trained, the corresponding model parameters can be saved in the edge server, which can reduce the data transmission between the edge device and the edge server, reduce problems such as data loss or transmission errors caused by the instability of the wide area network between the edge device and the edge server, reduce network congestion problems, and effectively improve the training efficiency and accuracy of the local model.

[0079] Optionally, before the edge server trains the back-end model according to the intermediate result, it further includes:

[0080] Score each of the training tasks according to a preset scoring mechanism, sort each of the training tasks in descending order according to the scoring results, and set a priority queue.

[0081] According to the priority queue, use the GPU to train the back-end model one by one.

[0082] Optionally, the step of using the GPU to train the back-end model one by one according to the priority queue includes:

[0083] When the GPU video memory is insufficient, the remaining back-end models in the priority queue are scheduled to the CPU for training.

[0084] When the training of the back-end model in the GPU is completed, use the Watch mechanism to notify the CPU, and schedule the back-end model in the CPU to the GPU for training.

[0085] Specifically, after the edge server receives the training task of the back-end model, it is necessary to schedule the back-end model to the GPU for calculation. In order to maximize the utilization of the existing resources on the edge server, an edge server needs to provide hierarchical offloading services for multiple edge devices. However, the resources of the edge server are limited, especially the GPU resources. When the video memory is insufficient, the back-end model cannot be placed on the GPU for training. Although the number of service devices can be reasonably selected for the edge server during the federated learning task deployment phase, there may be some temporary tasks on the edge server that occupy computing resources such as the GPU. Therefore, in order to ensure the smooth operation of hierarchical training, the embodiment of the present invention proposes to use a scoring mechanism for resource allocation. By scoring the hierarchical training tasks, sorting the training tasks according to the scores, and scheduling them to the GPU for execution in turn. When the GPU resources are in short supply, they are scheduled to the CPU for execution. The hierarchical training task scoring mechanism is represented by formula (3):

[0086]

[0087] Among them, score k represents the score of the hierarchical training task of edge device k, and λ k represents the hierarchical strategy, represents the average training time of the device during non-hierarchical training.

[0088] As Figure 2 shown, obtain the hierarchical training tasks of each edge device, score each hierarchical training task according to formula (3), arrange the hierarchical training tasks from largest to smallest according to the scoring results, construct a priority queue, and when hierarchical training is required, take out the corresponding backend model from the head of the priority queue and establish corresponding worker threads to provide training capabilities. Additionally, this score indicates how long it is expected to take to train the backend model if it is placed on the edge server. The larger this time is, the more likely this edge device is to become a straggler, and the GPU should be preferentially used for acceleration. In addition, the score score k′ of the previous round's straggler k' should be set to infinity to ensure that the previous round's straggler must be at the head of the queue, ensuring that it can obtain GPU resources preferentially, thereby further reducing the impact of the straggler problem.

[0089] Furthermore, when situations such as insufficient GPU video memory or display anomalies are captured, the remaining tasks (i.e., backend models) in the priority queue are scheduled to be executed on the CPU. However, since the execution speed of the GPU is much faster than that of the CPU, in order to make full use of GPU resources, the embodiments of the present invention utilize the Watch mechanism. When the tasks in the GPU are trained and completed, the tasks running on the CPU are informed through the watch mechanism that the GPU is now idle, so that the tasks running on the CPU can be placed in the GPU for training, improving the utilization rate of the GPU. Of course, in most cases, through reasonable layout, the edge server can load all backend models onto the GPU.

[0090] In this embodiment, a scoring mechanism is used to score the training tasks, and the training tasks are sorted according to the scoring results and then a priority queue is constructed. The backend models are sequentially trained according to the priority queue, which can effectively reduce the impact of the straggler problem. At the same time, when GPU resources are scarce, the training tasks are scheduled to be executed on the CPU in order, effectively alleviating the operating pressure of the GPU, and model training can also be synchronized when GPU resources are scarce. By using the watch mechanism, when GPU resources are idle, the resources on the CPU are timely scheduled to be executed on the GPU, improving the overall utilization ability of the GPU and the execution efficiency of the training tasks.

[0091] Regarding the two-stage aggregation scheme, it should be noted that, as Figure 3As shown in the figure, the two-stage aggregation architecture breaks the single-center architecture in the federated learning aggregation process. The edge devices are divided into different edge servers according to regions. After the edge devices have undergone several rounds of local model training, they will transfer the local model parameters of the edge devices to the corresponding edge servers. The edge servers will perform intermediate aggregation on the local parameters of all subordinate edge devices to obtain the intermediate parameters at the edge server level, and then redistribute the intermediate aggregation parameters to the edge devices. After several rounds of intermediate aggregation, the intermediate model of the edge server is transferred to the central cloud, and the central cloud completes the global aggregation.

[0092] To better understand the two-stage aggregation process, some definitions of the aggregation process of classical federated learning are given first. Assume that the training data set for the entire federated learning is D = {(x j , y j ) | j = 1,..., |D|}, where |D| represents the number of samples in the entire data set, x j is the j-th sample, and y j is the label corresponding to sample j. Assume that a total of N edge devices participate in the aggregation, and the local data volume on edge device k is D k (D k ∈ D). The empirical loss of the local model on edge device k is expressed using formula (4):

[0093]

[0094] Among them, F(w k ) represents the empirical loss of the local model on edge device k, D k represents the local data volume on edge device k, and f j (w k ) represents the loss function of the model.

[0095] For model parameter updates, the gradient descent algorithm is used, and the parameter updates are expressed using formula (5):

[0096]

[0097] Among them, w k (t) represents the local model parameters of edge device k at the t-th round, η represents the learning rate, represents the gradient.

[0098] At the central cloud, it is necessary to aggregate the parameters from N devices, but what it performs is not gradient descent, but weighted average. The global empirical loss cannot be directly calculated. Similarly, the weighted average of the empirical losses of each edge device is used as the representation of the global empirical loss. The global model parameters and global empirical loss are expressed using formula (6) and formula (7) respectively:

[0099]

[0100]

[0101] Among them, w are the global model parameters. At the start of each local model training, w k = w. N represents the number of edge devices, D k represents the local data volume on edge device k, |D| represents the number of samples in the entire dataset, F(w) is the global empirical loss, and F(w k ) represents the empirical loss of the local model on edge device k.

[0102] To better utilize the architectural advantages of edge computing and at the same time reduce the communication frequency between edge devices and the central node, this paper introduces a two-stage aggregation architecture. The biggest difference between the two-stage aggregation architecture and the classical federated learning algorithm lies in the addition of an intermediate aggregation process. Similar to classical federated learning, for the two-stage aggregation algorithm, the setting of two-layer federated learning is continued. Suppose there are N edge devices and E edge servers corresponding to them. Edge server i is connected to N i edge devices, and the covered dataset is D i with a size of |D i |. Above the E edge servers is a central cloud node. The local model training stage of the two-stage aggregation algorithm is the same as that of classical federated learning, that is, the calculation methods described by formulas (4) and (5). For the intermediate aggregation stage, the weighted average method is used on edge server i to calculate the intermediate model parameter w i and the empirical loss F(w i ). The intermediate model parameter and the intermediate empirical loss are represented by formulas (8) and (9) respectively:

[0103]

[0104]

[0105] Among them, w i represents the intermediate model parameter, F(w i ) represents the intermediate empirical loss, N i represents the number of edge devices, D k represents the local data volume on edge device k, w k represents the local model parameter, |D i | is the dataset size, and F(w k ) represents the empirical loss of the local model on edge device k.

[0106] For the aggregation on the central cloud, the intermediate model parameters from E edge servers are weighted averaged according to the size |D i | of the dataset responsible for each edge server i. The global model parameters and the global empirical loss are represented by formulas (10) and (11) respectively:

[0107]

[0108]

[0109] where w represents the global model parameters, F(w) represents the global empirical loss, E represents the number of edge servers, |D i | is the size of the dataset, w i represents the intermediate model parameters, |D| represents the number of samples in the entire dataset, and F(w i ) represents the intermediate empirical loss.

[0110] For the learning objective of the two-stage federated learning, it is to find the optimal global parameter w * , such that the global empirical loss is minimized, which is represented by formula (12):

[0111] min w F(w) (12),

[0112] where w represents the global model parameters and F(w) represents the global empirical loss.

[0113] In actual training, in order to reduce the communication overhead, the local model training phase often performs multiple iterations on the local dataset before passing the model parameters to the central node for aggregation. As Figure 4 shown, corresponding to the two-stage aggregation algorithm, it is defined that after the edge device undergoes k1 = 2 local iterations, it will pass the local model to the edge server, and after the edge server undergoes k2 = 2 intermediate aggregations, it will pass the intermediate model to the central cloud node for global aggregation. That is to say, one global aggregation requires k = k1·k2 local model iterations and k2 intermediate aggregations of the edge server. Here, k1 is defined as the local iteration frequency and k2 is defined as the intermediate aggregation frequency. The embodiment of the present invention proposes a self-adjusting two-stage aggregation algorithm AdaAGG based on a heuristic algorithm to automatically determine the aggregation periods k1 and k2, thereby minimizing the overall training time. It should be noted that this algorithm does not directly determine the specific values of k1 and k2, but automatically controls the training frequency of each device by introducing the concept of the loss rate.

[0114] Optionally, the edge device and the edge server co-train the front-end model and the back-end model according to the dataset, and obtaining the local model parameters includes:

[0115] The edge device calculates the number of device training times according to the device loss rate. The edge device and the edge server jointly train the front-end model and the back-end model according to the number of device training times, obtain the local model parameters, and calculate and obtain the device training time. Then, the device training time and the local model parameters are sent to the edge server. Among them, the device loss rate is initialized and preset by the central cloud and sent to the edge server, and the edge server sends the device loss rate to the edge device.

[0116] Specifically, the central cloud initializes and presets the device loss rate, and transfers it to the edge device through the edge server. The edge device calculates the number of device training times according to the device loss rate. The edge device and the edge server jointly train the front-end model and the back-end model according to the number of device training times, which specifically includes: to meet the device loss rate θ∈(0, 1), each device needs to perform L(θ)=μlog(1 / θ) local iterations, where μ is a constant coefficient related to the data set and the model to be trained. To meet the local loss rate condition, the edge device k needs to continuously update the local parameters on the local data set according to formula (5) until formula (13) is satisfied. At this time, it can be considered that the local model training stage ends.

[0117]

[0118] Among them, w k (t) represents the local model parameters of the edge device k at the t-th round, θ represents the device loss rate, represents the gradient. It can be seen from formula (13) that the so-called device loss rate θ essentially represents the change degree of the gradient of the empirical loss function. The smaller the loss rate, the greater the degree of gradient descent, and at this time the local iteration times L(θ) will also increase. When the device loss rate is infinite, it means that any small change in the gradient before and after training will meet the loss rate condition. When the device loss rate is 0, it means that it is necessary to ensure that the subsequent training gradient no longer changes, indicating that the model has reached the global optimum, and the corresponding iteration times L(θ) will theoretically become infinite.

[0119] When hierarchical training is not performed (i.e., when the edge device training ends), the training time for one round of model training iteration in the edge device (i.e., the device training time) is expressed by formula (14):

[0120]

[0121] Among them, is the training time for one round of model training iteration in the edge device, c k is the edge device, assuming sk Calculate the number of CPU cycles required for a data sample (x j , y j ) for device k. The size of the local dataset of device k is |D k |. Therefore, the total CPU cycles required for one local iteration is s k |D k |. Assume that the CPU frequency of device k is which is restricted by the specific CPU model of device k.

[0122] In this embodiment, the edge device can use an efficient network to send local parameters to the edge server, reducing data transmission time. Introduce the device loss rate to calculate the frequency of model training and the device training time in the edge device, so as to increase the accuracy of local model parameter calculation, reduce the number of subsequent global model aggregations, and improve the training efficiency of federated learning.

[0123] Optionally, the edge server performs intermediate aggregation according to the local model parameters sent by each edge device to obtain intermediate model parameters, and further includes:

[0124] The edge server calculates the number of intermediate aggregations according to the server loss rate, performs intermediate aggregation according to each local model parameter and the number of intermediate aggregations to obtain the intermediate model parameters, calculates the intermediate aggregation time, and sends the intermediate model parameters and the intermediate aggregation time to the central cloud, where the server loss rate is initialized and preset by the central cloud.

[0125] Specifically, the central cloud initializes and presets the server loss rate. The edge server calculates the number of intermediate aggregations according to the server loss rate, performs intermediate aggregation according to each local model parameter and the number of intermediate aggregations, and calculates the intermediate aggregation time. Specifically, after the local model training stage ends, that is, after completing L(θ) local iterations, the local model parameters w k of all device k will be transmitted to its corresponding edge server i. It is defined that the communication between the edge device and the edge server satisfies the OFDMA (Orthogonal Frequency Division Multiple Access) protocol, and the communication rate is represented by formula (15):

[0126]

[0127] where r k represents the communication rate, the total bandwidth of edge server i is B i , α k:i represents the proportion of the bandwidth of edge device k in the total bandwidth, that is, the available bandwidth of edge device k is α k:i B i , h k is the channel gain of edge device k, ek The energy consumed for transmitting data to edge device k, ξ k is the background noise.

[0128] Calculate the parameter transmission time of edge device k using formula (16):

[0129]

[0130] where, is the parameter transmission time of edge device k, and the size of the local model parameter is d k , r k represents the communication rate.

[0131] After the edge server receives the local model training, it needs to perform intermediate parameter aggregation. After the aggregation is completed, edge server i will broadcast and transmit the new w i to all subordinate devices. Here, the two parts of the network transmission time are uniformly simplified to So the interval time of intermediate aggregation of edge server i (i.e., the intermediate aggregation time) is expressed using formula (17):

[0132]

[0133] where, represents the interval time of intermediate aggregation of edge server i, and the aggregation calculation time of edge server i is the number of local model iteration L(θ), is the training time for one round of model training in the edge device.

[0134] Subsequently, repeat the local model training phase until the intermediate model on the edge server reaches the edge loss rate ∈ (all edge servers maintain the same edge loss rate). The meaning of the edge loss rate is similar to the local loss rate, that is, through continuous intermediate aggregation until formula (18) is satisfied. At this time, it can be considered that the intermediate aggregation process ends.

[0135]

[0136] where, w i (t) represents the intermediate model parameter after t rounds of aggregation, ∈ represents the server loss rate, represents the gradient.

[0137] To meet the set server loss rate, for machine learning tasks, the number of intermediate aggregations I(∈, θ) that the edge server needs to perform is expressed using formula (19):

[0138]

[0139] Among them, θ represents the device loss rate, ∈ represents the server loss rate, and γ is a coefficient related to a specific training task. Formula (19) is affected by both the server loss rate ∈ and the device loss rate θ. The value of ∈ mainly affects the numerator part in I(∈, θ), that is, when θ is fixed, the change trend of the I(∈) function is the same as that of L(θ). Incorporate this parameter into the interval time of intermediate aggregation in the edge server i That is, measure this coefficient with the actual execution time of the training task.

[0140] Here, I(∈, θ) is equivalent to k2, that is, after intermediate aggregation I(∈, θ) times on the edge server, the edge server will upload the intermediate parameters w i to the central cloud for global aggregation.

[0141] In this embodiment, a two-stage aggregation algorithm is adopted, and the server loss rate is introduced to calculate the frequency and intermediate aggregation time of local model parameter aggregation in the edge server, so as to increase the accuracy of intermediate aggregation parameter calculation, reduce the number of subsequent global model aggregations, and improve the training efficiency of federated learning.

[0142] Optionally, the central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges, including:

[0143] The central cloud updates the global model according to the intermediate model parameters sent by each edge server, calculates the total aggregation time according to the intermediate aggregation time, calculates the optimal device loss rate and the optimal server loss rate according to the total aggregation time, and sends the model parameters of the updated global model, the optimal device loss rate and the optimal server loss rate to each edge server, and returns to execute the step of the edge server sending the device loss rate to the edge device, and iteratively updates the global model until the global model converges.

[0144] Specifically, calculating the total aggregation time according to the intermediate aggregation time includes: the edge server i needs to pass through the wide area network to transfer the intermediate parameters to the central cloud, and the global model aggregation period t cloid is represented by formula (20):

[0145]

[0146] where t cloud is the global model aggregation period, and I(∈, θ) is the number of intermediate aggregations required by the edge server, Indicates the interval time for intermediate aggregation in edge server i, is the transmission time, is the time overhead for the central cloud to perform global aggregation.

[0147] When the central cloud node receives the parameters from the edge server, it will perform weighted averaging of the parameters according to formula (10). Similar to the loss rate on the edge server, there is a central loss rate λ on the central cloud. When the edge server uploads parameters, it can ensure that the server loss rate has reached ∈. Therefore, formula (19) is extended. That is, in order to achieve the central loss rate λ, I(λ, ∈) global aggregations are required. So finally, when the global model of the central cloud reaches the central loss rate λ, the total time spent is defined as formula (21):

[0148] t total = I(λ, ∈)·t cloud (21),

[0149] where, t total is the total aggregation time, I(λ, ∈) is the number of central aggregations, t cloud is the global model aggregation period.

[0150] Substitute formula (17) and formula (20) into formula (21), and the total aggregation time is defined as formula (22):

[0151] t total = log(1 / λ)·W(∈, θ)

[0152]

[0153] When the loss rate decreases, the gradient corresponding to the empirical loss will also become smaller and smaller. When the loss rate approaches 0, it means that the empirical loss of the entire model is 0, indicating that it has reached the global optimal point. In the AdaAGG algorithm, the loss rate on the central cloud should be as low as possible in the end. Here, a positive number 1e -5 is taken for λ, and log(1 / λ) is converted into a constant, so that the total aggregation time t total is converted into a function W(∈, θ) related to the device loss rate θ and the edge loss rate ∈. According to this function, the optimal device loss rate and server loss rate are calculated, and the step of the edge server sending the device loss rate to the edge device is returned. The global model is iteratively updated until the global model converges, completing the federated learning training and obtaining the trained global model.

[0154] In this embodiment, a two-stage aggregation algorithm is adopted, and only the intermediate model parameters after intermediate aggregation transmitted by the edge server are received, which can reduce the amount of data transmitted to the central cloud node, reduce the probability of network congestion. At the same time, the loss rate is introduced as an iteration condition for model training, effectively increasing the accuracy of model parameter training, reducing the number of model iterative trainings, and increasing the efficiency of federated model training.

[0155] Optionally, calculating the optimal device loss rate and the optimal server loss rate according to the total aggregation time includes:

[0156] Construct a combinatorial optimization problem according to the total aggregation time, the device loss rate, and the server loss rate.

[0157] Use the simulated annealing algorithm to iteratively solve the combinatorial optimization problem to obtain the optimal device loss rate and the optimal server loss rate.

[0158] Specifically, in order to minimize the total aggregation time t total The optimization objective is defined as formula (23):

[0159] min ∈,θ W(∈, θ)

[0160] s.t. 0 < ∈ < 1

[0161] 0 < θ < 1

[0162] (23),

[0163] Thus, a decision problem is transformed into a combinatorial optimization problem. By selecting the optimal device loss rate and server, the time to reach the specified central loss rate is minimized, that is, the total aggregation time is minimized.

[0164] To solve the above combinatorial optimization problem, the simulated annealing algorithm is introduced. This algorithm is a heuristic algorithm that simulates the principle of metal annealing in reality and iteratively solves the optimization strategy based on Monte-Carlo. The simulated annealing algorithm starts from a certain high temperature, obtains a new combination by adding perturbations. If it is a good combination, it is accepted. For a bad combination, it is judged whether to accept according to the Metropolis criterion, so that the algorithm can jump out of the local optimal solution and find the global optimal solution. The specific process of using the simulated annealing algorithm to solve the optimal loss rate is as Figure 5 shown. Its core lies in that when a more appropriate loss rate is found to make the objective function smaller, the optimal loss rate combination is updated at this time; when ΔW is less than 0, there is a certain probability of accepting this loss rate combination at this time in order to find the global optimal solution.

[0165] In this embodiment, the problem of solving the loss rate is transformed into a combinatorial optimization problem, and the simulated annealing algorithm is used to solve the optimal loss rate combination, effectively improving the efficiency and accuracy of solving the loss rate, thereby improving the iteration rate of the global model. While ensuring the training accuracy of federated learning, the training rate is accelerated.

[0166] An embodiment of the present invention also provides a cloud-edge training platform, which supports the construction of the experimental scenario of the federated learning training method in the embodiment of the present invention.

[0167] Most of the current research in the cloud-edge scenario is carried out through simulation experiments, or the experimental environment is directly deployed on physical machines. The deployment of the experimental environment is difficult, and the research results are difficult to collect. On this basis, an embodiment of the present invention provides a cloud-edge training platform, which can be quickly deployed on the physical environment, provides the necessary scenario construction ability for cloud-edge experiments, reduces the workload of experimenters in building cloud-edge platforms and constructing software dependencies, and at the same time provides rich interfaces and UI interfaces, facilitating experimenters to effectively monitor the platform and facilitating the collection of required experimental data.

[0168] As Figure 6 shown, the cloud-edge training platform selects KubeEdge as the underlying virtualization support, enriches its functions and defines its configurations, and provides functions including fast deployment scripts, node status acquisition, graphical monitoring of nodes, service registration, etc., including cloud-edge containerization components, cloud-edge management components, and federated learning training components.

[0169] Among them, the cloud-edge containerization components mainly rely on the container orchestration solution provided by KubeEdge itself. Similar to Kubernetes, KubeEdge itself also depends on components such as Kubectl, Apiserver, and Etcd. Among them, Kubectl is the user instruction entry, and users operate on the cluster through Kubectl instructions; Apiserver is mainly used to receive user instructions, parse the instructions or YAML files, and parse them into Kubernetes resource objects (such as POD, Demployment, etc.); Etcd is used for status synchronization between the central cloud node and the edge node. It itself is a strongly consistent KV storage component based on the RAFT algorithm. On this basis, KubeEdge adds two components, one is Cloudcore and the other is Edgecore. These two components are essentially a deep modification of the Kubelet and Controller Schedule components in K8S. These two components are respectively deployed on the central cloud node and the edge node, and tasks are scheduled to the edge node for execution according to the scheduling policy. KubeEdge fully supports the API resources of Kubernetes, so we can create resources such as Pod, Deployment, and StatefulSet on the KubeEdge platform just like using Kubernetes. And KubeEdge communicates between CloudHub and EdgeHub based on the WebSocket and QUIC protocols, which is more adaptable to the complex and unstable network conditions between the cloud and the edge compared to the native Kubernetes solution. And KubeEdge also stores the corresponding resource information at the edge end, so that even if the communication between the cloud and the edge is disconnected, the containers can continue to run normally according to the locally cached resource information such as Pods, and there will be no situation where the node containers terminate when the edge node goes offline. The cloud-edge containerization components are built on 4 nodes, including one central cloud node and three computing nodes.

[0170] In order to achieve effective management of the cloud-edge platform, improve the availability of the platform, and reduce the operation and maintenance difficulty of the platform, a set of visual cloud-edge management suites is built, including cloud-edge platform visual management components, node resource monitoring components, and application log collection components. Through the above visual platform, users can perform functions such as platform management, resource monitoring, and application log collection through the front end according to their personal needs, which is convenient for users to effectively control the running status of the cloud-edge platform. Although there is a method to obtain node status in KubeEdge, this method uses Kubectl commands and is deeply bound to K8S, with low scalability. In actual scenarios, system administrators need to be able to monitor node resources in real time, and users should have the ability to flexibly collect and retrieve node status information. This platform provides two ways to obtain node status. If users only hope to perform a simple single status acquisition, they need to send a request to the corresponding node by calling the status query CLI. A lightweight status collection Agent is deployed on each node of the cloud-edge platform. After receiving the required status information by the Agent (by reading the proc sub-file system), it will return it to the user and save it as a KV pair stored in ETCD. The Key is "user + time string + UUID", and an expiration time is set for this KV pair (default is 12 hours). Of course, it also supports specifying --store when passing parameters to persist the result of this time. When users need to obtain time-series status, it is not a wise approach to frequently call the status acquisition interface. Therefore, Promethus API is integrated in this platform, and users can directly obtain time-series status information from the TSDB of Promethus through Python scripts. In addition to obtaining the corresponding status information, this system also needs a set of visualization devices to facilitate managers to view the resource status of nodes, so as to discover and analyze problems in a timely manner. After research, it is decided to use the open-source components Promethus plus Grafana to achieve this. Promethus can be essentially understood as a time-series database. There will be a node-export component on each node to regularly obtain the corresponding status according to the Metric defined by Prothemus. Promethus summarizes and stores the status information, while Grafana is a visualization platform used to convert resource information in various dimensions into dashboards. The platform management component runs in the form of containers on top of the Docker engine in this platform, and KubeEdge is used for container orchestration and scheduling. In order to improve the user experience of the platform, users will prefer to manage tasks in a visual way.Currently, there is no corresponding platform management UI component in the KubeEdge community. After research, this platform has built a set of platform management components suitable for KubeEdge based on KuBoard. This component controls the entire cloud-edge platform by converting the operations of users in the front end into instructions for the ApiServer of KubeEdge. It provides functions including cluster import, task and resource creation, task and resource deletion, task and resource reset, namespace management, container status visualization, online Bash, etc., facilitating users to manage clusters and tasks. As the number of tasks running on the cloud-edge platform continues to grow, each application (POD) generates a large amount of logs, which are defaultly directly output to the stdout or stderr of this POD. With the rapid growth of task logs, it soon becomes difficult to manage. What's more complex is that when problems occur, due to the complex call relationships between services and a large number of invalid logs, it is very difficult to quickly find the cause of the error. Although for Docker, logs can be optionally output to the corresponding directory of the container, but when the container stops, this log file will be cleared. Many times, users may print experimental phenomena in the form of logs, which is also not conducive to users' later analysis of experimental phenomena. Therefore, it is necessary to collect the logs of platform applications and provide users with the function of aggregating logs according to keywords. In this platform, a set of log collection platforms is built based on ELK (ElasticSearch, Logstash, Kibana). An additional container encapsulating the log collection component is started in each application's POD, which will collect the logs output by the containers in this POD, then upload them to the ElasticSearch server, and build an index with "logstash-[pod_name]-[pod_create_time]" as the key in ES. In this way, users can retrieve the log information in ES through Kibana, view all the log information of the application by the index name, or filter and retrieve the log information by log keywords to locate problems and analyze experimental phenomena. There is a problem in the architecture of KubeEdge, that is, there is no concept of edge cloud in KubeEdge, and all edge nodes equivalently interact with the central cloud control node. However, in many edge environments, especially in mobile scenarios, the edge is often a small cluster with autonomous capabilities, that is, the edge cloud. Inside the edge cloud, there is a control node of the edge cloud and several edge servers as working nodes. In many cloud-edge scenarios, users will schedule tasks to run on the specified edge cloud. At this time, KubeEdge cannot provide the corresponding capabilities. Therefore, in order to adapt to more cloud-edge scenarios, this platform needs to support the organizational structure of the edge cloud.There are two ways to partition the edge cloud. One is to directly perform physical isolation and change the architecture, but this approach is relatively complex and requires large-scale modifications to KubeEdge. The second method is to perform logical segmentation. Through an independent service registration component, each node is assigned a corresponding role. When the edge cloud concept is needed, users can query the service registration component to obtain the current node-to-role mapping relationship. This method is insensitive to KubeEdge. The independent server registration component manages, adds, and deletes mapping information, which can greatly reduce the complexity. This component selects zookeeper as the metadata storage component and uses JSON files for cloud-edge role registration. Service registration essentially parses the JSON file and registers Znodes in Zookeeper according to the hierarchical relationship. The central cloud is the top-level directory, and the following edge clouds are mounted as subdirectories in sequence according to the hierarchical relationship in the configuration file, finally forming a tree structure. The metadata stores the relevant information of the node, including IP address, registration time, node role, etc. Using service registration can logically partition the nodes. When writing the YAML file for KubeEdge tasks, the logical role name can be provided, and then the provided script is used for text processing. This is convenient for users and does not require memorizing the IP addresses corresponding to each node.

[0171] To provide a federated learning training environment, this platform packages the environment into a complete image based on Docker. To support GPUs, this component is encapsulated based on the official Nvidia image, and the CUDA version is 11.0. The deep learning component is mainly used to verify the federated learning acceleration algorithm in this paper. Of course, it can also be used for regular model training. The architecture is as Figure 7As shown in the figure. To achieve a more general federated learning function, the federated learning component of this platform is divided into three layers. The top layer is the model layer, which mainly provides framework support for users to write federated learning model codes. This component pre-integrates two mainstream deep learning frameworks, TensorFlow and PyTorch. At the same time, to meet the needs of distributed training, this component also integrates the ChainerMN framework, which can realize the construction of a flexible distributed computational graph. Below the model layer is the Lib layer, which mainly provides two functions: (1) First is the parameter server. The federated learning model training involved in this article needs to be coordinated across multiple nodes. Although some functions of distributed training have been implemented in current TensorFlow and PyTorch, there are still many usage limitations (such as not supporting CPU-based distributed training). Therefore, this component implements a set of parameter servers based on FastAPI (a high-performance web server framework) for the interaction of parameters and status information between nodes. At the same time, in order to decouple the sent data from the data type, the Pickle library is used in the parameter server to serialize the object to be sent, and the receiver deserializes the data to achieve a unified processing process for different data types. The parameter information consists of three parts: information length, information identifier, and information body. The information identifier is of String type, and the receiver needs to verify whether the information type is correct. The information type and the information body are encapsulated using a list and sent to the receiver after serialization. (2) Secondly, there are some utility class interfaces, including algorithm function interfaces such as FedAVG and simulated annealing, dataset processing and splitting interfaces, and logging-based log control interfaces, which facilitate the rapid writing of experimental codes. The bottom layer is the network communication layer. This platform supports multiple CNI components, including Calico, Flanne, and Weave. Through these container network components, network routing and network virtualization functions at the cross-border points can be provided.

[0172] Although the present disclosure is disclosed as above, the protection scope of the present disclosure is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present disclosure, and these changes and modifications will all fall within the protection scope of the present invention.

Claims

1. A method for accelerating federated learning training, characterized in that, Including: The central cloud constructs a global model and the initial model parameters corresponding to the global model, and sends the initial model parameters to the edge server; The edge server receives the initial model parameters and sends the initial model parameters to the edge device; The edge device constructs a local model according to the initial model parameters, and calculates the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm, where the training task is the hierarchical ratio of the local model; Calculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes: Obtain the average training time of a single sample of the local model in the edge device, compare the average training time of the single sample with the historical average training time. If the average training time of the single sample is greater than a preset multiple of the historical average training time and lasts for a preset number of times, then according to the dynamic hierarchical decision algorithm, recalculate the training tasks of the edge device and the edge server; Recalculating the training tasks of the edge device and the edge server according to the dynamic hierarchical decision algorithm includes: Obtain the status information of the edge device, and according to the status information, calculate the training tasks of the edge device and the edge server at the next moment by using the dynamic hierarchical strategy based on reinforcement learning, where the status information includes the total training time in a single round of edge device training, the training time of the edge device, the network transmission time, the proportion of the edge server training time in the total training time, and the training hierarchical strategy at this moment; The edge device and the edge server respectively construct a front-end model and a back-end model according to the training tasks, and the edge device and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters, and send the local model parameters to the edge server; The edge server performs intermediate aggregation according to the local model parameters sent by each edge device to obtain intermediate model parameters, and sends the intermediate model parameters to the central cloud; The central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges.

2. The method for accelerating federated learning training according to claim 1, characterized in that, The edge device and the edge server respectively construct a front-end model and a back-end model according to the training tasks, and the edge device and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters, including: The edge device performs hierarchical processing on the local model according to the training tasks, constructs the front-end model, and sends the hierarchically processed local model to the edge server, where the edge server constructs the back-end model according to the hierarchically processed local model; The edge device trains the front-end model according to the data set, obtains an intermediate result, and sends the intermediate result to the edge server. Wherein, the edge server trains the back-end model according to the intermediate result, obtains a gradient, and sends the gradient to the edge device; The edge device trains the front-end model according to the gradient and obtains local model parameters.

3. The method for accelerating federated learning training according to claim 2, characterized in that, Before the edge server trains the back-end model according to the intermediate result, it further includes: Scoring each training task according to a preset scoring mechanism, sorting each training task in descending order according to the scoring results, and setting a priority queue; According to the priority queue, the back-end model is trained one by one using the GPU.

4. The method for accelerating federated learning training according to claim 3, characterized in that, The training the back-end model one by one using the GPU according to the priority queue includes: When the GPU video memory is insufficient, the remaining back-end models in the priority queue are scheduled to the CPU for training; When the training of the back-end model in the GPU is completed, the CPU is notified using the Watch mechanism, and the back-end model in the CPU is scheduled to the GPU for training.

5. The method for accelerating federated learning training according to claim 1, characterized in that, The edge device and the edge server jointly train the front-end model and the back-end model according to the data set to obtain local model parameters, including: The edge device calculates the device training times according to the device loss rate. The edge device and the edge server jointly train the front-end model and the back-end model according to the device training times, obtain the local model parameters, and calculate the device training time. The device training time and the local model parameters are sent to the edge server. Wherein, the device loss rate is initialized and preset by the central cloud and sent to the edge server, and the edge server sends the device loss rate to the edge device.

6. The method for accelerating federated learning training according to claim 5, characterized in that, The edge server performs intermediate aggregation on the local model parameters sent by each edge device to obtain intermediate model parameters, and further includes: The edge server calculates the intermediate aggregation times according to the server loss rate, performs intermediate aggregation according to each local model parameter and the intermediate aggregation times to obtain the intermediate model parameters, calculates the intermediate aggregation time, and sends the intermediate model parameters and the intermediate aggregation time to the central cloud. Wherein, the server loss rate is initialized and preset by the central cloud.

7. The method for accelerating federated learning training according to claim 6, characterized in that, The central cloud updates the global model according to the intermediate model parameters sent by each edge server, and sends the model parameters of the updated global model to each edge server, and returns to execute the step of the edge server sending the model parameters to the edge device, and iteratively updates the global model until the global model converges, including: The central cloud updates the global model according to the intermediate model parameters sent by each edge server, calculates the total aggregation time based on the intermediate aggregation time, calculates the optimal device loss rate and the optimal server loss rate according to the total aggregation time, and sends the model parameters of the updated global model, the optimal device loss rate, and the optimal server loss rate to each edge server, and returns to execute the step of the edge server sending the device loss rate to the edge device, and iteratively updates the global model until the global model converges.

8. The method for accelerating federated learning training according to claim 7, characterized in that, The calculating the optimal device loss rate and the optimal server loss rate according to the total aggregation time includes: Constructing a combinatorial optimization problem according to the total aggregation time, the device loss rate, and the server loss rate; Iteratively solving the combinatorial optimization problem by using the simulated annealing algorithm to obtain the optimal device loss rate and the optimal server loss rate.