Multi-edge load balancing task scheduling method based on federal reinforcement learning
Through the multi-edge load balancing task scheduling method of federated reinforcement learning, the problem of load imbalance of edge servers in wireless metropolitan areas is solved, efficient load balancing and low-cost task scheduling are achieved, and task response time is shortened.
Patent Information
- Application Number
- CN202510340848.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
In wireless metropolitan area networks, the traditional centralized decision-making model has problems of high communication costs and long decision-making time in the load balancing research in multi-edge systems, and it is difficult to cope with the actual situation of large number of edge servers, wide distribution, and rapid load dynamics.
Using a multi-edge load balancing task scheduling method based on federated reinforcement learning, the Markov decision-making process is constructed, and local training is performed using DQN algorithm and experience playback pool, and combined with federated learning is used to aggregate the model to achieve distributed load balancing.
Effectively balance the task load at each edge, shorten the maximum task response time in the edge system, reduce communication costs, and improve system performance.
Smart Images

Figure CN120276819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, and particularly relates to a multi-edge load balancing task scheduling method based on federated reinforcement learning. Background Art
[0002] In recent years, with the continuous development of mobile communication technology and the Internet, the penetration rate of mobile devices has increased significantly, and the number of mobile devices per capita has increased sharply. However, the computing power, battery, and data storage capacity of personal mobile devices are limited, which leads to limitations in performance and energy consumption when mobile devices process complex tasks, and it is difficult to meet the requirements of users for low-latency and high-reliability tasks. To overcome the resource limitations of mobile devices, Mobile Cloud Computing (MCC) emerged as a solution. By migrating computationally intensive tasks to remote cloud servers, it achieves more powerful computing and storage capabilities, reduces the burden on mobile devices, and extends their service life. However, remote cloud services may experience communication delays due to geographical distance and network fluctuations, affecting the Quality of Service (QoS), especially during high-load periods when the QoS drops more significantly.
[0003] To address the shortcomings and deficiencies of MCC, Mobile Edge Computing (MEC) was proposed as an emerging paradigm. MEC extends the traditional two-layer architecture of mobile devices - remote cloud by deploying computing resources to network edge nodes, forming a three-layer architecture of mobile devices - edge - remote cloud. The edge nodes are used as small data centers and are deployed close to users to provide low-latency computing and storage services through wireless networks. When users process tasks with tolerable latency, the edge server can migrate non-urgent tasks to the remote cloud to relieve its own load pressure by utilizing the computing resources of the remote cloud.
[0004] Currently, some scholars have started to study the deployment of edge nodes in Wireless Metropolitan Area Networks (WMANs) to build a collaborative system architecture. However, load balancing and task request allocation between edge nodes have become key issues. The traditional nearest neighbor allocation strategy may lead to uneven load distribution when the number of users surges, increasing the task response time and even causing task execution failures. Therefore, dynamically adjusting the task allocation strategy based on factors such as edge node resource utilization, network topology, task type, and priority to achieve load balancing is the key to improving system performance. At the same time, the communication overhead between edge nodes and the complexity of task migration also need to be considered. To achieve effective load balancing, distributed algorithms and machine learning techniques may be required for task request allocation and load management.
[0005] In the field of mobile cloud computing, scholars are dedicated to researching how to improve the task offloading performance between mobile devices and remote clouds. For example, the mCloud framework proposed by Zhou selects the optimal task offloading scheme by obtaining the resource utilization rates of mobile devices and cloud servers, thus enhancing the performance of cloud servers. The computation offloading framework designed by MAUI allows developers to autonomously choose the offloading granularity and requires annotating the methods that need to be remotely offloaded to reduce the task response time.
[0006] However, when a large number of tasks are migrated to the remote cloud, it may lead to network congestion or server overload, further reducing the QoS. To address this challenge, Satyanarayanan et al. proposed a three-layer task offloading architecture including an edge layer based on MCC. An intermediate layer, called the edge, was added between the original mobile device - remote cloud. Edge nodes, as computer clusters near mobile devices, provide computing and data resource services. However, the uneven distribution of task request quantities caused by the differences in user densities in different regions makes the scheduling problem of edge load balancing a research hotspot.
[0007] Currently, the centralized decision-making mode is used in the research of load balancing in multi-edge systems. By obtaining the load, bandwidth, and computing resource information of edge nodes in real time, a unified strategy or search algorithm is adopted to determine the optimal load balancing scheme. Ramasubbareddy et al. proposed a pre-request mechanism based on the edge task response time to allocate task offloading through a central controller. Zhang proposed a fair offloading scheduling algorithm based on the service utility index. Jia et al. proposed a scheduling model based on a distributed genetic algorithm and a fast heuristic algorithm to reduce the maximum task response time and optimize the application performance. However, considering the actual situation, the number of edge servers in a wireless metropolitan area network is large and widely distributed. The centralized decision-making method requires the interconnection of each node, which not only increases the communication cost but also lengthens the decision-making time, making it difficult to solve the actual situation of the continuously dynamic changes in the loads of each edge.
[0008] There are also some works aiming to pre-train models through machine learning and deep learning algorithms to reduce the overall decision-making time of the edge network system. Gomez et al. proposed a load balancing method based on machine learning techniques, while Li et al. designed a model based on a convolutional neural network (CNN) to predict network traffic for load balancing. However, these methods require a large amount of real-world data to train high-precision models, and it is difficult to obtain sufficient data for training in practice, making it difficult to ensure that the load balancing schemes given by the models are effective.
[0009] In view of this, the present invention will use the DQN algorithm in reinforcement learning. The DQN algorithm introduces the concept of an experience replay pool, whose function is to store the experience data of the agent in the environment in a buffer area, randomly sample the samples during training, and allow the same experience data to be used for training multiple times, thereby improving the utilization rate of the data. In addition, the present invention adds federated learning on the basis of reinforcement learning, which can effectively handle the problems of uneven data distribution and concept drift among different participants, and proposes a multi-edge load balancing scheduling method applicable to wireless metropolitan area networks. Summary of the Invention
[0010] The purpose of the present invention is to propose a multi-edge load balancing task scheduling method based on federated reinforcement learning, which solves the problem of multi-edge load balancing in a distributed computing environment, in order to better balance the task load of each edge and shorten the maximum task response time in the edge system.
[0011] To achieve the above object, the technical solution of the present invention is: a multi-edge load balancing task scheduling method based on federated reinforcement learning, specifically including the following steps:
[0012] Step 1: Construct a Markov decision process based on the multi-edge load balancing task, and define a quadruple <S, A, T, R>, where S is the state space, A is the action space, T is the state transition function, and R is the reward function;
[0013] Step 2: Generate an independent agent, a target network model, a policy network model, and an experience replay pool required by the DQN algorithm for each edge node, as well as a global agent and a global policy network model required by federated learning;
[0014] The policy network model is used to estimate the Q value of executing different actions in different states; the target network structure is the same as that of the policy network;
[0015] Step 3: Iterative training of the local model:
[0016] Step 3.1: Each independent agent selects an action according to the current state of the edge node and executes it. The environment returns the next state and the reward value according to the action given by the agent, and stores the current state, action, reward, and the next state as experience data in the experience pool in advance;
[0017] Step 3.2: Update the current state according to the returned next state, and return to execute Step 3.1 until the capacity of the experience pool reaches the upper limit; randomly select a set of historical scheduling policy data from the experience pool, perform target network training on the independent agent and update the policy network parameters;
[0018] Step 4: When the number of local model iteration training reaches the first preset value, aggregate the policy network parameters of each independent agent to generate a federated global policy network model, and transmit the aggregated global policy network parameters back to each independent agent as the target network parameters for the next round of local model iteration training of each independent agent, and return to Step 3;
[0019] Step 5: Repeat Step 3 and Step 4 until the global policy network model converges, and obtain a Q-value prediction federated model for the edge node task offloading operation for multi-edge load balancing task scheduling.
[0020] Preferably, the construction of the Markov decision process based on the multi-edge load balancing task is specifically as follows:
[0021] Let the number of edge servers in the edge network system E be N, that is, E = {e1, e2, …, e N}, where e i represents the i-th edge node, then S i ∈ S represents the state space of the edge node e i :
[0022]
[0023] Among them, λ i represents the initial task volume of the edge e i ; represents the actual task load rate of the edge e i and its adjacent edges; F i represents the current load balancing scheduling scheme of the edge e i ;
[0024] A i ∈ A represents the action space of the edge node e i :
[0025]
[0026] Among them, the action means that the edge e i increases the task volume offloaded to the edge e N , and the action means that the edge e i reduces the task volume offloaded to the edge e N ; the action None means not to perform a scheduling operation; each scheduling operation transfers a fixed task volume δ;
[0027] T(s i , a i ) represents that the edge node e i selects the scheduling action a i in the state s iThe next state returned by the post - environment;
[0028] to represent the edge node e i at state s i select the scheduling action a i The reward value given by the post - environment:
[0029]
[0030] wherein, represents the new state returned by the environment, represents the edge node e i at state s i select the scheduling action a i enter the legal state after that, the reduction value of the maximum task response time in the edge network.
[0031] Preferably, the task response time of edge e i includes the task waiting time for all tasks to be executed on the edge and the total time for adjacent edges to transmit tasks is expressed as
[0032] Preferably, the task waiting time of each edge e i is composed of the task queuing time and the task execution time, and its calculation formula is:
[0033]
[0034] wherein, using the Erlang C formula, according to the number of resources n i of the server used by edge e i and the traffic intensity to calculate the probability that a unit task cannot be processed immediately and needs to wait; v i represents the service rate of edge e i ;
[0035] The task delay time of each edge e i is representing the total transmission time required when adjacent edges transmit tasks to this edge;
[0036]
[0037] wherein, L i represents the set of adjacent edges of edge e i ; represents edge e j transmit a unit task volume to edge ei The time required; f(j, i) represents the amount of tasks transmitted from edge e j to e i per unit time.
[0038] Preferably, when performing a certain action causes the task amount of a certain edge node to exceed the schedulable range, it is regarded as an illegal state.
[0039] Preferably, the policy network model adopts a deep neural network.
[0040] Preferably, in step 3, the target network of the independent agent continuously updates the model parameters as the local model is iteratively trained. When the number of local model iterative training reaches the second preset value, the parameters of the policy network of the independent agent are replaced with the parameters of the target network of the independent agent.
[0041] Preferably, in step 3, each independent agent randomly selects an action using the ε-greedy policy or selects and executes the action with the largest Q value according to the action-Q value table predicted by the target network.
[0042] Preferably, the policy networks of all independent agents adopt the same neural network structure, and the aggregation of the global policy network model adopts the federated averaging method:
[0043]
[0044] where θ g and θ i are the network parameters of the global model and the local model of edge node e i respectively.
[0045] Preferably, the target network of the independent agent updates the network parameters using the Huber loss function.
[0046] Compared with the prior art, the present invention has the following beneficial effects:
[0047] The present invention is a load balancing scheduling method (FRL-MCLB) based on federated reinforcement learning and predictive feedback control mechanism. The FRL-MCLB method combines federated learning and reinforcement learning, and performs unsupervised learning based on the historical operation data and scheduling operations of local edge nodes; allows agents to train their respective local models in a distributed manner, and then send the local models to the central aggregation unit to build a global model; the present invention solves the problem of multi-edge load balancing in a distributed computing environment, in order to better balance the task load of each edge and shorten the maximum task response time in the edge system. Different from centralized decision-making, FRL-MCLB only needs each agent to upload model parameters, which greatly reduces the communication cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 Schematic diagram of the edge network system structure in an embodiment of the present invention;
[0049] Figure 2 Specific scenario diagram of six edge network systems in an embodiment of the present invention;
[0050] Figure 3 Comparison diagram of the effect of load balancing scheduling operations of four algorithms in different scenarios on reducing the maximum task response time of the system in an embodiment of the present invention. Specific implementation manner
[0051] The following combines the attached Figures 1-3 to specifically describe the technical solution of the present invention.
[0052] The present invention proposes a multi-edge load balancing task scheduling method based on federated reinforcement learning, which specifically includes the following steps:
[0053] Step 1: Construct a Markov decision process based on the multi-edge load balancing task, and define a quadruple <S, A, T, R>, where S is the state space, A is the action space, T is the state transition function, and R is the reward function;
[0054] Step 2: Generate an independent agent, a target network model, a policy network model, and an experience replay pool required by the DQN algorithm for each edge node, as well as a global agent and a global policy network model required by federated learning;
[0055] The policy network model is used to estimate the Q value of executing different actions in different states; the target network structure is the same as that of the policy network;
[0056] Step 3: Iterative training of the local model:
[0057] Step 3.1: Each independent agent selects an action according to the current state of the edge node and executes it. The environment returns the next state and the reward value according to the action given by the agent, and stores the current state, action, reward, and the next state as experience data in the experience pool in advance;
[0058] Step 3.2: Update the current state according to the returned next state, and return to execute Step 3.1 until the capacity of the experience pool reaches the upper limit; randomly select a set of historical scheduling policy data from the experience pool, perform target network training on the independent agent, and update the policy network parameters;
[0059] Step 4: When the number of local model iteration training reaches the first preset value, aggregate the policy network parameters of each independent agent to generate a federated global policy network model, and transmit the aggregated global policy network parameters back to each independent agent as the target network parameters for the next round of local model iteration training of each independent agent, then return to Step 3;
[0060] Step 5: Repeat Step 3 and Step 4 until the global policy network model converges, and obtain a Q-value prediction federated model for edge node task offloading operations for multi-edge load balancing task scheduling.
[0061] Assume that the service provider has established an edge network system as shown in Figure 1 in a certain area within the city. These edge servers are interconnected through the wireless network in the wireless metropolitan area network, and the edges can communicate and transfer data with each other. The user's mobile device can also directly access the nearby wireless network access point, and then enjoy the computing and storage resource services provided by the edge server. When the task executed by the user device needs to be offloaded to the edge server for execution to improve the application performance, the edge network can spontaneously perform load balancing to minimize the task response time as much as possible. When the mobile device offloads a task, it will default to finding the edge server closest to itself. Assume that the task to be executed can be freely divided into multiple parts. The edge server can, according to its current operating conditions and the load conditions of the edge network, split the initial task into multiple subtasks and distribute them to other edge servers for parallel execution. After the execution is completed, the initial edge server will integrate and return the final operation result to the user's mobile device.
[0062] Assume that the N edge servers in this area are respectively E = {e1, e2, …, e N}, where e i represents the i-th edge node. The edge node directly connected to a certain edge through the network is called an adjacent edge. Each edge e i has at most N - 1 adjacent edges. To better describe the state of the edge, this method gives Definition 1 and Definition 3 to describe the static data and dynamic data of edge e i respectively.
[0063] Definition 1: The static data of edge e i consists of a triple:
[0064] c i = {v i , λ i , L i}, i ∈ {1, 2, …, N} (1)
[0065] where c i represents edge ei Static data, v i Denote the edge e i The service rate, where the number of servers within the edge and their respective service rates are ignored, and only the overall service rate that the edge e i can provide is considered, that is, the total amount of tasks that the edge ei can process per unit time; λ i Denote the edge e i The initial task arrival volume, that is, the total amount of tasks unloaded by nearby mobile devices received by e i per unit time (this concept will be represented by "initial task" hereinafter).
[0066] Definition 2: 1 ≤ k ≤ N - 1, L i Denote the set of adjacent edges of the edge e i where Denote the k-th edge adjacent to e i .
[0067] Definition 3: The dynamic data of the edge e i consists of a triple:
[0068] p i = {D i , F i , w i}, i ∈ {1, 2, …, N} (2)
[0069] where p i represents the dynamic data of the edge e i .
[0070] Definition 4: In the process of mutual communication and task distribution among edge servers in the network, there will be a certain network delay, and this delay will change in real time according to the different traffic in the network. Therefore, the network delay matrix of the edge system in this method is represented by D ∈ R N×N as follows:
[0071]
[0072] where denotes the time required for the edge e i to transmit a unit amount of tasks to the edge e j . This is a symmetric matrix, so In particular, when i = j, When the edge e i and the edge e j are not directly adjacent,
[0073] Definition 5: Denote the edge e i The set of network delays between it and other edges.
[0074] Definition 6: The load balancing scheme for the entire edge network is represented by F ∈ R N×N as follows:
[0075]
[0076] where f(i, j) is used to represent the amount of tasks transmitted from edge e i to e j per unit time. Specifically, when i = j, represents the amount of tasks actually processed by edge e i per unit time.
[0077]
[0078] Its meaning is that the actual amount of tasks of edge e i is equal to the initial amount of tasks of the edge minus the amount of tasks scheduled to adjacent edges plus the amount of tasks scheduled to this edge by adjacent edges; when edge e i is not adjacent to edge e j , the value of is 0.
[0079] Definition 7: represents the load balancing scheme of edge e i .
[0080] Definition 8: W = {w1, w2,..., w n},
[0081] where W represents the actual task load rate of the edge system per unit time, and w i represents the current actual task load rate of edge e i (hereinafter, this concept will be represented by "load rate"). When the initial tasks arrive at the edge, the initial amount of tasks on the edge is equal to the actual amount of tasks. After the system performs load balancing, the edge will schedule the initial tasks to adjacent edges, and the actual amount of tasks processed on the edge will change, and the load rate will also change accordingly.
[0082] Definition 9: The task waiting time i of each edge e consists of the task queuing time and the task execution time, and its calculation formula is:
[0083]
[0084] where, The Erlang C formula is used, which calculates the probability that a unit task cannot be immediately processed and needs to wait based on the number of resources n of the server used at the edge i and the traffic intensity to calculate the probability that a unit task cannot be immediately processed and needs to wait
[0085] Definition 10: Assume that the tasks scheduled to adjacent edges can be arbitrarily divided into data packets of the same size, so that the network delay generated by transmitting a unit task volume between a pair of edges is the same. Then for each edge e i the task delay time is which represents the total transmission time required when the adjacent edge transmits the task to this edge
[0086]
[0087] Definition 11: Use to represent the task response time of edge e i including the task waiting time for all tasks to be completed on the edge plus the total time for the adjacent edge to transmit the task
[0088] The symbol definitions involved in the present invention are shown in Table 1
[0089] Table 1
[0090]
[0091]
[0092] The present method gives the following formal definition for the "multi-edge system load balancing problem": First, deploy N edge nodes on the wireless metropolitan area network in a certain area, and the service rate provided by each edge node is v i , when the edge provides computing and storage resource services for mobile devices in this area, it receives user request tasks. And the initial task arrival volume λ of each edge i is different. In order to make full use of the resources of each edge node, it is necessary to give the best load balancing scheme F to balance the load of the edge system and improve the user experience in this area. Therefore, the goal of the present method is to find the respective load scheme F i for each edge e i such that the maximum task response time in the edge network is minimized
[0093] In this embodiment, the specific construction of the Markov decision process based on multi-edge load balancing tasks is as follows
[0094] S i ∈S represents the edge node e iState space:
[0095]
[0096] where λ i represents the initial task volume of edge e i . represents the actual task load rate of edge e i and its adjacent edges; F i represents the current load balancing scheduling scheme of edge e i .
[0097] A i ∈A represents the action space of edge node e i :
[0098]
[0099] where the action represents an increase in the task volume unloaded to edge e i from edge e N , the action represents a decrease in the task volume unloaded to edge e i from edge e N ; the action None represents no scheduling operation; each scheduling operation transfers a fixed task volume δ;
[0100] T(s i ,a i ) represents the state transition function, where s i represents the current state of edge e i , a i represents the action selected by the DQN algorithm in the current state s i ; that is, T(s i ,a i ) represents the next state returned by the environment after edge node e i selects the scheduling action a i in state s i ;
[0101] Suppose edge e1 has two adjacent edges e2 and edge e3, and the current state is
[0102]
[0103] At this time, the selected action is that edge e1 schedules a task volume to edge e3, and the scheduled amount is δ. Therefore, at this time, the value returned by the T(s i ,a i ) state transition function is
[0104]
[0105] is used to represent the edge node e i select the scheduling action a in the state s i and the reward value given by the environment after that: i
[0106]
[0107] where, represents the new state returned by the environment, represents the edge node e i select the scheduling action a in the state s i and enter the legal state i and the reduction value of the maximum task response time in the edge network.
[0108] In this embodiment, DQN (Deep Q-Network) applies a deep neural network (Deep Neural Networks) to the Q-Learning algorithm, realizing the modeling and decision-making of high-dimensional state spaces in complex environments. DQN uses a deep neural network as the policy network of the agent to estimate the Q-values of performing different actions in different states. An experience replay pool is used to store the interaction experience data, including states, actions, rewards, and the next state, so that random sampling can be performed during training to improve the utilization rate of samples. A target network (Target Network) is introduced in DQN to calculate the target Q-value. The target network has the same structure as the policy network, but its parameters are continuously updated during the iteration process. When the preset number of iterations is reached, the parameters of the policy network are replaced with those of the target network, thereby reducing the fluctuation of the target value during the training process. The ε-greedy policy (Epsilon-Greedy Policy) is used. When selecting an action using the Q-value estimated by the current policy network, the optimal action is selected with a certain probability policy, thus achieving a balance between exploration and exploitation.
[0109] In this embodiment, the policy networks of all independent agents adopt the same neural network structure, and the aggregation of the global policy network model adopts the federated averaging method:
[0110]
[0111] where, θ g and θ i are the network parameters of the global model and the local model of the edge node e i respectively.
[0112] This method is a load balancing scheduling method based on federated reinforcement learning and predictive feedback control mechanism (FRL-MCLB). The FRL-MCLB method combines federated learning and reinforcement learning, and conducts unsupervised learning based on the historical operation data and scheduling operations of local edge nodes. The method allows agents to train their respective local models in a distributed manner, and then send the local models to the central aggregation unit to build a global model. Different from centralized decision-making, FRL-MCLB only needs each agent to upload model parameters, which greatly reduces the communication cost. In FRL-MCLB, the local models of all agents have the same neural network structure, and they are aggregated to form a global model by federated averaging method. This training process includes the following two parts:
[0113] 1. Distributed training: Each agent customizes a standard deep Q-network (DQN) to train the local model. After local training, the model parameters are extracted from each agent and sent to the central aggregation unit.
[0114] 2. Federated aggregation: When the central unit receives the model parameters of each agent, it aggregates these parameters to generate a global model. Then, the global model parameters are distributed to all agents to update their local models. The present invention uses the federated averaging algorithm (FedAvg) for model aggregation
[0115] Based on the above definition, the present invention uses FRL-MCLB to train a federated model for evaluating the Q-values of all edge node task offloading operations. The main steps are shown in Algorithm 1 of Table 2. The DQN parameters are randomly initialized (line 3) and trained on the dataset (line 5). In each run-time interaction (line 6), an action a is randomly selected with a probability of ∈ i . The reward is calculated and a new state is generated (lines 9-10). The experience is stored in the replay buffer (line 11). The network parameters are updated using mini-batch data and the Huber loss function (lines 14-15). The local network parameters are periodically uploaded to the central unit (line 18). The central unit uses FedAvg to fit the global Q-network (line 19). The network parameters are distributed to the edge node agents to update the local network (lines 19-21). Finally, the algorithm continues to train until convergence.
[0116] Table 2
[0117]
[0118]
[0119] Experimental simulation:
[0120] Regarding the load balancing problem of multi-edge systems in wireless metropolitan area networks, this method uses a load balancing scheduling method based on federated reinforcement learning and predictive feedback control mechanism (hereinafter referred to as "FRL-MCLB method"). This section will evaluate the feasibility and effectiveness of the FRL-MCLB method through simulation and experimental studies.
[0121] Both the FRL-MCLB method and the comparative algorithms were implemented in the Python 3.8 environment and run on a Windows 11 system equipped with an Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz CPU and 16 GiB of RAM.
[0122] The parameter settings of the FRL-MCLB method are as follows: the size of the pre-stored data in the experience pool REPLAY_MEMORY_SIZE = 2000 DQN, the size of the samples obtained by the agent BATCH_SIZE = 64, the learning rate of the DQN algorithm LEARNING_RATE = 0.001, and the decay factor of the reward function GAMMA = 0.99.
[0123] The simulation experiment of this method is based on the distribution coordinates of the telecom 5G wireless base stations in XX City, where the maximum number of adjacent edges for each edge node is limited to 3. Considering the cases of the number of edges N = 5 and N = 15, three scenarios are created respectively for comparative experiments. Six specific scenarios are as Figure 2 shown, where the blue nodes represent the edges, and the values on the edges represent the network delay of the unit task volume between the edge servers.
[0124] The experiment sets the service rate v of each edge i from the normal distribution N(15, 6), and the task arrival rate also satisfies the normal distribution of N(15, 6), but the upper limit is set to the edge service rate v i - 0.25. The unit task transfer time between the edges is within the interval [0.10, 0.19].
[0125] The experiment randomly divides the effective simulation data (about 5000 pieces) into a training set (75%) and a test set (25%). To avoid insufficient sample richness at the beginning of training, data is first filled into the experience replay pool according to the size of REPLAY_MEMORY_SIZE. During iterative training, the environment determines the reward value reward of the action according to the set reward function. The number of iterations of DQN is set to 3000 times, and the target network parameters of the edges are federally called back every 50 iterations.
[0126] The experiment compares the load balancing scheduling performance of the FRL-MCLB method, the greedy algorithm and the random migration algorithm (RMA). Ten independent repeated experiments were conducted in six different scenarios, and then analyzed from the perspective of the system's maximum task response time and program execution time. The greedy algorithm strategy used in this method is to select the edge with the largest task response time for task scheduling each time, while the strategy used by the RMA algorithm is to randomly divide the range of tasks to be scheduled for each edge and select the final scheduled task amount.
[0127] Table 3 lists the comparison of the maximum task response time of the edge network system after scheduling by the FRL-MCLB method and the comparison algorithm and without scheduling.
[0128] Table 3
[0129]
[0130] Figure 3 It more intuitively reflects the degree of improvement in the load balancing scheduling operation of the FRL-MCLB method in reducing the maximum task response time of the system in different scenarios. It can be seen that this method can significantly reduce the maximum task response time of the edge, and compared with the greedy algorithm, it is shorter in most scenarios, and the maximum improvement ratio reaches 29.52%, while the maximum improvement ratio compared with the random migration algorithm reaches 42.63%. The reason for this is that because the FRL-MCLB method will continue to trial and error and accumulate experience, the operation that will increase the maximum task response time will result in a decrease in the accumulated reward value. As the number of model iterations increases, the experience becomes richer, and the algorithm will try its best to avoid scheduling operations that increase the maximum task response time.
[0131] In summary, this method combines federated learning, reinforcement learning and predictive feedback control mechanism to design a new load balancing scheduling algorithm to solve the problem of multi-edge load balancing in a distributed computing environment, so as to better balance the task load of each edge and shorten the maximum task response time in the edge system. At the same time, through experimental comparison and evaluation and analysis, the superiority and effectiveness of this method are proved, which ensures the service quality of users.
[0132] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions do not exceed the scope of the technical solution of the present invention, belong to the protection scope of the present invention.
Claims
1. A multi-edge load balancing task scheduling method based on federated reinforcement learning, characterized in that Specifically, it includes the following steps: Step 1: Construct a Markov decision process based on the multi-edge load balancing task, and define a quadruple <S, A, T, R>, where S is the state space, A is the action space, T is the state transition function, and R is the reward function; Step 2: Generate an independent agent, a target network model, a policy network model, and an experience replay pool required by the DQN algorithm for each edge node, as well as a global agent and a global policy network model required by federated learning; The policy network model is used to estimate the Q-values of executing different actions in different states; the target network structure is the same as that of the policy network; Step 3: Iterative training of the local model: Step 3.1: Each independent agent selects and executes an action according to the state of the current edge node. The environment returns the next state and the reward value according to the action given by the agent, and pre-stores the state, action, reward, and the next state as experience data in the experience pool; Step 3.2: Update the current state according to the returned next state, and return to execute Step 3.1 until the capacity of the experience pool reaches the upper limit; randomly select a set of historical scheduling policy data from the experience pool, perform target network training on the independent agent, and update the policy network parameters; Step 4: When the number of iterations of the local model training reaches the first preset value, aggregate the policy network parameters of each independent agent to generate a federated global policy network model, and transmit the aggregated global policy network parameters back to each independent agent as the target network parameters for the next round of local model iterative training of each independent agent, and return to Step 3; Step 5: Repeat Step 3 and Step 4 until the global policy network model converges, and obtain a federated model for predicting the Q-values of the edge node task offloading operations, which is used for multi-edge load balancing task scheduling.
2. The multi-edge load balancing task scheduling method based on federated reinforcement learning according to claim 1, wherein The construction of the Markov decision process based on the multi-edge load balancing task is specifically as follows: Let the number of edge servers in the edge network system E be N, i.e., E = {e1, e2, …, e N}, where e i represents the i-th edge node, then S i ∈ S represents the state space of the edge node e i : Among them, λ i represents the initial task volume of edge e i ; represents the actual task load rate of edge e i and its adjacent edges; F i represents the current load balancing scheduling scheme of edge e i ; A i ∈A represents the edge node e i 's action space: Among them, the action represents the edge e i increases the amount of tasks unloaded to the edge e N ; the action represents the edge e i decreases the amount of tasks unloaded to the edge e N ; the action None means no scheduling operation is performed; each scheduling operation transfers a fixed amount of tasks δ T(s i ,a i ) represents the next state returned by the environment after the edge node e i selects the scheduling action a i in the state s i ; to represent the edge node e i in state s i to select the scheduling action a i the reward value given by the environment after that: wherein, represents the new state returned by the environment, represents the edge node e i selects the scheduling action a i in the state s i to enter the legal state and then the reduction value of the maximum task response time in the edge network.
3. The multi-edge load balancing task scheduling method based on federated reinforcement learning according to claim 2, wherein Edge e i The task response time of includes the task waiting time for all tasks to be completed on the edge and the total time for adjacent edge transmission tasks 4. The method for multi-edge load balancing task scheduling based on federated reinforcement learning according to claim 3, wherein The task waiting time of each edge e i is composed of the task queuing time and the task execution time, and its calculation formula is: The task waiting time of each edge e is composed of the task queuing time and the task execution time, and its calculation formula is: Among them, Using the Erlang C formula, based on the edge e i the number of resources n of the server used i and the traffic intensity to calculate the probability that a unit task cannot be immediately accepted for processing and needs to wait; v i represents the service rate of the edge e i of the service rate; For each edge e i the task latency time is which represents the total transmission time required when the adjacent edge transmits the task to this edge; Among them, L i represents the set of adjacent edges of edge e i . represents the time required for edge e j to transmit a unit task volume to edge e i ; f(j, i) represents the task volume transmitted by edge e j to e i per unit time.
5. The method for multi-edge load balancing task scheduling based on federated reinforcement learning according to claim 2, wherein When executing a certain action causes the task volume of a certain edge node to exceed the schedulable range, it is regarded as an illegal state.
6. The multi-edge load balancing task scheduling method based on federated reinforcement learning according to claim 1, wherein The policy network model uses a deep neural network.
7. The multi-edge load balancing task scheduling method based on federated reinforcement learning according to claim 1, wherein In Step 3, the model parameters of the target network of the independent agent are continuously updated with the iterative training of the local model. When the number of iterations of the local model training reaches the second preset value, the parameters of the policy network of the independent agent are replaced with the parameters of the target network of the independent agent.
8. The multi-edge load balancing task scheduling method based on federated reinforcement learning according to claim 1, characterized in that In Step 3, each independent agent randomly selects an action using the ε-greedy strategy or selects the action with the largest Q-value according to the action-Q value table predicted by the target network and executes it.
9. The method for multi-edge load balancing task scheduling based on federated reinforcement learning according to claim 6, wherein The policy networks of all independent agents adopt the same neural network structure, and the aggregation of the global policy network model adopts the federated averaging method: Among them, θ g and θ i are the network parameters of the global model and the local model of the edge node e i respectively.
10. The method for multi-edge load balancing task scheduling based on federated reinforcement learning according to claim 9, wherein The target network of the independent agent updates the network parameters using the Huber loss function.