Industrial edge computing system load balancing method based on deep reinforcement learning

By introducing central scheduler and deep reinforcement learning into the edge computing system, the problem of load imbalance is solved, efficient task scheduling and system balance are achieved, and service quality and real-time are improved.

CN120336009APending Publication Date: 2025-07-18WEST ANHUI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415692.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

When facing the high dynamic and randomness of task generation in the industrial environment, the existing edge computing scheduling mechanism leads to load imbalance, affecting service quality and computing resource utilization. In addition, the traditional optimization methods have high computational complexity and unstable strategy convergence, making it difficult to meet the real-time and stability requirements of industrial scenarios.

Method used

The load balancing method based on deep reinforcement learning is adopted, and tasks are scheduled from edge servers with high load to edge servers with low load through the central scheduler. The PPO algorithm is used to generate scheduling actions, and the system load balancing is achieved through long-term discount utility model optimization strategy, reducing the computational complexity of state and action space, and reducing coordination delays between servers.

Benefits of technology

It significantly improves the service quality of edge computing systems, reduces the average running delay by more than 15%, achieves higher stability and adaptability, and supports real-time scheduling of more than 20 edge servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336009A_ABST
    Figure CN120336009A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial edge computing system load balancing method based on deep reinforcement learning. According to the method, real-time scheduling of more than 20 edge servers is supported by converting discrete task queue length into calculation time deviation, reducing state dimension, limiting actions to be whole-quantity scheduling, avoiding invalid action exploration, balancing instant and future benefits through discount factors, improving strategy stability and reducing coordination overhead among the servers. Compared with an existing distributed method and a centralized optimization algorithm, the method has the advantages that the calculation complexity is reduced by reconstructing the state space and the action space, the problem of explosion of the action space under a large-scale system is solved, the coordination delay between edge servers is reduced, the average operation delay of the system is reduced by more than 15%, and the system reliability is improved. And the service quality of the edge computing system is obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of communication and edge computing, and particularly relates to a load balancing method for an industrial edge computing system based on deep reinforcement learning. Background Art

[0002] In recent years, with the rapid development of the fifth-generation mobile communication technology (5G) and the wide application of the industrial Internet of Things (IIoT) in intelligent manufacturing, mobile edge computing (MEC), as a distributed computing architecture that sinks computing and storage capabilities to the network edge and close to the data source, has become an important technical means to promote the intelligent upgrade of industry. In the industrial production process, terminal devices can unload a large amount of task data generated in real time to an edge server (ES) deployed in the factory area through the 5G network for local processing, thereby effectively reducing computing latency, alleviating network congestion, and improving data processing efficiency. Currently, this technology has been applied in many industrial scenarios such as equipment predictive maintenance, production quality monitoring, and intelligent control. For example, Intel has deployed edge servers inside its manufacturing workshop for equipment data collection and analysis, achieving rapid fault detection and intervention.

[0003] However, most of the existing edge computing scheduling mechanisms adopt a distributed scheduling method, that is, the edge servers perform autonomous negotiation-based scheduling based on strategies such as game theory and heuristic algorithms. Due to the high dynamics and randomness of task generation in the industrial environment, the loads of each edge server may be seriously unbalanced at different times, affecting the overall service quality and computing resource utilization rate. In addition, although some studies have introduced a central scheduler model to connect multiple edge servers in the region through a fiber-optic network and uniformly coordinate task allocation, in the face of a high-dimensional state space and complex scheduling actions, traditional optimization methods (such as particle swarm optimization and greedy algorithms) still have problems such as high computational complexity, unstable strategy convergence, and poor adaptability. With the continuous maturity of deep reinforcement learning technology, some studies have begun to try to use methods such as DQN and Actor-Critic to achieve cross-slot scheduling optimization, but due to factors such as a large state dimension, a complex action space, and poor training stability, it is still difficult to meet the dual requirements of real-time and stability in industrial scenarios. Therefore, there is an urgent need for a load balancing scheduling method with high stability, strong generalization ability, and the ability to efficiently adapt to dynamic task changes to better serve the intelligent scheduling needs in the industrial edge computing system. Summary of the Invention

[0004] In order to overcome the shortcomings and deficiencies of the prior art, the purpose of the present invention is to provide a load balancing method for an industrial edge computing system based on deep reinforcement learning.

[0005] The present invention is implemented as follows. A load balancing method for an industrial edge computing system based on deep reinforcement learning, the method comprising the following steps:

[0006] S1. At the beginning of each time slice, each edge server reports to the central scheduler the number of tasks and computing capabilities in the current queue;

[0007] S2. The central scheduler defines the system state as the deviation between the computing time of each edge server and the average computing time of the system, and generates a task scheduling action through the proximal policy optimization algorithm;

[0008] S3. The central scheduler schedules tasks from edge servers with higher loads to edge servers with lower loads according to the action decision, and ensures that the scheduled task volume is an integer and satisfies the queue capacity constraint;

[0009] S4. Based on the long-term discounted utility model to optimize the policy, the central scheduler realizes the load balancing of the system by minimizing the variance of the system running time.

[0010] Preferably, in step S1, at the beginning of each time slice, each edge server reports the following information to the central scheduler:

[0011] (1) The number of tasks Q in the current queue i and the average computing amount Z of each task;

[0012] (2) The resources of each edge server, such as CPU frequency, memory, etc.;

[0013] (3) Obtaining the computing time of each edge server

[0014] (4) State construction: The central scheduler calculates the average computing time of the system and generates a continuous state vector: where N is the total number of edge servers.

[0015] Preferably, in step S2, the central scheduler generates a scheduling action based on the PPO algorithm, which specifically includes the following contents:

[0016] (1) Policy network inference: Input the state s t , and output the action probability distribution π(a|s t );

[0017] (2) Action constraint: The action space for edge server i defined by the central scheduler is the integer task volume a i ∈A, and satisfies

[0018] (3) Queue capacity constraint: The task volume of each edge server after scheduling needs to satisfy Qi +a i ≥ 0 and Q i +a i ≤ Q max .

[0019] Preferably, in step S3, the central scheduler issues a scheduling instruction to each edge server to perform task migration, including:

[0020] (1) Computing time update: According to the actual scheduling result, each server updates its local computing time T i , and feeds it back to the central scheduler;

[0021] (2) Reward function calculation: The system reward r is defined as: where C max is the maximum deviation allowed by the system.

[0022] Preferably, in step S4, the central scheduler optimizes the policy through a discounted utility model to perform the following operations:

[0023] Trajectory collection: Collect (s t , a t , r t , s t+1 ) data for multiple time slices to form an experience pool;

[0024] Advantage estimation: Use Generalized Advantage Estimation (GAE) to calculate the advantage value

[0025] δ t = r(t) + βV(s t+1 ; θ v ) - V(s t ; θ v );

[0026]

[0027] where V(s t ; θ v ) represents the value function estimation of the current state s t , V(s t+1 ; θ v ) represents the value function estimation of the next state s t+1 , and λ is the GAE decay factor;

[0028] Policy update: Optimize the policy network through the clipped objective function:

[0029]

[0030] Calculate the value function loss:

[0031] LV (θ) = (r(t) - V(s t ; θ v )) 2 ;

[0032] Perform gradient descent to update the parameter θ v , and loop until convergence.

[0033] The present invention overcomes the deficiencies of the prior art and provides a load balancing method for an industrial edge computing system based on deep reinforcement learning. The core improvements of this method include:

[0034] (1) State space reconstruction: Convert the discrete task queue length into a computational time deviation to reduce the state dimension;

[0035] (2) Action space constraint: Limit the action to integer quantity scheduling to avoid exploration of invalid actions;

[0036] (3) Long-term utility optimization: Balance immediate and future rewards through a discount factor to improve policy stability;

[0037] (4) Central scheduler design: Reduce the coordination overhead between servers and support real-time scheduling of more than 20 edge servers.

[0038] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects: Compared with the existing distributed methods and centralized optimization algorithms, the present invention reduces the computational complexity by reconstructing the state space and action space, solves the problem of action space explosion in large-scale systems, and at the same time reduces the coordination delay between edge servers, reducing the average running delay of the system by more than 15%, significantly improving the quality of service of the edge computing system. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the operation process of the PPO deep reinforcement learning algorithm of the present invention after training;

[0040] Figure 2 is the training performance of the PPO deep reinforcement learning algorithm of the present invention under different numbers of ES;

[0041] Figure 3 is the performance comparison between the PPO deep reinforcement learning algorithm of the present invention and the particle swarm optimization algorithm and the greedy algorithm;

[0042] Figure 4 is the running time of multiple ESs of the present invention method under different methods in the same time slot. DETAILED DESCRIPTION OF THE INVENTION

[0043] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] An embodiment of the present invention discloses a load balancing method for an industrial edge computing system based on deep reinforcement learning. The method includes the following specific steps:

[0045] S1. At the beginning of each time slice, each edge server reports the number of tasks and computing capabilities in the current queue to the central scheduler.

[0046] In step S1, at the beginning of each time slice, each edge server reports the following information to the central scheduler:

[0047] (1) The number of tasks Q in the current queue i and the average computing amount Z of each task i ;

[0048] (2) The resources of each edge server, such as CPU frequency, memory, etc.;

[0049] (3) Obtain the computing time of each edge server

[0050] (4) State construction: The central scheduler calculates the average computing time of the system and generates a continuous state vector: where N is the total number of edge servers.

[0051] S2. The central scheduler defines the system state as the continuous deviation between the computing time of each edge server and the average computing time of the system, and generates a task scheduling action through the Proximal Policy Optimization algorithm (PPO).

[0052] In step S2, the PPO algorithm parameter configuration is specifically as follows:

[0053] (1) Network structure:

[0054] Policy network (Actor): It includes 2 fully connected layers, with 500 neurons in each layer. The activation function is ReLU, and the output layer uses Softmax to generate the action probability distribution.

[0055] Value function network (Critic): It has the same structure as the policy network, and the output is the state value function estimate.

[0056] (2) Training parameters:

[0057] Discount factor γ = 0.99, GAE parameter λ = 0.95, clipping parameter ∈ = 0.2.

[0058] Batch size = 64, number of training epochs = 5, learning rate α = 3×10 -4 。

[0059] The entropy regularization coefficient β = 0.001 is used to enhance the policy exploration ability.

[0060] In step S2, the central scheduler generates scheduling actions based on the PPO algorithm, which specifically includes the following:

[0061] (1) Policy network inference: Input the state s t , and output the action probability distribution π(a|s t );

[0062] (2) Action constraint: The action space for edge partitioner i is defined as the integer task volume a i ∈A, and it satisfies For example, if a1 = +3, then 3 tasks need to be transferred from edge server 1 and assigned to other edge servers (such as a2 = -1, a3 = -2).

[0063] (3) Queue capacity constraint: The task volume of each edge server after scheduling needs to satisfy Q i +a i ≥0 and Q i +a i ≤Q max 。

[0064] S3. The central scheduler schedules tasks from edge servers with higher loads to those with lower loads according to the action decision, and ensures that the scheduled task volume is an integer and satisfies the queue capacity constraint;

[0065] In step S3, the central scheduler issues scheduling instructions to each edge server to execute task migration, including:

[0066] (1) Computation time update: According to the actual scheduling results, each server updates its local computation time T i , and feeds it back to the central scheduler;

[0067] (2) Reward function calculation: The system reward r is defined as: Among them, C max is the maximum deviation allowed by the system.

[0068] S4. Optimize the policy based on the long-term discounted utility model, and the central scheduler achieves load balancing for the multi-time slice system by maximizing the long-term discounted revenue.

[0069] In step S4, the long-term discounted utility is:

[0070]

[0071] Among them, t0 is the initial time, s0 is the initial state, k is the time slice index, r(t0 + k + 1) is the return of time slice t0 + k + 1, and β is the discount factor.

[0072] In step S4, the central scheduler optimizes the policy through the discounted utility model as follows:

[0073] Trajectory collection: Collect (s t , a t , r t , s t+1 ) data of multiple time slices to form an experience pool;

[0074] Advantage estimation: Calculate the advantage value using Generalized Advantage Estimation (GAE)

[0075] δ t = r(t) + βV(Sm t+1 ; θ v ) - V(Sm t ; θ v )

[0076]

[0077] Among them, V(s t ; θ v ) represents the value function estimation of the current state s t , V(s t+1 ; θ v ) represents the value function estimation of the next state s t+1 , and λ is the GAE decay factor;

[0078] Policy update: Optimize the policy network through the clipped objective function:

[0079]

[0080] Calculate the value function loss:

[0081] L V (θ) = (r(t) - V(Sm t ; θ v )) 2

[0082] Execute gradient descent to update the parameter θ v , and loop until convergence.

[0083] Application Example

[0084] 1. Training process

[0085] When training with the PPO algorithm, the parameters of each neuron in the PPO network are obtained after training for multiple time slices. The process is as follows: At the beginning of a time slice, the edge server reports the number of tasks in the existing queue to the central server. Then, the central server selects an action based on the output of the network. This action is to schedule among multiple servers, with the goal of making each server run more evenly. Since tasks arrive randomly in each time slice, the running time of task on edge server i in this time slice is:

[0086]

[0087] where Q a represents the original number of tasks in the queue, a o is the number of tasks transferred out by server i, a c is the number of tasks transferred to server i by other servers, and λ i is the number of tasks arriving at server i in this time slot.

[0088] The variance of the running time of each edge server minus the average running time of each server in this time slot is the cost of the system. Subtracting this cost from the maximum tolerable deviation of the system gives the system's revenue. Based on this revenue, the neural network is trained using the following learning method, and then the system enters the second time slice. Similarly, the central server performs similar operations.

[0089] Until the training algorithm converges, at this time the connection weights of each neuron in the PPO deep reinforcement learning have been determined. At this time, after scheduling in each time slice, the system will approximately reach equilibrium in each time slice. That is to say, during training, multiple time slices are used for training. After the network parameters are determined, scheduling in one time slice can make the system reach approximate equilibrium.

[0090] 2. Performance Comparison

[0091] As Figure 1 shown, Figure 1 shows the operation process of the PPO algorithm of the present invention after training. Among them, in each time slice, the edge server reports the tasks in its respective queue to the central scheduler. Then, the scheduler issues scheduling instructions to each edge server according to the PPO policy module, and then each edge server performs scheduling to achieve load balancing.

[0092] The present invention first compares the training performance of the PPO deep reinforcement learning algorithm (PPO-DLR) under different numbers of edge servers (Edge Sever ES). The results are as Figure 2 shown, Figure 2 which shows the long-term cumulative reward during the training process within 500 time slices. From Figure 2As can be seen, when there are 20 ESs in the system, the algorithm has the fastest convergence speed and converges in 200 time slices. In contrast, when there are 28 ESs, the algorithm has the slowest convergence, and convergence occurs at approximately 350 time slices. Figure 2 It shows that the fewer the number of ESs in the system, the faster the algorithm converges. This is because the fewer the ESs, the less the system input, thus reducing the training time.

[0093] To evaluate the performance of this algorithm, the present invention compares it with the Particle Swarm Optimization (PSO) algorithm and the Greedy algorithm, and the results are as Figure 3 shown. Figure 3 It shows the average running time of 100 Episodes. Due to the randomness of task arrivals among different ESs, there will be some fluctuations in the average running time of the system. In addition, when using PPO deep reinforcement learning, the average running time of the system is the shortest, about 2.5 milliseconds. When using particle swarm optimization, the average running time of the system is the second lowest, about 4 milliseconds. Therefore, the PPO deep reinforcement learning algorithm of the present invention makes the average running time of the system the shortest.

[0094] In the embodiments of the present invention, 10 ESs are randomly selected from these groups for comparison, and the results are as Figure 4 shown. Figure 4 It shows the running time of multiple ESs in the same time slot under different methods. As can be seen from Figure 4 it, when the PPO deep reinforcement learning algorithm of the present invention is adopted, the running time of each ES is usually consistent, while when using the PSO and greedy algorithms, there are significant differences in the running time, and the largest difference is observed under the greedy algorithm. In addition, Figure 4 it shows that both the PSO and greedy algorithms require more time, mainly due to the accumulation of tasks in previous time slots, resulting in longer running times for multiple ESs. Therefore, by modeling the system state as the deviation between the computing time of ESs and the average computing time of the system and using a central scheduler for task scheduling, the present invention successfully reduces the state and action algorithms. The PPO deep reinforcement learning algorithm not only achieves a more balanced task allocation but also significantly reduces the overall workload of ESs, thus optimizing the overall performance of the system.

[0095] In summary, by modeling the system state as the deviation between the computing time of ESs and the average computing time of the system and using a central scheduler for task scheduling, the present invention successfully reduces the state and action algorithms. The PPO deep reinforcement learning algorithm not only achieves a more balanced task allocation but also significantly reduces the overall workload of ESs, thus optimizing the overall performance of the system.

[0096] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A load balancing method for an industrial edge computing system based on deep reinforcement learning, characterized in that, The method includes the following steps: S1. At the beginning of each time slice, each edge server reports the number of tasks and computing power in the current queue to the central scheduler; S2. The central scheduler defines the system state as the deviation between the computing time of each edge server and the average computing time of the system, and generates a task scheduling action through the proximal policy optimization algorithm; S3. The central scheduler schedules tasks from edge servers with higher loads to edge servers with lower loads according to the action decision, and ensures that the scheduled task volume is an integer and satisfies the queue capacity constraint; S4. Based on the long-term discounted utility model optimization strategy, the central scheduler realizes the load balancing of the system by minimizing the variance of the system running time.

2. The method according to claim 1, characterized in that, In step S1, at the beginning of each time slice, each edge server reports the following information to the central scheduler: (1) The number of tasks Q in the current queue i and the average computational amount Z of each task; (2) The resources of its own edge server; (3) Obtain the computing time of each edge server (4) State construction: The central scheduler calculates the average computing time of the system and generates a continuous state vector: where N is the total number of edge servers.

3. The method according to claim 1, wherein In step S2, the specific content of the central scheduler generating a scheduling action based on the PPO algorithm includes the following: (1) Policy network inference: Input the state s t , and output the action probability distribution π(a|s t ); (2) Action Constraint: The action space of the central scheduler for edge server i is defined as the integer task volume a i ∈A, and satisfies (3) Queue capacity constraint: The task volume of each edge server after scheduling needs to satisfy Q i +a i ≥0 and Q i +a i ≤Q max .

4. The method according to claim 1, characterized in that In step S3, the central scheduler issues a scheduling instruction to each edge server to perform task migration, including: (1) Calculation time update: Based on the actual scheduling results, each server updates its local calculation time T i , and feeds it back to the central scheduler; (2) Reward function calculation: The system reward r is defined as: where C max is the maximum deviation allowed by the system.

5. The method according to claim 1, characterized in that In step S4, the central scheduler performs the following operations through the discounted utility model optimization strategy: Trajectory collection: Collect data of (s t , a t , r t , s t+1 ) for multiple time slices to form an experience pool; Advantage Estimation: Calculate the advantage value using Generalized Advantage Estimation (GAE) : Among them, V(s t ; θ v ) represents the value function estimate of the current state s t , and V(s t+1 ; θ v ) represents the value function estimate of the next state s t+1 . λ is the GAE decay factor; Policy update: Optimize the policy network by clipping the objective function; Calculate the value function loss: L V (θ) = (r(t) - V(s t ; θ v )) 2 ; Perform gradient descent to update the parameter θ v , and loop until convergence.