Micro-service Fault Tolerance Scheduling Method Based on Reinforcement Learning in Computing Power Network

By designing a microservice fault-tolerant scheduling method based on reinforcement learning in an edge computing environment, combining PB model and FAWS algorithm, training the SAC model offline and fine-tuning online, the problem of inconsistent device failure and data distribution in the edge computing environment is solved, and efficient fault-tolerant scheduling and environmental adaptability of tasks are achieved.

CN119324927BActive Publication Date: 2025-07-22JINAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411357646.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-07-22
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

In computing power networks, in the edge computing environment, it is difficult for the existing technology to effectively deal with instantaneous equipment failures, resulting in a downgrade of user service experience. At the same time, there are problems of cold start and data distribution inconsistent, affecting the efficiency and adaptability of task offloading.

Method used

A microservice fault-tolerant scheduling method based on reinforcement learning is designed, combining PB model and FAWS heuristic algorithm, and by training SAC models offline and fine-tuning online, the task offload strategy is optimized to solve the problem of device failure and inconsistent data distribution.

Benefits of technology

Improve the fault tolerance and adaptability of tasks in an edge computing environment, ensure that tasks are completed on time, and improve the robustness of the system and task offloading efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119324927B_ABST
    Figure CN119324927B_ABST
Patent Text Reader

Abstract

The present invention discloses a microservice fault-tolerant scheduling method based on reinforcement learning in a computing power network environment; it includes analyzing the task fault-tolerant scheduling constraints based on the PB (Primary-Backup) model for the problems of network host and link device failures in the network environment, designing a heuristic scheduling algorithm FAWS to improve the fault tolerance of the system and ensure the timely completion of tasks. To improve the performance and adaptability of the model, an offline-to-online task fault-tolerant scheduling framework based on reinforcement learning is designed. For the cold start problem during model training, the task running logs of FAWS are used to offline train a discrete soft actor-critic reinforcement learning model. For the problem of inconsistent data distribution among servers in the edge environment, when deploying the DRL model to edge servers, the GAE and PPO methods are jointly used to online fine-tune the DRL model to further improve the decision-making ability of the model and obtain the optimal scheduling decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and particularly to a fault-tolerant scheduling algorithm for a microservice workflow based on reinforcement learning in an edge computing environment with edge server and link device failures, in the face of complex task requests and heterogeneous edge server resources. Background Art

[0002] A computing power network (Computing First Network) refers to connecting computing resources scattered at different locations through a network to form a unified and collaborative computing resource pool to provide powerful computing capabilities and flexible resource scheduling. The core goal of the computing power network is to efficiently integrate and utilize distributed computing resources to meet the needs of large-scale computing tasks. In recent years, with the popularization of the Internet of Things (IoT) and the rapid development of communication technologies, new computing and communication technologies have promoted the emergence of more and more innovative mobile application programs and services, such as augmented reality, virtual reality, autonomous driving, and mobile healthcare. The computing power network has gradually pushed computing capabilities to the edge. The demand for computing and storage resources traditionally provided by cloud servers for these mobile application programs and intelligent devices has increased sharply, and the amount of data generated by them has increased unprecedentedly. At the same time, they also face the challenge of operating efficiently under strict latency and capacity constraints. The traditional centralized cloud computing architecture cannot fully solve these problems. Multi-access Edge Computing (MEC) in the computing power network is a promising computing paradigm to solve such problems. The basic principle of MEC is to extend cloud computing capabilities to network edge hosts close to users, which can significantly relieve network congestion and reduce service latency because edge computing nodes are very close to users and have sufficient computing capabilities. In addition, through cooperation between edge nodes, resource utilization can be improved and the quality of experience (QoE) can be enhanced.

[0003] In multi-access edge computing scenarios, most of the basic operation units (such as edge servers and transmission link devices) are exposed to the natural environment. Due to the relatively scattered physical locations of these nodes, it is difficult for edge computing clusters to provide a maintenance and management level comparable to that of cloud computing centers. These components are prone to transient faults caused by electromagnetic interference or cosmic radiation. The occurrence time of such faults is short, and generally does not damage the hardware. However, when transient faults or software crashes occur, it will bring a degraded service experience to users. Most of the existing work on task offloading in the edge computing paradigm usually considers minimizing the latency and energy consumption of task execution. They ignore the security issues existing in each operation unit (such as edge servers and transmission link devices) in the edge network system, especially ignoring the research on the fault perception of link devices (such as switches and repeaters). In addition, traditional heuristic, meta-heuristic, and model-based task offloading methods are mostly not suitable for edge environments where resource availability is constantly changing. When the environment changes, these mathematical models may need to be updated accordingly. Algorithms based on reinforcement learning can solve the above problems to a certain extent. However, this method of online learning at the edge has problems of cold start and inconsistent edge data distribution, and has weak adaptability to unexpected perturbations or invisible situations (i.e., new environments), such as changes in applications, the number of tasks, or data rates. Therefore, their sample efficiency is very low, and they require complete retraining to learn updated strategies for new environments, so they are very time-consuming.

[0004] In summary, there is an urgent need for a safe and effective fault tolerance mechanism to address the problem of instantaneous device faults encountered in task offloading in the computing power network scenario and to ensure the service experience of users. In addition, it is necessary to design a task scheduling framework based on reinforcement learning to adapt to the changing edge environment and solve the problems of cold start and distribution shift during the training process. Summary of the Invention

[0005] Aiming at the problems existing in the above background technology, with the goal of minimizing task calculation latency and considering the faults of network hosts and devices at the same time, the present invention designs a microservice fault perception and security scheduling method based on reinforcement learning in multi-access edge computing. Then, on this basis, a fault-tolerant reinforcement learning task offloading framework from offline to online is designed, which is applicable to edge computing systems with multiple edge servers, and solves the problems of model training cold start and edge-side data distribution shift. Specifically, it includes the following steps:

[0006] Step 1. Most large-scale compute-intensive tasks can be modeled as a directed acyclic graph (DAG). Considering the failures of edge servers and transmission links in the system, constraints on task fault-tolerant scheduling are obtained. Combining with the PB (Primary-Backup) model, an objective function for task offloading is established based on task computation latency, a fault-tolerant FAWS heuristic scheduling algorithm is designed, an offloading strategy for user tasks is obtained, and the log information of each task scheduling is recorded. The specific steps are as follows:

[0007] 1) Mobile applications based on microservices can be modeled as a directed acyclic graph DAG. Assume that a DAG arriving at an edge server s at time slot t i follows a Bernoulli distribution. Consider the DAG formed by mobile applications arriving at edge server s within time slot t as W = (R, E), where R = {n1, n2,... n i ,..., n n} represents the task set, and E = {e i,j | e i,j = (n i , n j )} represents the dependency relationships between tasks. Among them, let DP(n i ) denote the set of direct predecessor tasks of task n i , and correspondingly, DS(n i ) represents the set of direct successor tasks of task n i . Define the task set with an empty set of predecessor tasks as the entry task set n entry , and the task set with an empty set of successor tasks is the exit task set n exit . Assume the scheduling plan of tasks A = {a1, a2,..., a n}, where a i represents the offloading location of task n i . The failures in the system can be modeled as a Poisson distribution, and the PB model is adopted to solve the problems of server and host failures. The PB model realizes the fault-tolerant function by assigning a primary replica and a backup replica to each task and scheduling the primary and backup replicas to edge servers under different subnets respectively. Denote the primary replica of task n i as and the backup replica as . For ease of understanding, use to represent any replica task.

[0008] 2) Use UT i Y (t) to represent the upload latency of uploading task from the task initiating server a0 to server a i at time slot t, and calculate it according to the following formula:

[0009]

[0010] Among them, the parameter represents the data volume of task in time slot t, and TP(a0) represents the data transfer rate of edge server a0. represents the fixed transfer delay between a0 and a i . Obviously, if the task is scheduled to run on the edge server that initiates the task, then UT i Y (t) = 0. Use ET i Y (t) to represent the computing delay of task on edge server a i , which is calculated according to the following formula:

[0011]

[0012] Among them, the parameter represents the data volume of the task executed in time slot t . F(a i ) represents the processing capacity of edge server a i , with the unit of CPU cycles per second; ρ represents the number of cpu cycles required to process 1 bit. Use to represent the transfer delay caused by transmitting the operation result of task to its successor task . It is calculated according to the following formula:

[0013]

[0014] Among them, the parameter a i is the edge server where task is located, and a j is the edge server where task is located. td ij represents the data volume size transmitted between tasks. TP(a i ) represents the data transfer rate of a i . represents the fixed transfer delay between a i and a j . If task and its successor task run on the same edge server or task is an egress task, then the transfer delay is 0. Use to represent the start execution time of task in time slot t. It is calculated according to the following formula:

[0015]

[0016] Among them, the parameter AT is the arrival time of the DAG task (arrive time), and the arrival time of the first DAG task is defined as 0. Ava() is the idle time of the edge server. Similarly, use to represent the task The execution completion time within time slot t is calculated as follows:

[0017]

[0018] 3) Define as the edge server where the primary copy is located, as the edge server where the backup copy is located. is the subnet where the edge server that offloads the primary copy is located, is the subnet where the edge server that offloads the backup copy is located. Analyze the dependency relationship of the real workflow tasks arriving in time slice t. The task scheduling location constraint based on the PB model is expressed as follows:

[0019] C1:

[0020] C2:

[0021] Constraint C1 means that the backup copy i of the currently scheduled task n cannot be in the same subnet as the primary copy ; Constraint C2 means that the backup copy i of task n cannot be in the same subnet as the primary copy k |n k ∈ DP(n i )(direct predecessor task) is in the same subnet.

[0022] According to the analysis of the above steps, the primary copy i of task n can only start execution after receiving the output data of the predecessor task (that is, under normal execution). Therefore, according to formula (4), the start time of the primary copy is expressed as follows:

[0023] C3:

[0024] Similarly, it is necessary to consider the relationship between the start time of the primary copy i of task n and the backup copy of the predecessor task. In order to minimize the delay as much as possible, the constraint between the two is expressed as follows:

[0025] C4:

[0026] In the case of formula (9), the execution of has a certain time overlap with the execution of To avoid subsequent task execution failures, the start time of any backup copy

[0027] C5:

[0028] Based on the analysis in Step 1, by considering the factors of edge server and network device failures in the edge environment, with the goal of minimizing the total computing latency of DAG tasks, the time when task n i finally completes its execution is calculated as follows:

[0029]

[0030] Therefore, the multi-task scheduling optimization problem in the edge service system is established:

[0031]

[0032] s.t. C1, C2, C3, C4, C5

[0033] 4) First, implement the heuristic algorithm FAWS for DAG task scheduling according to the greedy idea. Since there are dependencies between tasks, it is necessary to convert the DAG into a sequence to determine the offloading order of each task. First, sort the tasks in ascending order of the rank value of each task. The rank value is calculated as follows:

[0034]

[0035] where rank() represents the rank value of a certain task, a i represents any edge node, and N represents the number of edge nodes. By calculating the rank value of each task in the DAG, the tasks are input into the DAG scheduling priority queue Q in descending order of the rank value. In this way, the offloading order of the tasks is obtained. Immediately afterwards, under the fault tolerance constraint, with the goal of minimizing the overall completion time of task execution, according to the order of the scheduling priority queue Q, select the edge server with the earliest completion time as the offloading target for the primary and backup copies of each task according to the greedy idea.

[0036] Step 2: Based on the task scheduling log information obtained by running the fault-tolerant scheduling algorithm, integrate the transitions data that can be used for reinforcement learning training. Offline train a discrete Soft Actor-Critic (SAC) model with the transitions data. Deploy the trained model to each edge node for online decision-making. The specific steps are as follows:

[0037] 1) First, analyze the task scheduling log information. Let the dataset sampled from the buffer be B, denoted as Update the model.

[0038] 2) Offline train a discrete SAC model with the transitions data. In the present invention, two Q networks are designed to reduce the estimation bias of the value function. Denote the parameters of the two Q networks as φ1 and φ2 respectively, and the target network parameters of φ1 and φ2 as respectively. θ is the parameter of the actor network. To achieve online fine-tuning, an additional offline critic network is trained, denoted with the parameter ω. The soft value and the target value are calculated as follows:

[0039]

[0040] where ɑ is the entropy regularization parameter used to control the importance of entropy, and γ is the discount factor. Due to the distribution shift between the heuristic method strategy for collecting data and the offline learning strategy, the CQL regularizer is used to optimize the Q network. Update the q network parameters by minimizing the following equation using gradient descent, λ c is the CQL weight. Calculate as follows:

[0041]

[0042] To facilitate subsequent online fine-tuning, an additional critic network is trained. Use gradient descent to minimize the MSE to update the critic network parameters, and at the same time use gradient descent to maximize the MSE to update the actor network parameters. Calculate as follows:

[0043]

[0044] 3) Finally, deploy the trained model to each edge server;

[0045] Step 3: The edge server regularly collects the real-time data of task offloading, combines the PPO-Clip algorithm to perform online fine-tuning on the DRL model, and finally uses the fine-tuned model to make real-time offloading decisions for tasks to obtain the optimal choice. The specific steps are as follows:

[0046] 1) During the offline training process, an additional critic network was trained. By retaining the offline-trained actor-critic network and migrating it to the online fine-tuning architecture, a dataset C was sampled batch by batch from the online data buffer pool. The Generalized Advantage Estimation (GAE) was used to balance the bias of value estimation and the variance of rewards. By designing the advantage function, the overestimation problem was solved. The update of GAE was calculated as follows:

[0047]

[0048] Similar to the offline training update process, the critic network parameters were updated by minimizing the Mean Squared Error (MSE) through gradient ascent, which was calculated as follows:

[0049]

[0050] 2) The Proximal Policy Optimization (PPO)-Clip algorithm was used to update the actor network. The PPO objective function consists of two parts: the clipped objective function and the entropy regularization term. By restricting the update speed of the network parameters, the actor network parameters were updated by maximizing the objective function through gradient ascent, which was calculated as follows:

[0051]

[0052] 3) The fine-tuned model was deployed to each edge node to make real-time decisions on the primary and backup copies of the task;

[0053] Compared with the prior art, the advantages of this method are as follows:

[0054] 1. Considering the failure problems of the basic operation units (especially network devices) in the edge environment, analyzing the characteristics of fault-tolerant dependent task scheduling based on the PB model, offloading tasks to appropriate edge servers, ensuring the timely completion of tasks, and improving the fault tolerance of the system.

[0055] 2. Designing an offline-to-online task scheduling framework based on reinforcement learning, making offloading decisions for the primary and backup copies respectively, solving the data distribution problem of model training while achieving fault-tolerant task scheduling, and improving the robustness of the model to environmental changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the flowchart of the method of the present invention;

[0057] Figure 2 is the system architecture diagram of the present invention;

[0058] Figure 3 is the flowchart of the offline training SAC agent of the present invention;

[0059] Figure 4It is the flowchart of the online fine-tuning agent of the present invention; Specific embodiments

[0060] The following combines the description Figure 1 to further elaborate on the technical solution of the present invention.

[0061] As Figure 1 shown, it is the flowchart of the microservice fault perception and security scheduling method based on reinforcement learning. This method includes the following steps:

[0062] 4) In step one, as Figure 2 shown in the system architecture diagram, there are N edge servers in the system, and the N edge servers provide computing services for all mobile users. The link devices (such as switches and routers) in the system divide the system into multiple partitions or subnets. Each switch connects one or more edge nodes to form a subnet. Each edge server in the system is jointly deployed through an access point (AP). The edge server is connected to a certain number of user devices, and user tasks access the edge system through the AP. The mobile application based on microservices can be modeled as a directed acyclic graph DAG. Let DAG be W=(R, E), where R={n1, n2,...n i ,...,n n} represents the task set, and E={e i,j |e i,j =(n i ,n j )} represents the dependency relationship between tasks. Among them, let DP(n i ) represent the set of direct predecessor tasks of task n i , and correspondingly, DS(n i ) represents the set of direct successor tasks of task n i . Define the task set with an empty set of predecessor tasks as the entry task set n entry , and the task set with an empty set of successor tasks is the exit task set n exit . Assume the scheduling plan of the task is A={a1, a2,...,a n}. Among them, a i represents the offloading location of task n i . The faults in the system can be modeled as a Poisson distribution, and the PB model is used to solve the problems of server and host faults. The PB model realizes the fault tolerance function by assigning a primary copy and a standby copy to each task and scheduling the primary and standby copies to edge servers under different subnets respectively. Denote the primary copy of task n i as and the standby copy as . For easy understanding, use to represent any copy task. Use UT i Y(t) represents that the task in time slot t is uploaded from the task initiation server a0 to the server a i . The upload delay is calculated as follows:

[0063]

[0064] where the parameter represents the data volume of the task within time slot t , and TP(a0) represents the data transmission rate of the edge server a0. represents the fixed transmission delay between a0 and a i . Obviously, if the task is scheduled to run on the edge server that initiates the task, then UT i Y (t) = 0. Let ET i Y (t) represent the computing delay of the task on the edge server a i . It is calculated as follows:

[0065]

[0066] where the parameter represents the data volume of the task executed in time slot t . F(a i ) represents the processing capacity of the edge server a i , with the unit of CPU cycles per second; ρ represents the number of cpu cycles required to process 1 bit. Let represent the transmission delay caused by transmitting the operation result of the task to its subsequent task in time slot t. It is calculated as follows:

[0067]

[0068] where the parameter a i is the edge server where the task is located, a j is the edge server where the task is located, td ij represents the data volume transmitted between tasks. TP(a i ) represents the data transmission rate of a i . represents the fixed transmission delay between a i and a j . If the task and its subsequent task run on the same edge server or the task is an exit task, then the transmission delay is 0. Let Indicates the task The start execution time within time slot t. It is calculated according to the following formula:

[0069]

[0070] Among them, the parameter AT is the arrival time of the DAG task (arrive time), and the arrival time of the first DAG task is defined as 0. Ava() is the idle time of the edge server. Similarly, use Indicates the task The completion execution time within time slot t, which is calculated according to the following formula:

[0071]

[0072] Define As the edge server where the primary replica is located, As the edge server where the backup replica is located. Is the subnet where the edge server where the primary replica is offloaded is located, Is the subnet where the edge server where the backup replica is offloaded is located. Analyze the dependency relationship of the real workflow tasks arriving in time slice t. The task scheduling location constraint based on the PB model is expressed as follows:

[0073] C1:

[0074] C2:

[0075] Constraint C1 means that the backup replica i Of the currently scheduled task n Cannot be in the same subnet as the primary replica Constraint C2 means that the backup replica i Of task n Cannot be in the same subnet as the primary replica k |n k ∈DP(n i )(Direct predecessor task) of the primary replica In the same subnet.

[0076] According to the previous analysis, design the backup replica i Of task n Can only start execution after the primary replica Finishes execution. This is because considering that the probability of frequent system failures is small, then the probability of starting the backup replica to execute the task is also small. The primary replica i Of task n It can only start to execute after receiving the output data of the previous task (that is, under normal execution). Therefore, according to formula (4), the start time of the primary replica is expressed as follows:

[0077] C3:

[0078] Similarly, it is necessary to consider task n i 's primary replica and the backup replica of the previous task in terms of the start time. In order to minimize the time delay as much as possible, the constraint between the two is expressed as follows:

[0079] C4:

[0080] In the case of formula (9), 's execution and 's execution have a certain time overlap. In order to avoid subsequent task execution failures, according to formula (4), the start time of any backup replica is expressed as follows:

[0081] C5:

[0082] Based on the analysis in step one, by considering the factors of edge server and network device failures in the edge environment, with the goal of minimizing the total computing delay of DAG tasks, therefore, the moment when task n i finally completes execution is calculated as follows:

[0083]

[0084] Therefore, the multi-task scheduling optimization problem in the edge service system is established:

[0085]

[0086] s.t. C1, C2, C3, C4, C5

[0087] First, implement the heuristic algorithm FAWS for DAG task scheduling according to the greedy idea. Since there are dependencies between tasks, it is necessary to convert the DAG into a sequence to determine the offloading order of each task. First, sort the tasks in ascending order according to the rank value of each task. The rank value is calculated as follows:

[0088]

[0089] where rank() represents the rank value of a certain task, a iDenote any edge node as \(v\), and \(N\) as the number of edge nodes. By calculating the rank value of each task in the DAG, the tasks are input into the DAG scheduling priority queue \(Q\) in descending order of the rank value. In this way, the offloading order of the tasks is obtained. Immediately following that, under the fault tolerance constraint, with the goal of minimizing the overall completion time of task execution, according to the order of the scheduling priority queue \(Q\), and following the greedy idea, the edge node with the earliest completion time is selected as the offloading target for the primary and backup replicas of each task respectively. If the task meets the fault tolerance constraint, it can proceed to the next step; otherwise, a new offloading target is selected for the task. In addition to meeting the fault tolerance avoidance principle, the task also needs to satisfy the resource constraint of the edge node. Finally, by capturing the environmental information before and after task scheduling, combined with the offloading decision of the DAG task, elements such as state, action, next_state, and reward required for offline training of reinforcement learning are constructed. And the standardized data is configured into a transition and stored in the log Buffer;

[0090] In step two, data is sampled in batches from the buffer Buffer, and these data are used to train a discrete soft actor-critic reinforcement learning model, which can effectively solve the problem of inconsistent data distribution in training the DRL model at the edge, and improve the stability and sample efficiency of the model. The algorithm process of offline training the discrete SAC with heuristic data is as Figure 3 shown. First, data is taken out in batches from the Buffer, and the relevant data is standardized and normalized. Let the dataset sampled from the buffer be \(B\), denoted as \(1\leq i\leq|B|\), and the model is updated. Then two Q networks are used to avoid value overestimation. First, the value of soft value is calculated, then the target Q value is calculated, and finally the Q network is updated. The softvalue and targetvalue are calculated as follows:

[0091]

[0092] where \(\alpha\) is the entropy regularization parameter, used to control the importance of entropy, and \(\gamma\) is the discount factor. Due to the distribution shift between the heuristic method strategy for collecting data and the offline learning strategy, the CQL regularizer is used to optimize the Q network. CQL ensures its conservative estimate by regularizing the Q value function to avoid overestimation during the value function learning process, thereby reducing instability and sub-optimality of the strategy. The q network parameters are updated by minimizing the following equation using gradient descent, \(\lambda\) c is the CQL weight. It is calculated as follows:

[0093]

[0094] To facilitate subsequent online fine-tuning, an additional critic network was trained. The critic network parameters were updated by minimizing the MSE using gradient descent, while the actor network parameters were updated by maximizing the MSE using gradient descent. It is calculated as follows:

[0095]

[0096] Finally, the trained model was deployed to each edge server;

[0097] In step three, each edge server accumulates a large amount of task running data during execution. These real-time data need to be collected regularly, and the DRL model is fine-tuned to update the scheduling model so that the model can better adapt to the local data distribution characteristics. The algorithm flow for online training of the DRL model is as Figure 4 shown. During offline training, an additional critic network was trained. By retaining the offline-trained actor-critic network and migrating it to the online fine-tuning architecture; then, a dataset C was sampled batch by batch from the online data buffer pool. Since there is a distribution shift problem between offline data and online data during the process of online fine-tuning the model, and the Q network and the critic network are pre-trained based on offline data, this will lead to an overestimation of the observed results not seen in the online stage. GAE is used to balance the bias of value estimation and the variance of rewards. By designing the advantage function, the overestimation problem is solved. The update of GAE is calculated as follows:

[0098]

[0099] GAE balances the variance and bias by smoothing the advantage estimates of different step sizes, so that the advantage function can be better estimated and the critic becomes more robust. Similar to the offline training update process, the critic network parameters are updated by minimizing the MSE using gradient ascent. It is calculated as follows:

[0100]

[0101] The PPO-Clip algorithm is used to update the actor network. The PPO objective function consists of two parts: the clipped objective function and the entropy regularization term. By restricting the magnitude of the policy update, it prevents the policy from changing too much during each update, ensuring the stability of the algorithm and helping to converge to a better policy. By restricting the update speed of the network parameters, the actor network parameters are updated by maximizing the objective function using gradient ascent. It is calculated as follows:

[0102]

[0103] The fine-tuned model is sent to each edge node to make real-time decisions on the primary and backup copies of the task. The process of real-time decision-making is as follows: If the scheduled task is the primary copy, the actor network returns (action, prob, val) based on the currently input state, and obtains the offloading plan for task scheduling. If the scheduled task is the backup copy, the actor network first excludes actions that do not meet the constraints according to the fault tolerance requirements based on the currently input state, and then selects the optimal action as the offloading plan. The data returned by the scheduled task is combined with the environmental information to form a new transition, which is then input into the buffer for experience replay sampling.

Claims

1. A microservice fault-tolerant scheduling method based on reinforcement learning in a computing power network environment, characterized in that: First, analyze the fault-tolerant scheduling constraints based on task dependencies, design a heuristic scheduling algorithm FAWS with the goal of minimizing the total task execution latency, and then design an offline-to-online task scheduling framework based on reinforcement learning to make offloading decisions for primary and backup replicas respectively. The specific steps are as follows: Step 1: Model large compute-intensive tasks as a directed acyclic graph (DAG). Considering the failures of edge servers with heavy traffic and transmission links, obtain the constraints on task fault-tolerant scheduling. Combine with the PB Primary-Backup model, establish the objective function for computing task offloading according to the edge computing latency, design a fault-tolerant FAWS heuristic scheduling algorithm, obtain the offloading strategy for user tasks, and record the log information of each task scheduling: 1) In step 1, define the DAG as W = (R, E), where R = {n1, n2,... n i ,..., n n} represents the task set, and E = {e i,j | e i,j = (n i , n j )} represents the dependency relationship between tasks. Among them, let DP(n i ) represent the set of direct predecessor tasks of task n i , and correspondingly, DS(n i ) represents the set of direct successor tasks of task n i . Define the task set with an empty set of predecessor tasks as the entry task set n entry , and the task set with an empty set of successor tasks is the exit task set n exit . Assume the scheduling plan of tasks A = {a1, a2,..., a n}, where a i represents the offloading location of task n i . The faults in the system can be modeled as a Poisson distribution. The PB model is used to solve the problems of server and host faults. The PB model assigns a primary replica and a backup replica to each task, and schedules the primary and backup replicas to edge servers under different subnets respectively to achieve the fault tolerance function. Denote the primary replica of task n i as and the backup replica as . For ease of understanding, use to represent any replica task, and use to represent the upload delay of uploading task from the task initiating server a0 to server a i at time slot t, which is calculated according to the following formula: Among them, the parameter represents the data volume of task in time slot t. TP(a0) represents the data transmission rate of edge server a0. represents the fixed transmission delay between a0 and a i . Obviously, if the task is scheduled to run on the edge server that initiates the task, then is used to represent the computing delay of task on edge server a i , which is calculated according to the following formula: Among them, the parameter represents the amount of task data executed in the t time slot , and F(a i ) represents the processing capacity of edge server a i , with the unit of CPU cycles per second; ρ represents the number of cpu cycles required to process 1 bit, and represents the data transmission of the task operation result to its subsequent task in the t time slot, which brings the transmission delay and is calculated as follows: Among them, parameter a i is the edge server where the task is located, and a j is the edge server where the task is located. td ij represents the amount of data transferred between tasks. TP(a i ) represents the data transfer rate of a i . represents the fixed transmission delay between a i and a j . If task and its successor task run on the same edge server or task is an egress task, the transmission delay is 0. Let represent the start execution time of task in time slot t, which is calculated according to the following formula: Among them, the parameter AT is the arrival time of the DAG task (arrive time), and the arrival time of the first DAG task is defined as 0. Ava() is the idle time of the edge server. Similarly, represents the task The moment when the execution is completed within the time slot t is calculated according to the following formula: Step 2: Based on the task scheduling log information obtained by running the fault-tolerant scheduling algorithm, integrate the transitions data that can be used for reinforcement learning training. Offline train a discrete soft actor-critic (SAC) model with the transitions data, and deploy the trained model to each edge node for online decision-making; Step 3: Edge servers regularly collect the real-time data of task offloading, fine-tune the DRL model using the PPO-Clip algorithm, and finally use the fine-tuned model to make real-time offloading decisions for tasks.

2. The microservice fault tolerance scheduling method based on reinforcement learning according to claim 1, wherein: Definition is the edge server where the primary copy is located, is the edge server where the backup copy is located, is the subnet where the edge server for offloading the primary copy is located, is the subnet where the edge server for offloading the backup copy is located. Analyze the dependency relationship of the real workflow tasks arriving at time slice t. The task scheduling location constraint based on the PB model is expressed as follows: Constraint C1 states that the backup copy i of the currently scheduled task n cannot be in the same subnet as the primary copy ; Constraint C2 states that the backup copy i of task n cannot be in the same subnet as the primary copy of any k |n k ∈ DP(n i ) that is a direct predecessor task .

3. The microservice fault tolerance scheduling method based on reinforcement learning according to claim 1, characterized in that: In step 1, based on the analysis of task dependencies, design task n i Backup copy of Only in the primary copy After the execution is completed, the execution can start. i Master copy Execution can only start after receiving the output data of the previous task. Therefore, according to formula (4), the startup time of the master copy is expressed as follows: Similarly, it is necessary to consider task n i 's main copy and the backup copy of the previous task The relationship in terms of startup time. To minimize latency as much as possible, the constraint between the two is expressed by the following formula: In the case of formula (9), The execution of has a certain time overlap with the execution of. To avoid subsequent task execution failures, any backup copy is set according to formula (4) The start time is expressed by the following formula:

Citation Information

Patent Citations

  • Edge computing adaptive task arranging and scheduling method based on meta-reinforcement learning

    CN118093141A

  • KR20240072551A