A workflow task scheduling method
By building a multi-workflow task scheduling system on non-dedicated edge servers, using deep reinforcement learning algorithms and Markov decision-making process model, the problem of high service-level protocol violation rate in multi-workflow scheduling is solved, and the service quality is improved.
Patent Information
- Application Number
- CN202111402858.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-24
AI Technical Summary
When scheduling multiple workflows on non-dedicated edge servers, it is difficult for the prior art to effectively minimize the violation rate of workflow service-level protocols, resulting in a decline in service quality.
Using an algorithm based on deep reinforcement learning, by constructing an optimization model for multi-workflow task scheduling problems and a Markov decision-making process model, using the PRDDQN algorithm to solve the scheduling location of the task, establish a workflow task scheduling system, including a workflow information collection module and a task scheduling module, and work together to minimize the violation rate of service-level protocols.
It effectively improves the service quality of non-dedicated edge servers, and improves the efficiency and accuracy of task scheduling by minimizing the violation rate of workflow service-level protocols.
Smart Images

Figure CN114035927B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile edge computing, and in particular, to a workflow task scheduling method. Background Art
[0002] Due to the rapid development of scientific computing, scientific workflow applications have become big data applications that require large-scale infrastructure to execute within a reasonable time. However, the inherent resource limitations of mobile devices cannot meet their normal operation. Therefore, how to schedule workflows on heterogeneous resources has become an urgent problem to be solved. Previous studies usually schedule the scientific workflows generated by applications to cloud computing platforms with powerful computing resources. However, the cloud computing platform is too far from users, which will greatly increase communication costs and energy consumption, which is obviously fatal for some applications requiring low latency and low energy consumption. Therefore, mobile edge computing, as a complementary computing platform between mobile devices and remote clouds, has been widely used to solve this problem. In this distributed architecture, large-scale services originally processed by the central node are cut into smaller parts and assigned to edge nodes closer to users for processing, which greatly reduces latency and energy consumption.
[0003] Driven by the significant progress of mobile communication technology, more and more users start to execute workflow applications in the mobile edge computing environment. In the future, most mobile edge servers will become non-dedicated edge servers that can provide computing services for multiple users at the same time. The competition between resources will cause some workflows offloaded to the edge server to not be completed on time, which will lead to violations of the workflow service level agreement, thereby reducing the service quality of the edge server. Therefore, scheduling multiple parallel workflows on non-dedicated edge servers is a huge challenge.
[0004] Existing research usually focuses on offloading single or multiple workflows on edge servers, while there are few literatures on scheduling multiple workflows on non-dedicated edge servers. Therefore, we propose an algorithm based on deep reinforcement learning to schedule multiple workflows on non-dedicated edge servers, aiming to minimize the violation rate of the workflow service level agreement, thereby improving the service quality of non-dedicated edge servers. Summary of the Invention
[0005] Aiming at the problems existing in the prior art, the present invention provides a workflow task scheduling method for a non-dedicated edge server environment, which can minimize the violation rate of the workflow service level agreement by minimizing the violation rate of the workflow service level agreement.
[0006] The workflow task scheduling method provided by the present invention includes the following steps:
[0007] S1: Offload multiple workflows to a non-dedicated edge server;
[0008] S2: The workflow scheduler in the edge server identifies tasks that can be executed immediately and adds them to the task queue to be scheduled;
[0009] S3: Build an optimization model for the multi-workflow task scheduling problem and a Markov decision process model;
[0010] S4: Solve the scheduling positions of the workflow tasks in the task queue to be scheduled, and add the workflow tasks to be scheduled to the corresponding CPU waiting queue;
[0011] S5: Detect whether all workflow tasks have been scheduled. If so, go to step S1; otherwise, go to step S2.
[0012] Furthermore, in step S3, the optimization model of the multi-workflow task scheduling problem includes minimizing the violation rate of the workflow service level agreement under the constraint function;
[0013] The constraint function can ensure that each task can only appear once in the task waiting queue, ensure that each position in the waiting queue on the CPU can only be occupied by one task at a time, ensure that each task can only start execution after all its predecessor tasks are completed, and ensure that each task is completed before its successor task starts execution.
[0014] Furthermore, in step S3, a Markov decision process model is established for the multi-workflow task scheduling problem. The elements of the Markov decision process model include the state space, the action space, and the reward function.
[0015] Furthermore, the state space includes:
[0016] S i,j,CPU ={LD i,j ,Type i,j ,PP1,ST i,j,1 ,…,PP l ,ST i,j,l}i=0,1,…,mj=0,1,…,n
[0017] Where LD i,j is the task instruction length of task t i,j , Type i,j is the task type of task t i,j , PP l represents the processing performance of CPU l , and ST i,j,l represents the earliest start execution time of task t i,j on CPU l .
[0018] Furthermore, the action space includes:
[0019] A = {<i, j, k>} where i = 0, 1, …, m; j = 0, 1, …, n; k = 0, 1, …, l
[0020] Denoted as task t i,j and the receiving task t i,j of the processor CPU k of the triple
[0021] Furthermore, the steps of obtaining the reward function include:
[0022] Schedule all tasks to the CPU with the worst performance, and obtain the number V of workflows that violate the service-level agreement w ;
[0023] Schedule all tasks to the CPU with the best performance, and obtain the number V of workflows that violate the service-level agreement b ;
[0024] Through V w and V b Obtain the final reward, and the final reward r includes:
[0025]
[0026] Furthermore, in step S4, the PRDDQN algorithm is used to solve the scheduling position of the workflow tasks, and the steps of the PRDDQN algorithm include:
[0027] S41: Randomly initialize all parameters ω of the current Q-network, and initialize the parameters ω' of the target Q-network Q' = ω; Initialize the default data structure of the experience replay pool Sumtree, and the priority P of all leaf nodes of all Sumtrees j is 1; Initialize the agent that interacts with the environment and the environment;
[0028] S42: The agent observes the state S of the current environment t , and obtains its feature vector
[0029] S43: Use as the input in the current Q-network, and obtain the Q-values corresponding to all actions of the Q-network; Then use the ε-greedy method to select the action a corresponding to the maximum Q-value from the current Q-values t ;
[0030] S44: The agent executes the action a t , obtains the reward r t , and obtains whether the final state final is reached t ;
[0031] S45: Use the experience replay method with the maximum priority P t = max i<t P i Store the six - tuple ( a t , r t , γ t , final t , ) into the Sumtree;
[0032] S46: The agent continues to observe the next state;
[0033] S47: Sample from the Sumtree to train the Q - network;
[0034] S48: If the tasks of all workflows are successfully scheduled, go to step S41; otherwise, go to step S43 until the preset number of iteration rounds is reached.
[0035] Furthermore, in step S41, the experience replay pool Sumtree is a tree - shaped structure. Each leaf node stores the priority P of each sample. Each non - leaf node has only two branches, and the value of each non - leaf node is the sum of the two branches.
[0036] Furthermore, in step S43, the specific steps for the agent to select an action are: The agent selects a random action a with a probability of ε t , otherwise selects the action with a probability of 1 - ε
[0037] Furthermore, in step S47, the steps of sampling from the Sumtree include:
[0038] S471: Based on the probability Take out m samples ( a j , r j , γ j , final j , ) from the Sumtree;
[0039] S472: Calculate the loss function weight: ω j =(N × p(j)) -β / max i ω i ;
[0040] S473: Calculate the current target Q - value y j :
[0041]
[0042] S474: Update all the parameters ω of the Q-network through backpropagation of gradients of the neural network, where the loss function is
[0043] S475: Recalculate the TD errors of all samples: Then update the priorities P of all nodes in the Sumtree j = |δ j |;
[0044] S476: Update the target Q-network parameters ω' = ω every certain number of steps.
[0045] By providing a multi-workflow task scheduling method for a non-dedicated edge server environment, the present invention first establishes a Markov decision process model for the multi-workflow scheduling problem, then uses the PRDDQN algorithm to solve the scheduling positions of tasks, and finally schedules the tasks to the waiting queues of the corresponding servers, thereby minimizing the violation rate of the service level agreement of the workflow, and can effectively improve the service quality of the non-dedicated edge server. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The present invention will be described in more detail below based on embodiments with reference to the drawings. Among them:
[0047] Figure 1 is a schematic flowchart of the workflow task scheduling method for a non-dedicated edge server environment in the present invention;
[0048] Figure 2 is a system architecture diagram of the workflow task scheduling method for a non-dedicated edge server environment in the present invention;
[0049] Figure 3 is a storage structure diagram of the Sumtree of the experience replay pool of the experience replay method used in the PRDDQN algorithm in the present invention;
[0050] Figure 4 is an instance deployment diagram of the workflow task scheduling method for a non-dedicated edge server environment in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] To clearly illustrate the inventive concept of the present invention, the present invention will be described below with reference to embodiments.
[0052] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "horizontal", "top", "bottom", etc. are all based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present invention. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0053] As Figure 1 shown, a workflow task scheduling method for a non-dedicated edge server environment provided by the present invention specifically includes the following steps:
[0054] S1: Unload multiple workflows to a non-dedicated edge server;
[0055] S2: The workflow scheduler in the edge server identifies tasks that can be immediately executed and adds them to the task queue to be scheduled;
[0056] S3: Construct an optimization model for the multi-workflow task scheduling problem and a Markov decision process model;
[0057] S4: Solve the scheduling positions of the workflow tasks in the task queue to be scheduled, and add the workflow tasks to be scheduled to the corresponding CPU waiting queue;
[0058] S5: Detect whether all workflow tasks have been scheduled. If so, go to step S1; otherwise, go to step S2.
[0059] In step S3, the optimization model of the multi-workflow task scheduling problem includes minimizing the violation rate of the workflow service level agreement under the constraint function;
[0060] The optimization model established for the scheduling problem of multi-workflow tasks includes:
[0061]
[0062] The constraint function specifically includes:
[0063]
[0064]
[0065]
[0066]
[0067] x i,j,s,k ∈{0,1} (6)
[0068] V i ∈ {0, 1} (7)
[0069] Among them, the objective function (1) is to minimize the violation rate of the workflow service - level agreement. WF and SIZE(WF) represent the set of workflows and the number of workflows respectively; V i indicates whether the service - level agreement of workflow i is violated.
[0070] The constraint function (7) defines the value range of V i . The value of V i can be determined by the following method: when the total execution time of workflow i exceeds the user - specified deadline Deadline i , the value of V i is 1; otherwise, the value of V i is 0.
[0071] The constraint function (2) ensures that each task can only appear in the task waiting queue once. Among them, CPU and SIZE(CPU) represent the set of processor resources and the number of CPUs respectively, and TQ and SIZE(TQ) represent the set of task waiting queues and the length of the task waiting queue; x i,j,s,k is a decision variable, which indicates whether task j in workflow i is scheduled to the s - th position of the waiting queue of processor k.
[0072] The constraint function (6) defines the value range of x i,j,s,k : when the scheduling is successful, x i,j,s,k = 1; otherwise, x i,j,s,k = 0. T i represents the set of tasks in workflow i.
[0073] The constraint function (3) ensures that each position in the waiting queue on the CPU can only be occupied by one task at a time. SIZE(T i ) represents the number of tasks in workflow i.
[0074] The constraint function (4) ensures that each task can only start execution after all its predecessor tasks are completed. ST i,j,k represents the start execution time of task j in workflow i on processor k. pr(t i,j ) represents the set of predecessor tasks of task t i,j .
[0075] The constraint function (5) ensures that each task is completed before its successor tasks start execution. su(t i,j ) represents the set of successor tasks of task t i,j .
[0076] In the process of establishing the Markov decision process model in step S3, due to the Markov property between workflows, the multi-workflow scheduling problem can be modeled as a Markov decision process. The elements of the Markov decision process model include the state space, action space, and reward function, etc., which are defined as follows:
[0077] S31, state space: In the multi-workflow scheduling problem, the violation rate of the optimized target workflow can be regarded as an agent that can interact with the environment. When the agent interacts with the environment, the environment actually includes the state of tasks and the state of the CPU. More specifically, the state of task t i,j is represented by the task instruction length LD i,j and the task type Type i,j ; the state of the CPU k includes its processing performance and the waiting time of the new task scheduled to the CPU k . When a new task is assigned to the CPU k for execution, the waiting time is actually the start time of this task. Therefore, when task t i,j is scheduled to the processor set CPU, the current state perceived by the agent can be expressed as:
[0078] S i,j,CPU ={LD i,j , Type i,j , PP1, ST i,j,1 ,…, PP l , ST i,j,l} i = 0, 1, …, m j = 0, 1, …, n
[0079] where PP l represents the processing performance of the CPU l , and ST i,j,l represents the earliest start execution time of task t i,j on the CPU l .
[0080] S32, action space: An action refers to the behavior of allocating a task to a processor. Therefore, an action can be expressed as a triple of task t i,j and the processor CPU i,j receiving task t k :
[0081] A = {<i, j, k>} i = 0, 1, …, m j = 0, 1, …, n k = 0, 1, …, l
[0082] S33, Reward Function: After taking an action, the agent needs to obtain the reward feedback from the current state according to the reward function. Since the optimization goal is to minimize the violation rate of the service level agreement of the workflow, it is necessary to obtain the completion time MS of each workflow i to calculate whether it exceeds the deadline i . However, the MS of the workflow cannot be obtained until the workflow is completed i . Therefore, before all tasks in each workflow are completed, the value of the reward function is set to 0. When the last task in the workflow is completed, the agent is given an appropriate reward. To simplify the calculation, the present invention uses the Min-Max normalization method to reduce the value of V i to an appropriate range. First, all tasks are scheduled to the CPU with the worst performance to obtain the number V of workflows violating the service level agreement w , and then all tasks are scheduled to the CPU with the best performance to obtain the number V of workflows violating the service level agreement b . Finally, the final reward is obtained through V w and V b :
[0083]
[0084] In step S4, the PRDDQN algorithm is used to solve the scheduling position of the workflow tasks in the task queue to be scheduled. The steps of the PRDDQN algorithm include:
[0085] S41: Randomly initialize all parameters ω of the current Q-network, and initialize the parameters ω′ = ω of the target Q-network Q′; Initialize the default data structure of the experience replay pool Sumtree, and the priority P of all leaf nodes of all Sumtrees j is 1; Initialize the agent interacting with the environment and the environment;
[0086] S42: The agent observes the state S of the current environment t , and obtains its feature vector
[0087] S43: Use as the input in the current Q-network to obtain the Q-values corresponding to all actions of the Q-network; Then use the ε-greedy method to select the action a corresponding to the maximum Q-value from the current Q-values t ;
[0088] S44: The agent executes the action a t , obtains the reward r t , and obtains whether the final state final is reached t ;
[0089] S45: Use the experience replay method with the maximum priority P t = max i<t P i to store the six-tuple ( a t , r t , γ t , final t , ) into the Sumtree;
[0090] S46: The agent continues to observe the next state;
[0091] S47: Sample from the Sumtree every certain number of steps to train the Q-network;
[0092] S48: If the tasks of all workflows are successfully scheduled, go to step S41; otherwise, go to step S43 until the preset number of iteration rounds is reached.
[0093] When initializing the experience replay, the specific storage structure of the experience replay pool Sumtree is described as follows:
[0094] Refer to Figure 3 . Overall, the Sumtree is a tree structure. Each leaf node stores the priority P of each sample, and each non-leaf node has only two branches. The value of each non-leaf node is the sum of the two branches, so the top node of the Sumtree is the sum of all S leaf nodes. When sampling, divide the value of the top node by the number of samples to be sampled, divide it into several intervals, and then randomly select a number in each interval. Finally, search the Sumtree from top to bottom according to this number, and compare the priority P of the two branch nodes with this number. If it is greater than this number, select the left branch node and continue to go down. Otherwise, select the right branch node, subtract the value of the right branch node from this number, and then continue to go down. Until reaching the leaf node, select the data corresponding to this leaf node as the sample for this sampling.
[0095] In step S43, the specific steps for the agent to select an action are: The agent selects a random action a with a probability of ε t , otherwise selects an action with a probability of 1 - ε
[0096] In step S47, the sampling process is as follows:
[0097] S471: Based on the probability take out m samples from the Sumtree ( a j , r j , γj , final j , )。
[0098] S472: Calculate the loss function weight: ω j = (N × p(j)) -β / max i ω i 。
[0099] S473: Calculate the current target Q value y j :
[0100]
[0101] S474: Update all parameters ω of the Q network through the gradient backpropagation of the neural network, where the loss function is
[0102] S475: Recalculate the TD error (degree of priority learning) of all samples: Then update the priority P of all nodes in the Sumtree j = |δ j |。
[0103] S476: Update the target Q network parameter ω' = ω every certain number of steps.
[0104] The workflow task scheduling method in the present invention can effectively improve the service quality of non-dedicated edge servers by minimizing the violation rate of the service level agreement of the workflow.
[0105] The present invention also provides a workflow task scheduling system. Combining Figure 2 with the system architecture diagram of the workflow task scheduling method for non-dedicated edge server environments, the workflow task scheduling system includes a base station of non-dedicated edge servers and multiple users, where the non-dedicated edge server contains L heterogeneous CPU processors.
[0106] See Figure 2 , there are many users in a region. Each user can generate different workflows using different devices at different times. There is a base station with a non-dedicated edge server in the region. The edge server contains L heterogeneous CPU processors. Users can submit workflows to the server through the cellular mobile network or the WIFI network.
[0107] Combining Figure 4 , the implementation and deployment of the method involved in the present invention rely on core components such as a workflow information collection module and a task scheduling module. During the execution of the method, the collaborative work of each component can be divided into a preparation stage, a running stage, and a scheduling stage. The work in each stage is as follows:
[0108] Preparation stage:
[0109] 1: Add a multi-workflow task scheduling decision maker to the non-dedicated edge server;
[0110] 2: Deploy the workflow information collection module and the task scheduling module to the multi-workflow task scheduling decision maker;
[0111] 3: Use the stress test program to pre-measure the task request processing saturation value S of each CPU. The multi-workflow scheduling decision maker creates an idle resource vector for each CPU to form a task waiting queue, and initializes the element value of the idle resource vector of each CPU to S;
[0112] 4: The workflow information collection module creates a workflow task buffer queue.
[0113] Running stage:
[0114] When the user submits a workflow instance, the collaborative work among the core components of the multi-workflow scheduling decision maker includes:
[0115] 1: The workflow information collection module inserts the workflow submitted by the user into the buffer queue;
[0116] 2: The workflow information collection module identifies the tasks that can be executed immediately and sends them to the task scheduling module;
[0117] 3: The task scheduling module constructs a multi-workflow scheduling optimization model based on the task instance information;
[0118] 4: The task scheduling module uses the PRDDQN algorithm to generate a scheduling result, so as to obtain the scheduling position of each task;
[0119] 5: The task scheduling module inserts the tasks into the task waiting queue according to the scheduling result.
[0120] Scheduling stage:
[0121] 1: The CPU schedules tasks according to the order of the task waiting queue;
[0122] 2: When the task execution ends, delete the first element of the task waiting queue of the corresponding CPU, and insert the new task at the end of the task waiting queue of the corresponding CPU;
[0123] 3: The workflow information collection module regularly collects the execution logs and metadata of the executed task instances from the databases of each CPU. The metadata of the task instance includes: task instruction length, task execution time, and the user who submitted the task;
[0124] 4: Return the execution log and metadata to the non-dedicated edge server, and one monitoring cycle ends.
[0125] The workflow task scheduling method for the non-dedicated edge server environment in the present invention specifically adopts the following steps:
[0126] Step 1: In one embodiment, the edge server for multi-workflow scheduling can be a base station server or a cell computer room server, etc. As Figure 2 shown, in the scenario of the present invention, the edge server has l heterogeneous CPU processors, represented by CPUS = {cpu1, cpu2,..., cpu l}, and the processing performance of each processor is pp k . Each user in this area can generate different scientific workflows using different devices at different times. Therefore, the present invention considers scheduling the scientific workflows generated by users from a discretization perspective, dividing the service time T of the edge server into discrete time intervals t, t ∈ {1, 2,..., n}, where n represents the number of time intervals.
[0127] Step 2: Regard the edge server as a multi-workflow scheduling decision maker, and deploy two components on it, namely the workflow information collection module and the task scheduling module. At each time interval t, the workflow information collection module collects the scientific workflows submitted by multiple users in real time and sorts them according to the earliest deadline first (EDF) principle. The number of task nodes included in the workflow is set to K, and a topological arrangement is performed on 1 - K, and the corresponding directed acyclic graph G = <V, E> is generated with this arrangement order as the topological sorting result of the directed acyclic graph. The set of nodes V = {v1, v2,..., v k} in the graph is used as the set of task nodes in the workflow, and the set of directed edges E = {e ij |v i ∈V, v j ∈V} in the graph is used as the set of dependency relationships between task nodes in the workflow. The task information extracted by the workflow information collection module includes task size, task type, deadline, etc.
[0128] Step 3: The workflow information collection module forwards the task information that can be executed immediately to the task scheduling module. The task scheduling module uses the collected task information to establish a multi-workflow scheduling model, takes the violation rate of the service level agreement of the workflow as the objective function, uses the PRDDQN algorithm to solve the objective function, allocates CPU resources to the task nodes according to the calculated optimal solution, and adds them to the CPU task waiting queue of the corresponding one. Among them, when sampling samples from the experience replay pool, the specific description of the sampling process of the PRDDQN algorithm is as follows: See Figure 3, Overall, the Sumtree is a tree structure. Each leaf node stores the priority P of each sample, and each non-leaf node has only two branches. The value of each non-leaf node is the sum of the two branches, so the top node of the Sumtree is the sum of all S leaf nodes. When sampling, divide the value of the top node by the number of samples to be sampled, divide it into several intervals, and then randomly select a number in each interval. Finally, search the Sumtree from top to bottom according to this number, and compare the priority P of the two branch nodes with this number. If it is greater than this number, select the left branch node and continue to go down. Otherwise, select the right branch node, subtract the value of the right branch node from this number, and then continue to go down. Until reaching the leaf node, select the data corresponding to the leaf node as the sample for this sampling. Through the above process, the speed of the neural network training model can be accelerated, effectively saving time costs.
[0129] Step 4: When all tasks of a workflow are executed, the task scheduling module returns the task execution log data to the workflow information collection module. When a new workflow arrives at the edge server, calculate the task waiting queue to which the new task belongs and its position in the task waiting queue, and insert it into the task waiting queue.
[0130] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0131] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A workflow task scheduling method, characterized in that, It includes the following steps: S1: Unload multiple workflows to non-dedicated edge servers; S2: The workflow scheduler in the edge server identifies tasks that can be immediately executed and adds them to the task queue to be scheduled; S3: Construct an optimization model for the multi-workflow task scheduling problem and a Markov decision process model; S4: Solve the scheduling positions of the workflow tasks in the task queue to be scheduled, and add the workflow tasks to be scheduled to the corresponding CPU waiting queue; S5: Detect whether all workflow tasks have been scheduled. If so, go to step S1; Otherwise, go to step S2; In step S3, the optimization model of the multi-workflow task scheduling problem includes minimizing the violation rate of the workflow service level agreement under the constraint function; The constraint function can ensure that each task can only appear once in the task waiting queue, ensure that each position in the waiting queue on the CPU can only be occupied by one task at a time, ensure that each task can only start execution after all its predecessor tasks are completed, and ensure that each task is completed before its successor task starts execution; In step S3, establish a Markov decision process model for the multi-workflow task scheduling problem. The elements of the Markov decision process model include the state space, the action space, and the reward function; The state space includes: S i,j,CPU = {LD i,j , Type i,j , PP1, ST i,j,1 , …, PP l , ST i,j,l} i = 0, 1, …, m j = 0, 1, …, n Among them, LD i,j is the task instruction length of task t i,j , Type i,j is the task type of task t i,j , PP l represents the processing performance of the CPU l , ST i,j,l represents the earliest start execution time of task t i,j on the CPU l ; The action space includes: A = {<i,j,k>} i = 0,1,…,m j = 0,1,…,n k = 0,1,…,l Denoted as task t i,j and the receiving task t i,j of the processor CPU k triple; The steps to obtain the reward function include: Schedule all tasks to the CPU with the worst performance, and obtain the number V of workflows that violate the service-level agreement w ; Schedule all tasks to the CPU with the best performance, and obtain the number V of workflows that violate the service level agreement b ; Through V w and V b Obtain the final reward, and the final reward r includes: In step S4, use the PRDDQN algorithm to solve the scheduling position of the workflow task. The steps of the PRDDQN algorithm include: S41: Randomly initialize all parameters ω of the current Q-network, and initialize the parameters ω of the target Q-network Q'; initialize the default data structure of the experience replay pool Sumtree, and set the priority P of all leaf nodes of all Sumtrees to 1; initialize the agent that interacts with the environment and the environment. ′ = ω; Initialize the default data structure of the experience replay pool Sumtree, and the priority P of all leaf nodes of all Sumtrees j is 1; Initialize the agent that interacts with the environment and the environment; S42: The agent observes the state S of the current environment t , and obtains its feature vector S43: Use in the current Q-network as the input to obtain the Q-values corresponding to all actions of the Q-network; then use the ε-greedy method to select the action a corresponding to the maximum Q-value from the current Q-values t ; S44: The agent executes action a t , and obtains a reward r t , and obtains whether the final state final is reached t ; S45: Use the experience replay method with the highest priority to store the six-tuple into the Sumtree; S46: The agent continues to observe the next state; S47: Sample from the Sumtree to train the Q network; S48: If all tasks of all workflows are scheduled successfully, go to step S41. Otherwise, go to step S43 until the preset number of iteration rounds is reached.
2. The workflow task scheduling method according to claim 1, wherein, In step S41, the experience replay pool Sumtree is a tree structure. Each leaf node stores the priority P of each sample. Each non-leaf node has only two branches, and the value of each non-leaf node is the sum of the two branches.
3. The workflow task scheduling method according to claim 1, wherein In step S43, the specific steps for the agent to select an action are as follows: the agent selects a random action a with a probability of ε t , otherwise, with a probability of 1 - ε, it selects an action 4. The workflow task scheduling method according to claim 1, wherein In step S47, the steps of sampling from the Sumtree include: S471: Based on probability Take out m samples from the Sumtree S472: Calculate the loss function weight: S473: Calculate the current target Q value y j : S474: Update all parameters ω of the Q-network through backpropagation of gradients of the neural network, where the loss function is S475: Recalculate the TD error for all samples: Then update the priorities P of all nodes in the Sumtree j = |δ j |; S476: Update the target Q-network parameter ω every certain number of steps ′ = ω.
Citation Information
Patent Citations
Workflow scheduling method based on deep Q neural network in edge computing environment
CN112905312A
Optimized service deployment method in mobile edge computing
CN113296909A