A Speculative Execution Scheduling Method for Heterogeneous MapReduce Clusters Based on Reinforcement Learning

By adopting a speculative execution scheduling method based on reinforcement learning in Hadoop MapReduce, dynamically update node weights and distinguish between fast/slow nodes of task types, the problems of inaccurate task estimation and low resource scheduling efficiency in heterogeneous clusters are solved, and more efficient data processing is achieved.

CN113867944BActive Publication Date: 2025-07-01BEIJING INST OF COMP TECH & APPL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111106821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2025-07-01
Estimated Expiration
2041-09-22

AI Technical Summary

Technical Problem

The existing Hadoop MapReduce speculates that the execution algorithm has problems such as inaccurate task remaining time estimation and low resource scheduling efficiency in heterogeneous cluster environments, especially the failure to effectively distinguish between Map Task and Reduce Task fast/slow nodes, resulting in inaccurate judgment of straggler migration and wasted system resources.

Method used

The heterogeneous MapReduce cluster speculative execution scheduling method is adopted based on reinforcement learning, and the node weight is dynamically updated through the Q-learning algorithm, identifies straglers and distinguishes between Map Task and Reduce Task fast/slow nodes. Only when specific conditions are met are met are migrated to the fast nodes, and backup tasks are started.

Benefits of technology

It improves the accuracy of estimation of the remaining running time of the task, improves the resource utilization rate of heterogeneous MapReduce clusters, and significantly improves the efficiency of large-scale data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113867944B_ABST
    Figure CN113867944B_ABST
Patent Text Reader

Abstract

The present invention relates to a speculative execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning, belonging to the field of big data processing. The present invention adopts a method for dynamically updating node weights based on Q-learning reinforcement learning, and realizes the adaptive adjustment of node weights based on historical information, effectively improving the estimation accuracy of the remaining running time of tasks; for the discrimination of whether a straggler is migrated, two conditions, namely the backup task ratio constraint and the running time constraint after migration, need to be satisfied simultaneously before the straggler can start the backup task; at the same time, the fast nodes of map tasks and the fast nodes of reduce tasks are combined, which improves the resource utilization rate of heterogeneous MapReduce clusters. The simulation test results based on typical data sets show that, compared with the existing algorithms, the algorithm proposed in this paper significantly improves the processing efficiency for large-scale data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of big data processing, and particularly relates to a speculative execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning. Background Art

[0002] Hadoop MapReduce is a framework for distributed and parallel processing of large-scale data. In a distributed cluster environment, due to reasons such as load imbalance or uneven resource distribution, the running speeds of multiple tasks of a Job may be inconsistent, slowing down the execution progress of the Job. Hadoop adopts the Speculative Execution mechanism, speculates the "lagging" task (straggler) according to specific rules, starts a backup task for it, runs it simultaneously with the original task, and selects the result output by the task that finishes first as the final result.

[0003] To solve the problems existing in the Hadoop 1.0.0 version, Hadoop 0.21.0 adopted the speculative execution mechanism based on the LATE (Longest Approximate Time to End) algorithm proposed by Zaharia et al. The LATE algorithm estimates the remaining completion time of a task based on its current running speed, and the task with the largest remaining completion time is the straggler, and a backup task is started on a fast node. However, the LATE algorithm has the following problems: 1) The weights M1, M2, R1, R2, R3 in each stage of Map Task and Reduce Task are fixed values, which are 1, 0, 1 / 3, 1 / 3, 1 / 3 respectively. However, the weights in each stage of the same task running on different nodes are not exactly the same, especially in a heterogeneous environment, the fixed weights lead to inaccurate estimation of the remaining completion time of the task, and the straggler is prone to misjudgment, and the system starts unnecessary backup tasks, resulting in low resource scheduling efficiency; 2) The LATE algorithm only divides nodes into fast nodes and slow nodes, and does not distinguish between nodes that are fast in executing Map Task and nodes that are fast in executing Reduce Task. In actual situations, some nodes are fast in executing Map Task but slow in executing Reduce Task.

[0004] To address the above problems of the LATE algorithm, Quan Chen et al. proposed the SAMR (Self-adaptive MapReduce Scheduling Algorithm) algorithm, which adaptively adjusts the weights of each stage of Map Task and Reduce Task through historical information to improve the accuracy of estimating the remaining completion time of tasks. Nodes are divided into Map Task fast nodes and Reduce Task fast nodes, and backup tasks are started on different fast nodes according to the type of straggler; SAMR performs better than the LATE algorithm in heterogeneous environments. The ESAMR algorithm adaptively adjusts the weights of each stage of Map Task and Reduce Task using the K-means algorithm.

[0005] The K-means algorithm is an unsupervised learning method that cannot accurately calculate weights. Mandana Farhang et al. proposed a speculation execution mechanism based on ANN (SEWANN, Speculative Execution with ANN), which uses the historical information (weights, amount of data processed) of tasks already executed on nodes as the input of ANN, and has a significant improvement in the accuracy of weight calculation compared with the K-means algorithm. However, the SEWANN algorithm has the following problems: 1) It does not distinguish between fast / slow nodes of Map Task and Reduce Task, and fast nodes are executing Map Task or Reduce Task; 2) The migration discrimination of stragglers does not consider the running time after migrating to fast nodes, which will result in invalid migrations and waste system resources.

[0006] To address the problems of the above algorithms, this paper proposes a speculation execution scheduling algorithm SERL (Speculative Execution with Reinforcement Learning) for heterogeneous MapReduce clusters based on reinforcement learning. Summary of the Invention

[0007] (1) Technical Problems to be Solved

[0008] The technical problem to be solved by the present invention is how to provide a speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning to solve the following problems of the SEWANN algorithm: 1) It does not distinguish between fast / slow nodes of Map Task and Reduce Task, and fast nodes are executing Map Task or Reduce Task; 2) The migration discrimination of stragglers does not consider the running time after migrating to fast nodes, which will result in invalid migrations and waste system resources.

[0009] (2) Technical Solution

[0010] To solve the above technical problems, the present invention proposes a speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning, and the method includes the following steps:

[0011] S1. Update the weights of each node in the heterogeneous MapReduce cluster according to historical information;

[0012] S2. Determine whether the currently running task i is a straggler. If so, mark it as straggler task i;

[0013] S3. Determine whether each node in the heterogeneous MapReduce cluster is a slow node;

[0014] S4. For straggler task i, determine whether to migrate it to a fast node for execution. If the condition is met, start a backup task on the fast node; otherwise, continue to run task i on the original node.

[0015] Further, the step S1 specifically includes:

[0016] S11. After the heterogeneous MapReduce cluster is started, the TaskTracker reads the historical information on the node, and the historical information includes the weight and the input data volume;

[0017] S12. The TaskTracker updates the node weight information by using the Q-learning reinforcement learning algorithm and starts the task to run;

[0018] S13. Report the running information of the completed task to the TaskTracker;

[0019] S14. The TaskTracker saves the historical information of the completed task to the node.

[0020] Further, the step S2 specifically includes:

[0021] S21. Calculate the progress value PS of task i i ;

[0022] S22. Calculate the progress rate PR of task i i ;

[0023] S23. Calculate the remaining completion time TTE of task i i ;

[0024] S24. Calculate the average remaining completion time of all the currently running tasks;

[0025] S25. Determine whether task i is a straggler.

[0026] Furthermore, for task i, its progress value PS i is:

[0027] Map process:

[0028]

[0029] Reduce process:

[0030]

[0031] where M1 and M2 are the weights of the map and sort stages of the map process respectively, and R1, R2, and R3 are the weights of the shuffle, sort, and reduce stages of the reduce process respectively; SubPS i is the progress value of task i in the current running stage, where N fi is the number of key / value pairs that task i has processed in the current running stage, and N ai is the total number of key / value pairs that task i needs to process in this stage.

[0032] Furthermore, for task i, the progress rate PR i is:

[0033]

[0034] where T i is the time that task i has been running;

[0035] For task i, its remaining completion time TTE i is:

[0036]

[0037] The average remaining completion time of all running tasks is:

[0038]

[0039] where L is the number of running tasks;

[0040] For task i, if the following conditions are met, it is determined to be a straggler,

[0041] TTE i - ATTE > ATTE * STT

[0042] Among them, STT is a constant, and STT ∈ [0, 1].

[0043] Furthermore, the step S3 specifically includes the following steps:

[0044] S31. Calculate TT i The average progress rate TrR of the upper map task and reduce task mi , TrR ri ; TT i is the i-th TaskTracker / node;

[0045] S32. The average progress rate ATrR of the map tasks on all nodes in the system m , and the average progress rate ATrR of the reduce tasks on all nodes r ;

[0046] S33. Determine whether TT i is a slow node running the map task or a slow node running the reduce task.

[0047] Furthermore, the average progress rate of the map tasks on TT i is:

[0048]

[0049] Among them, M is the number of map tasks running on TT i , and PR j is the progress rate of the j-th map task on TT i ;

[0050] The average progress rate of the reduce tasks on TT i is:

[0051]

[0052] Among them, R is the number of reduce tasks running on TT i , and PR j is the progress rate of the j-th reduce task on TT i ;

[0053] Furthermore, the average progress rate of the map tasks on all nodes in the system is:

[0054]

[0055] Among them, N is the number of all nodes in the system;

[0056] The average progress rate of reduce tasks on all nodes in the system is:

[0057]

[0058] where N is the number of all nodes in the system;

[0059] For TT i if the following conditions are met, then TT i is a slow node running map tasks:

[0060] TrR mj <(1 - STrC) * ATrR m

[0061] where STrC is a constant and STrC ∈ [0, 1];

[0062] For TT i if the following conditions are met, then TT i is a slow node running reduce tasks:

[0063] TrR rj <(1 - STrC) * ATrR r .

[0064] Furthermore, the specific steps of step S4 are as follows:

[0065] S41. Determine whether the number of backup tasks exceeds the specified ratio. If not, execute step S42; otherwise, straggler task i is executed on the original node;

[0066] S42. Determine whether the running time exceeds TTE i after straggler task i is migrated to the corresponding fast node. If not, straggler task i can be migrated to the corresponding fast node for running; otherwise, straggler task i is executed on the original node; Fast nodes include fast nodes running map tasks or fast nodes running reduce tasks. After slow nodes are identified, nodes other than slow nodes are fast nodes.

[0067] Furthermore, for straggler task i, whether to migrate to a fast node needs to meet the following two conditions:

[0068] One is that the number of backup tasks does not exceed the specified ratio, that is, it satisfies

[0069] BackupNum < BP * TaskNum

[0070] Among them, BackupNum is the number of running backup tasks, and TaskNum is the number of all running tasks; BP is the proportionality constant of the number of backup tasks to the number of all tasks, and BP ∈ [0, 1];

[0071] Second, according to the type of straggler task i, after migrating to the corresponding fast node, the running time does not exceed TTE i , that is, it satisfies

[0072] fTTE < TTE i

[0073] Among them, fTTE is the average running time of the completed tasks on the fast node, where fTTE j is the running time of the completed task j on the fast node, and U is the number of completed tasks on the fast node;

[0074] The stragglers that satisfy the above two conditions at the same time can be migrated to the fast node for running; otherwise, the running node of the straggler task i is not migrated.

[0075] (III) Beneficial effects

[0076] The present invention proposes a speculative execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning. Aiming at the problems of low estimation accuracy of the remaining time of tasks and inability to support heterogeneous cluster environments in the existing Hadoop MapReduce speculative execution algorithm, this paper proposes a speculative execution scheduling algorithm SERL for heterogeneous MapReduce clusters based on reinforcement learning. It mainly consists of 4 steps: First, the Q-learning reinforcement learning method is used to dynamically and adaptively adjust the weights of each node in the cluster based on historical information; then, compare the remaining completion time of the task with the average remaining completion time of all running tasks in the cluster to identify the stragglers; at the same time, the nodes in the cluster are divided into fast / slow nodes for map tasks and fast / slow nodes for reduce tasks. The stragglers of the map task type can be migrated to the fast nodes of the map task, improving the running efficiency after migration; finally, a discrimination is made on whether the stragglers are migrated. Only the stragglers that meet the two conditions at the same time can start the backup task, improving the utilization rate of cluster resources. The simulation test results based on typical data sets show that, compared with the existing algorithms, the algorithm proposed in this paper significantly improves the processing efficiency of large-scale data.

[0077] The present invention adopts a method for dynamically updating node weights based on Q-learning reinforcement learning, which realizes the adaptive adjustment of node weights based on historical information, effectively improving the estimation accuracy of the remaining running time of tasks.

[0078] To determine whether to migrate a straggler, two conditions need to be satisfied simultaneously: the backup task ratio constraint and the running time constraint after migration. Only when both conditions are met can the straggler start the backup task. At the same time, by combining fast nodes in map tasks and fast nodes in reduce tasks, this method improves the resource utilization rate of heterogeneous MapReduce clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 is the overall process of speculative execution scheduling based on reinforcement learning of the present invention;

[0080] Figure 2 is the process of updating node weights;

[0081] Figure 3 is the basic structure of the reinforcement learning algorithm;

[0082] Figure 4 is the flowchart for identifying stragglers;

[0083] Figure 5 is the flowchart for identifying slow nodes;

[0084] Figure 6 is the flowchart for determining whether to migrate. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0085] To make the objectives, contents, and advantages of the present invention clearer, the following further describes the detailed embodiments of the present invention in conjunction with the accompanying drawings and embodiments.

[0086] The overall process of the speculative execution scheduling algorithm SERL for heterogeneous MapReduce clusters based on reinforcement learning is as follows Figure 1 As shown, it mainly includes four steps: node weight update, straggler identification, slow node identification, and determination of whether to migrate. Among them, the historical information of nodes includes weights and input data volumes, which are stored in xml format on each node of the cluster.

[0087] S1. Update the weights of each node in the heterogeneous MapReduce cluster according to historical information;

[0088] S2. Determine whether the currently running task i is a straggler. If so, mark it as straggler task i;

[0089] S3. Determine whether each node in the heterogeneous MapReduce cluster is a slow node;

[0090] S4. For straggler task i, determine whether to migrate it to a fast node for execution. If the condition is met, start a backup task on the fast node; otherwise, continue to run task i on the original node.

[0091] The following provides a detailed introduction to each step.

[0092] S1. Node weight update

[0093] After the heterogeneous MapReduce cluster is started, the node weight update process is as follows Figure 2 shown, mainly including the following 4 steps:

[0094] S11. TaskTracker reads the historical information (weight, input data volume) on the node;

[0095] S12. TaskTracker uses the Q-learning reinforcement learning algorithm to update the node weight information and starts the task running;

[0096] S13. Report the running information (weight, running time) of the completed task to TaskTracker;

[0097] S14. TaskTracker saves the historical information of the completed task to the node.

[0098] Q-learning is a model-free reinforcement learning algorithm based on the Markov decision process, as follows Figure 3 shown. The agent is in an environment, and each state is the agent's perception of the current environment; the agent can only affect the environment through actions. When the agent executes an action, the environment will transfer to another state with a certain probability; at the same time, the environment will feedback a reward to the agent according to the potential reward function. The goal of reinforcement learning is to find an optimal policy to enable the agent to obtain as much reward as possible from the environment.

[0099] The update process of Q-learning is:

[0100] Q(s, a) ← Q(s, a) + α(r + γ · max a' Q(s', a') - Q(s, a))

[0101] where Q(s, a) is the benefit of taking action a in state s at a certain moment, α is the learning rate, r is the reward; γ is the reward decay coefficient, γ ∈ [0, 1], and the closer γ is to 1, the greater the influence of the subsequent state; maxa' Q(s', a') is the maximum Q(s', a') value in the next state s'.

[0102] S2. Straggler identification

[0103] The straggler identification process is as follows Figure 4 as shown, including the following steps:

[0104] S21. Calculate the progress value PS of task i i ;

[0105] S22. Calculate the progress rate PR of task i i ;

[0106] S23. Calculate the remaining time to completion TTE of task i i ;

[0107] S24. Calculate the average remaining time to completion of all running tasks;

[0108] S25. Determine whether task i is a straggler.

[0109] For task i, its progress value PS i is:

[0110] Map process:

[0111]

[0112] Reduce process:

[0113]

[0114] where M1 and M2 are the weights of the map and sort stages of the map process respectively, and R1, R2, and R3 are the weights of the shuffle, sort, and reduce stages of the reduce process respectively. SubPS i is the progress value of task i in the current running stage, where N fi is the number of key / value pairs that task i has processed in the current running stage, and N ai is the total number of key / value pairs that task i needs to process in this stage.

[0115] For task i, the progress rate PR i is:

[0116]

[0117] where Ti is the running time of task i.

[0118] For task i, its remaining time to completion TTE i is:

[0119]

[0120] The average remaining time to completion of all running tasks is:

[0121]

[0122] where L is the number of running tasks.

[0123] For task i, if the following condition is met, it is determined to be a straggler,

[0124] TTE i - ATTE > ATTE * STT

[0125] where STT is a constant and STT ∈ [0, 1].

[0126] S3. Slow node identification

[0127] The slow node identification process is as follows Figure 5 shown, including the following steps:

[0128] S31. Calculate TT i (On the i-th TaskTracker / node) the average progress rate TrR of map tasks and reduce tasks mi 、TrR ri ;

[0129] S32. The average progress rate ATrR of map tasks on all nodes in the system m , and the average progress rate ATrR of reduce tasks on all nodes r ;

[0130] S33. Determine whether TT i is a slow node running map tasks or a slow node running reduce tasks.

[0131] TT i (On the i-th TaskTracker / node) the average progress rate of map tasks is:

[0132]

[0133] where M is TT iThe number of map tasks running on, PR j is TT i The progress rate of the j-th map task on

[0134] TT i (For the i-th TaskTracker / node), the average progress rate of reduce tasks is:

[0135]

[0136] where R is the number of reduce tasks running on TT i and PR j is the progress rate of the j-th reduce task on TT i

[0137] The average progress rate of map tasks on all nodes in the system is:

[0138]

[0139] where N is the number of all nodes in the system.

[0140] The average progress rate of reduce tasks on all nodes in the system is:

[0141]

[0142] where N is the number of all nodes in the system.

[0143] For TT i , if the following conditions are met, then TT i is a slow node for running map tasks:

[0144] TrR mj <(1 - STrC)*ATrR m

[0145] where STrC is a constant and STrC ∈ [0, 1].

[0146] For TT i , if the following conditions are met, then TT i is a slow node for running reduce tasks:

[0147] TrR rj <(1 - STrC)*ATrR r

[0148] S4. Migration Judgment

[0149] The migration judgment process is as follows​Figure 6 As shown in the figure, it includes the following steps:

[0150] S41. Determine whether the number of backup tasks exceeds the specified ratio. If not, execute step S42; otherwise, the straggler task i is executed on the original node.

[0151] S42. Determine whether the running time exceeds TTE after the straggler task i is migrated to the corresponding fast node (the fast node running the map task or the fast node running the reduce task). i If not, the straggler task i can be migrated to the corresponding fast node to run; otherwise, the straggler task i is executed on the original node.

[0152] For the straggler task i, whether to migrate to the fast node needs to meet the following two conditions:

[0153] One is that the number of backup tasks does not exceed the specified ratio, that is, it satisfies

[0154] BackupNum < BP * TaskNum

[0155] where BackupNum is the number of running backup tasks, TaskNum is the number of all running tasks; BP is the proportional constant of the number of backup tasks to the number of all tasks, BP ∈ [0, 1], and the default value is 0.1.

[0156] The second is that according to the type of the straggler task i (map task or reduce task), after migrating to the corresponding fast node (after identifying the slow node, the nodes other than the slow node are fast nodes; the fast node running the map task or the fast node running the reduce task), the running time does not exceed TTE i That is, it satisfies

[0157] fTTE < TTE i

[0158] where fTTE is the average running time of the completed tasks on the fast node, (where fTTE j is the running time of the completed task j on the fast node), and U is the number of completed tasks on the fast node.

[0159] The straggler that meets the above two conditions at the same time can be migrated to the fast node to run; otherwise, the running node of the straggler task i is not migrated.

[0160] In view of the problems of the existing Hadoop MapReduce speculative execution algorithm, such as low estimation accuracy of the remaining time of tasks and inability to support heterogeneous cluster environments, this paper proposes a speculative execution scheduling algorithm SERL for heterogeneous MapReduce clusters based on reinforcement learning. It mainly consists of four steps: First, the Q-learning reinforcement learning method is adopted to dynamically and adaptively adjust the weights of each node in the cluster based on historical information; then, the remaining completion time of the task is compared with the average remaining completion time of all running tasks in the cluster to identify stragglers; at the same time, the nodes in the cluster are divided into fast / slow nodes for map tasks and fast / slow nodes for reduce tasks. Stragglers of the map task type can be migrated to fast map task nodes, improving the running efficiency after migration; finally, a discrimination is made on whether to migrate stragglers. Only stragglers that meet both conditions can start backup tasks, improving the resource utilization rate of the cluster. The simulation test results based on typical data sets show that, compared with the existing algorithms, the algorithm proposed in this paper significantly improves the processing efficiency of large-scale data. Weight update based on small sample learning is the next research direction.

[0161] The advantages of the present invention are as follows:

[0162] Adopt a dynamic update method of node weights based on Q-learning reinforcement learning, and adaptively adjust the node weights based on historical information, effectively improving the estimation accuracy of the remaining running time of tasks;

[0163] Make a discrimination on whether to migrate stragglers. Stragglers can start backup tasks only when they meet both the backup task ratio constraint and the running time constraint after migration; at the same time, combine fast map task nodes and fast reduce task nodes, which improves the resource utilization rate of heterogeneous MapReduce clusters.

[0164] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and deformations can be made, and these improvements and deformations should also be regarded as the protection scope of the present invention.

Claims

1. A speculative execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning, characterized in that The method includes the following steps: S1. Update the weights of each node in the heterogeneous MapReduce cluster according to historical information; S2. Determine whether the running task i is a straggler. If so, mark it as straggler task i; S3. Determine whether each node in the heterogeneous MapReduce cluster is a slow node; S4. For straggler task i, determine whether to migrate it to a fast node for execution. If the conditions are met, start a backup task on the fast node. Otherwise, continue to run task i on the original node; Wherein, The specific steps of step S3 include the following steps: S31. Calculate TT i The average progress rate TrR of the upper map task and reduce task mi , TrR ri ; TT i is the i-th TaskTracker / node; S32. Calculate the average progress rate ATrR of map tasks on all nodes in the computing system m , and the average progress rate ATrR of reduce tasks on all nodes r ; S33. Determine TT i as a slow node running a map task or a slow node running a reduce task; TT i The average progress rate of the upper map task is: Among them, M is the number of map tasks running on TT i and PR j is the progress rate of the j-th map task on TT i ; TT i The average progress rate of the upper reduce task is: where R is the number of reduce tasks running on TT i and PR j is the progress rate of the j-th reduce task on TT i .

2. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 1, wherein The specific steps of step S1 include: S11. After the heterogeneous MapReduce cluster is started, the TaskTracker reads the historical information on the node, and the historical information includes weights and input data volume; S12. The TaskTracker updates the node weight information using the Q-learning reinforcement learning algorithm and starts the task to run; S13. Report the running information of the completed task to the TaskTracker; S14. The TaskTracker saves the historical information of the completed task to the node.

3. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 2, characterized in that, The specific steps of step S2 include: S21. Calculate the progress value PS of task i i ; S22. Calculate the progress rate PR of task i i ; S23. Calculate the remaining completion time TTE of task i i ; S24. Calculate the average remaining completion time of all running tasks; S25. Determine whether task i is a straggler.

4. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 3, characterized in that, For task i, its progress value PS i is as follows: Map process: Reduce process: Among them, M1 and M2 are the weights of the map and sort stages in the map process respectively, and R1, R2, and R3 are the weights of the shuffle, sort, and reduce stages in the reduce process; SubPS i is the progress value of task i in the current running stage, where N fi is the number of key / value pairs that have been processed by task i in the current running stage, and N ai is the total number of key / value pairs that need to be processed by task i in this stage.

5. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 4, characterized in that, For task i, the progress rate PR i is as follows: where T i is the time that task i has been running; For task i, its remaining time to completion TTE i is as follows: The average remaining completion time of all running tasks is: Wherein, L is the number of running tasks; For task i, if the following conditions are met, it is determined as a straggler, TTE i -ATTE > ATTE * STT Wherein, STT is a constant, STT ∈ [0,1].

6. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 1, wherein The average progress rate of map tasks on all nodes in the system is: Wherein, N is the number of all nodes in the system; The average progress rate of reduce tasks on all nodes in the system is: Wherein, N is the number of all nodes in the system; For TT i , TT i is a slow node running the map task if the following conditions are met: TrR mj <(1 - STrC) * ATrR m Wherein, STrC is a constant, STrC ∈ [0,1]; For TT i , TT i is a slow node running the reduce task if the following conditions are met: TrR rj <(1 - STrC) * ATrR r 。 7. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 1 or 6, characterized in that The specific steps of step S4 include the following steps: S41. Determine whether the number of backup tasks exceeds the specified ratio. If not, execute step S42; otherwise, straggler task i is executed on the original node; S42. Determine whether the running time of straggler task i exceeds TTE after it is migrated to the corresponding fast node i . If not, straggler task i can be migrated to the corresponding fast node for running; otherwise, straggler task i is executed on the original node. The fast nodes include the fast nodes running map tasks or the fast nodes running reduce tasks. After the slow nodes are identified, the nodes other than the slow nodes are fast nodes.

8. The speculation execution scheduling method for heterogeneous MapReduce clusters based on reinforcement learning according to claim 7, characterized in that For straggler task i, whether to migrate to a fast node needs to meet the following two conditions: One is that the number of backup tasks does not exceed the specified ratio, that is, it satisfies BackupNum < BP * TaskNum Wherein, BackupNum is the number of running backup tasks, TaskNum is the number of all running tasks; BP is the proportional constant of the number of backup tasks to the number of all tasks, BP ∈ [0,1]; Second, after migrating to the corresponding fast node according to the type of straggler task i, the running time does not exceed TTE i , that is, it satisfies fTTE < TTE i Among them, fTTE is the average running time of the tasks completed on the fast nodes, where fTTE j is the running time of the completed task j on the fast node, and U is the number of tasks completed on the fast node; The straggler that meets the above two conditions at the same time can be migrated to the fast node to run; otherwise, the running node of straggler task i is not migrated.

Citation Information

Patent Citations

  • Task scheduling optimization method in resource imbalance Spark environment

    CN110413389A

  • Load prediction-based hadoop computing task speculative execution method

    WO2020248227A1