A distributed computing method and system based on MDS coding and flexible scheduling strategy

By adopting MDS encoding and flexible scheduling strategies in the master-slave distributed computing model and adjusting task encoding and scheduling according to system load, the delay problem caused by lagging nodes is solved, the task execution time is optimized, and the system efficiency is improved.

CN114756381BActive Publication Date: 2025-09-19NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210555877.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-09-19
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

In the master-slave distributed computing model, lagging nodes cause the overall task execution time to be extended. The existing heartbeat detection method increases communication overhead and re-execution of tasks increases delays. The existing encoding computing model fails to effectively optimize task delays under high load.

Method used

Adopting MDS encoding and flexible scheduling strategies, the task encoding scheme and scheduling strategy are adjusted according to the system load. The optimal encoding scheme and scheduling strategy are found through iteration, the task segmentation and allocation are optimized, the dependence on lagging nodes is reduced, and real-time adjustments are made to adapt to changes in the task arrival rate.

Benefits of technology

It effectively reduces the average computing time of tasks in distributed computing, reduces the task queuing time under high load, solves the problem of laggards, and improves the overall efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114756381B_ABST
    Figure CN114756381B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed computing method and system based on MDS encoding and a flexible scheduling strategy. This method is oriented towards a master-slave distributed computing framework. Based on different task arrival rates, and considering the non-negligible overhead of revoking redundant tasks, the method, by designing an appropriate model encoding scheme and task scheduling strategy, alleviates the stragglers problem in the distributed system while balancing the number of redundant tasks with the system load, thereby reducing the average execution time of the overall task. Furthermore, the method considers the situation in which new and old tasks exist after adjusting the model encoding scheme and task scheduling strategy. By designing a compatible solution that distinguishes task types, it is possible to avoid invalidating the calculation results of tasks before the adjustment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of distributed systems, and in particular relates to a distributed computing method and system based on MDS coding and flexible scheduling scheme. Background Art

[0002] In the existing master-slave distributed computing model, the master node needs to obtain the computation results of all worker nodes before it can integrate the overall task results. The entire task execution time is determined by the slowest worker node, which becomes the computing bottleneck. Low node efficiency significantly reduces the efficiency of the entire system. Even for homogeneous worker nodes, the time it takes to run the same task varies due to hardware instability and frequent resource contention. Some nodes can even take five to eight times the average execution time.

[0003] To alleviate the stragglers problem in distributed computing, distributed computing frameworks such as Hadoop and Spark use heartbeat detection to check for stragglers in the system. If a worker node fails to send a heartbeat message to the master node in a timely manner, the master node will reassign the task to another node for execution. However, this approach incurs additional communication overhead, and re-executing tasks increases overall computational latency. In recent years, researchers have proposed an encoding computation model. When a computation task satisfies the characteristics of a linear operation, the task can be sliced ​​and encoded, and then sent to different worker nodes for execution. In this model, the master node only needs to receive the results returned by some of the worker nodes to decode the overall task result. Since it does not need to wait for the results of stragglers, the overall task computation time is significantly reduced. In the encoding computation model, the task encoding scheme and task scheduling strategy have a significant impact on the overall task processing time. There is a need to further optimize the overall task latency of existing encoding computation models under different loads. Summary of the Invention

[0004] The purpose of this invention is to propose a method for optimizing the overall task delay based on the system computing load in a computing scenario where streaming linear tasks are executed under a master-slave distributed computing framework, so as to solve the problem of laggards in distributed computing and avoid the problem of greatly increasing task queuing time under high load.

[0005] The present invention also provides a corresponding distributed computing system.

[0006] In order to achieve the above-mentioned object of the invention, the present invention adopts the following technical solutions:

[0007] In a first aspect, a distributed computing method based on MDS coding and a flexible scheduling strategy comprises the following steps:

[0008] Based on the computing power of a single working node and the task arrival rate in the system, the average task computation time T in the system is obtained. The computing power includes the average time it takes for a working node to execute the entire task and the time it takes to cancel a task. The task arrival rate is the number of tasks arriving at the master node per unit time. The average task computation time is the time from when a task arrives at the master node to when the master node decodes the overall task result.

[0009] Iteratively search for a model encoding scheme and task scheduling strategy that minimizes the average computation time T of tasks in the system. The model encoding scheme is used to determine the number of blocks into which the task model is segmented, and the task scheduling strategy is used to determine which worker nodes the tasks are scheduled to execute.

[0010] Execute the code and place the task model on the worker node;

[0011] When the master node receives a task, it dispatches it to the worker node. The worker node obtains the task from the task queue, processes it, and returns the processing result to the master node. The master node collects the results from the worker nodes and decodes the overall task result.

[0012] Furthermore, the calculation formula for the average calculation time T of tasks in the system is as follows:

[0013]

[0014] Where k is the encoding scheme, its value represents the number of blocks of the model, r is the scheduling strategy, its value represents the number of scheduled tasks to the working nodes, n is the number of all working nodes, H i is the harmonic series of order i, and its formula is expressed as μ a is the average rate at which a single node processes tasks, and its calculation formula is: ρ is the load of a single node, and its calculation formula is: W i r,k is a fixed coefficient, and its calculation formula is in is the number of combinations, which means the total number of choices of j numbers from r different numbers. It is also a fixed coefficient, and its calculation formula is as follows:

[0015]

[0016] Furthermore, the iterative search for the model encoding scheme and task scheduling strategy that minimizes the average computation time T of tasks in the system includes:

[0017] Set the encoding scheme to a random integer, denoted as k, and set the task scheduling strategy to a random integer, denoted as r, satisfying 1≤k≤r≤n;

[0018] The fixed coding scheme is k, and all task scheduling strategies r are traversed to find the scheduling strategy that minimizes the average computing time T of the task. The optimal scheduling strategy is r * ;

[0019] Fixed task scheduling strategy is r * , traverse all encoding schemes k, find the encoding scheme that minimizes the average computing time T of the task, and record the optimal encoding scheme as k * ;

[0020] Set the encoding scheme k=k * , repeat the above operations of finding the optimal scheduling strategy and the optimal coding scheme until the obtained coding scheme k and scheduling strategy r can no longer reduce the average computing time T of the near task;

[0021] Get the coding scheme k=k * , task scheduling strategy r=r * .

[0022] Furthermore, executing the code and placing the task model on the worker node includes:

[0023] Divide the task model into k parts according to the obtained encoding scheme k. If the task model cannot be divided exactly, add special elements at the end of the model. For example, if the model is a matrix, you can add several rows of zero elements at the end.

[0024] For the k-partitioned model data, use (n, k) MDS code to generate n new coded data blocks, where n is the number of working nodes and satisfies n ≥ k;

[0025] The generated n encoded data blocks are saved on n working nodes respectively.

[0026] Furthermore, the master node dispatches the task to the worker node upon receiving the task, including:

[0027] When the master node receives a task from a user, it randomly selects r nodes from n working nodes according to the obtained scheduling strategy r to send the task, where r≤n;

[0028] The worker node that receives the task uses the queue to save the task.

[0029] Furthermore, the working node obtains tasks from the task queue and processes the tasks including:

[0030] When the task queue of a working node is not empty and the node is in an idle state, the node obtains tasks from the head of the queue;

[0031] When the task obtained is an untagged task, the worker node will calculate the task with the model stored on itself and send the calculation result to the master node after the calculation is completed. The untagged task specifically refers to the task that needs to be calculated with the model on the worker node;

[0032] When the acquired task is a marked task, the working node performs a cancel operation to remove the task from the task queue. The marked task refers to a redundant computing task that does not need to be calculated with the model.

[0033] Furthermore, the master node collects the results from the worker nodes and decodes the overall task result, including:

[0034] The master node stores all the returned results from the worker nodes. When the number of results for a task reaches k, the master node decodes the overall calculation result of the task according to the properties of the MDS code and then sends the result to the user.

[0035] For a task that receives k results, the master node will notify the rk worker nodes that have not completed the task to cancel the execution of the task. After receiving the notification, the worker node will find the task from the task queue and mark it.

[0036] Furthermore, the method further includes: adjusting the model encoding scheme and task scheduling strategy in real time according to the task arrival rate to make the execution of new and old tasks compatible.

[0037] Furthermore, the real-time adjustment of the model encoding scheme and task scheduling strategy according to the task arrival rate to ensure compatibility with the execution of new and old tasks includes:

[0038] Set a threshold for the change in task arrival rate. When the change in task arrival rate exceeds the threshold, redesign the model encoding scheme and task scheduling strategy.

[0039] Encode the task model using the new encoding scheme and save it to all working nodes;

[0040] The master node uses the new scheduling policy to schedule newly arrived tasks to the worker nodes, and these tasks are calculated with the new model in the worker nodes. For the unfinished tasks in the queue, they are calculated with the old model in the worker nodes.

[0041] The master node distinguishes between new and old task results and decodes the overall task result using different decoding methods;

[0042] When all old tasks in the task queue of a working node are completed, the working node deletes the old task model data.

[0043] In a second aspect, a distributed computing system based on MDS encoding and flexible scheduling strategies is provided. The work cluster consists of a master node and multiple worker nodes. The master node is responsible for scheduling tasks and collecting and integrating task results. The worker nodes are responsible for calculating subtask results and returning them to the master node. The task models in the worker nodes are encoded according to a predetermined encoding scheme. Upon receiving a task, the master node schedules the task to the worker node according to a predetermined scheduling strategy. The encoding scheme and scheduling strategy are obtained according to the following method:

[0044] Based on the computing power of a single working node and the task arrival rate in the system, the average task computation time T in the system is obtained. The computing power includes the average time it takes for a working node to execute the entire task and the time it takes to cancel a task. The task arrival rate is the number of tasks arriving at the master node per unit time. The average task computation time is the time from when a task arrives at the master node to when the master node decodes the overall task result.

[0045] Iteratively search for a model encoding scheme and a task scheduling strategy that minimize the average computation time T of tasks in the system. The model encoding scheme is used to determine the number of blocks into which the task model is divided, and the task scheduling strategy is used to determine which work nodes the tasks are scheduled to execute.

[0046] Beneficial effects: The present invention is based on a computing scenario in which streaming linear tasks are executed under a master-slave distributed computing framework. It adopts MDS code to encode the task model and uses a redundant task scheduling strategy. By approximating the average task delay under different task arrival rates, the model encoding scheme and the task scheduling strategy are adjusted to optimize the approximate task delay. This not only solves the problem of stragglers in distributed computing, but also avoids the problem of greatly increasing the task queuing time under high load, thereby reducing the average task computing time. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 This is a flow chart of distributed computing based on MDS codes and flexible scheduling strategies provided by an embodiment of the present invention;

[0048] Figure 2 is a schematic diagram of a distributed matrix computing architecture according to an embodiment of the present invention;

[0049] Figure 3 This is a flow chart of designing a coding scheme and a scheduling strategy in an embodiment of the present invention;

[0050] Figure 4(a) and 4(b) They are respectively flow charts of selecting the optimal coding scheme and selecting the optimal task scheduling strategy in the embodiments of the present invention;

[0051] Figure 5 This is a schematic diagram of a working node executing and canceling a task in an embodiment of the present invention;

[0052] Figure 6 It is a schematic diagram of adjusting the coding scheme and task scheduling strategy in an embodiment of the present invention. DETAILED DESCRIPTION

[0053] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to specific embodiments and accompanying drawings.

[0054] Existing encoding computation models send tasks to all worker nodes for execution to minimize computational latency. However, in scenarios where tasks arrive continuously, tasks need to be cached in the nodes' task queues. Scheduling tasks to all nodes increases system load, thereby increasing task queueing time. While more redundant tasks can allow the system to tolerate more stragglers, they also require the system to cancel more redundant tasks. This takes time for worker nodes to cancel these tasks, increasing the queue time for tasks following the task in question. When the task arrival rate is low, adding redundant tasks has a smaller impact on queueing time because fewer tasks are cached in the queue. When the task arrival rate is high, adding redundant tasks has a greater impact on queueing time because more tasks are cached in the queue. Research by the current named inventors has found that, when the task arrival rate is high, scheduling tasks to only a subset of worker nodes can reduce overall task computation time compared to scheduling them to all worker nodes. Therefore, it is important to select an appropriate number of redundant tasks for different task arrival rates to reduce overall task computation time.

[0055] The number of task divisions in a coding calculation also affects the overall task calculation time. When the number of task divisions is small, worker nodes take longer to execute a single task, the master node needs to receive fewer subtask results, and the system can tolerate more stragglers. Conversely, when the number of task divisions is large, worker nodes take less time to execute a single task, the master node needs to receive more subtask results, and the system can tolerate fewer stragglers. Therefore, the number of task divisions needs to be adjusted based on the number of stragglers in the system. At the same time, the number of task divisions is limited by the number of redundant tasks: the sum of the number of task divisions and the number of redundant tasks cannot exceed the total number of worker nodes.

[0056] Based on the above understanding, the present invention proposes a method for adjusting the task coding scheme and task scheduling strategy in real time according to the system operation status to optimize the overall task delay.

[0057] The following example uses matrix-vector multiplication to illustrate the distributed computing method proposed in this invention, based on MDS codes (maximum distance separable codes) and a flexible scheduling strategy. The mathematical expression for matrix-vector multiplication is y = Ax, where A is a very large two-dimensional matrix, x is a one-dimensional vector, the number of columns in A is equal to the length of vector x, and y is the result of the multiplication of A and x.

[0058] Before the system begins providing computing services, it is necessary to evaluate the computing power of the worker nodes and the average task arrival rate in the system. Generally speaking, evaluating the computing power of a single worker node primarily involves obtaining the average computation time of the node's computing tasks and the average cancellation time of the node for computing tasks. According to one embodiment, by simulating the task execution process, the entire model is stored in a single worker node, multiple tasks are generated and executed on the node, and the execution time of these tasks is calculated to obtain the average computation time of the tasks. According to one embodiment, by simulating the task cancellation process, multiple tasks are generated and sent to the worker nodes, and the worker nodes are responsible for canceling these tasks. The cancellation time of these tasks is calculated to obtain the average cancellation time of the tasks. According to one embodiment, evaluating the average task arrival rate in the system includes: simulating the system operation process, with the master node using a task queue to store tasks arriving in the system; the master node counting the number of tasks arriving in the queue per unit time, where the number of tasks can be obtained by subtracting the length of the task queue at time t from the length of the task queue at time t+1, where t represents a discrete time and t∈{0, 1, 2, ...}; and calculating the average rate at which tasks arrive in the master node's task queue over a period of time.

[0059] In the present invention, the time it takes for a single node to execute a task is modeled as an exponential distribution with a parameter μ, based on the average time it takes for a single node to execute the entire task. In this case, 1 / μ is the average time it takes for a single node to execute the entire task. In the present invention, the time it takes for a single node to cancel a task is modeled as an exponential distribution with a parameter μ. c exponential distribution, where 1 / μ c is the average time it takes for a single node to cancel a task; based on the task arrival rate λ, the number of tasks arriving at the system per unit time is modeled as a Poisson distribution with parameter λ; based on the task arrival rate λ, the task execution rate μ and the task cancellation rate μ c , when the task encoding scheme k and task scheduling strategy r are given, the approximation of the average computing time of the task can be obtained.

[0060] In this embodiment of the present invention, the results of Ax are calculated multiple times on different working nodes, and the average time taken for these calculations is recorded. Assuming the average time taken is 100 milliseconds, the average rate at which the node executes the calculation task Ax is 0.01 / millisecond, which is recorded as μ = 0.01. It is also necessary to evaluate the time taken for the working node to cancel tasks. The master node continuously sends tasks to the working node. The working node cancels each task and records the time taken to cancel the task. The average time taken to cancel the task is then calculated. Assuming the average time taken is 2 milliseconds, the average rate at which the node cancels a single task is 0.5 / millisecond, which is recorded as μ. c =0.5. We also need to evaluate the system's average task arrival rate. The master node uses a queue to store tasks sent by users, records the growth of the queue length every millisecond, and averages the recorded values ​​over time to obtain the average task arrival rate. Assume that the system's average task arrival rate is 0.01 tasks per millisecond, denoted by λ = 0.01.

[0061] When the rate μ of computing tasks of working nodes in the system is obtained, the rate μ of canceling tasks is obtained. c , when the task arrival rate is λ, for any encoding scheme k and task scheduling strategy r, the average task computing time T of the system can be evaluated, which is calculated as follows:

[0062]

[0063] Where k is the number of blocks of the model, r is the number of scheduled tasks to the working nodes, n is the number of all working nodes, H i is the harmonic series of order i, and its formula is expressed as μ a is the average rate at which a single node processes tasks, and its formula is expressed as ρ is the load of a single node, and its formula is expressed as W i r,k is a fixed coefficient, and its formula is expressed as in is the number of combinations, which means the total number of choices of j numbers from r different numbers. It is also a fixed coefficient, and its calculation formula is as follows:

[0064]

[0065] The calculation method proposed in the present invention can adjust the encoding scheme and scheduling strategy in real time according to different task arrival rates, optimize the average calculation time of the tasks evaluated above, and further optimize the actual task execution time.

[0066] Reference Figure 1As shown, the distributed computing method based on MDS code and flexible scheduling strategy proposed in the present invention includes the following steps:

[0067] S1. Design model encoding scheme and task scheduling strategy for the system;

[0068] S2, the system encodes and places the model matrix A according to the encoding scheme;

[0069] S3. The master node dispatches tasks to the worker nodes according to the scheduling strategy. After the worker nodes complete the tasks, they send the results to the master node.

[0070] S4. The master node collects the task results and decodes the overall task results;

[0071] S5. The working node cancels redundant tasks;

[0072] S6. When the change in the task arrival rate exceeds the threshold, redesign the model encoding scheme and task scheduling strategy.

[0073] Figure 2 The figure shows a schematic diagram of the distributed matrix computing system architecture according to an embodiment of the present invention. The computing system consists of a master node M and five working nodes W1, W2, W3, W4, and W5. User tasks continuously arrive at the system at a rate of λ. The matrix A is divided into three matrices of equal size, namely A1, A2, and A3, and encoded into five coding matrices according to the (5, 3) MDS code, namely These five encoded matrices are stored in different working nodes. When the user task arrives at the master node, the master node randomly selects four working nodes to schedule the task, such as Figure 1 In the example, task x1 is scheduled to worker nodes W1, W2, W3, and W4, while task x2 is scheduled to worker nodes W2, W3, W4, and W5. Worker nodes execute tasks in the task queue on a first-come, first-served basis and send the task results to the master node upon completion. After decoding a task result, a worker node notifies another worker node to cancel redundant tasks. When a worker node receives a redundant task from the task queue, it performs the cancellation operation. The following details the design of the encoding scheme and scheduling strategy, as well as the specific process of task execution and cancellation.

[0074] Figure 3 The design flow chart of matrix coding scheme and task scheduling strategy is given. First, the coding scheme k and scheduling strategy r are randomly generated, and they satisfy 1≤k≤r≤n. Then, according to k and r, they are substituted into the above formula for calculating the average task time T to obtain the average task computing time under the current scheme, which is recorded as T oldNext, we adjust the encoding scheme and scheduling strategy respectively. First, we adjust the encoding scheme k with the fixed scheduling strategy r. The specific process is shown in Figure 4(a). We try all possible encoding schemes and calculate the average computing time of the tasks under these encoding schemes and the fixed scheduling strategy r. Finally, we return the encoding scheme k that minimizes the computing time. * ; Secondly, the fixed coding scheme k * Adjust the scheduling strategy. The specific process is shown in Figure 4(b). Try all possible scheduling strategies and calculate the average computing time of tasks under these scheduling strategies and fixed coding scheme k. Finally, return the scheduling strategy r that minimizes the computing time. * Note that the fixed scheduling strategy is the initial scheduling strategy when it is executed for the first time, and the optimal scheduling strategy r obtained in the previous round is used in subsequent executions. * , and the above fixed coding scheme is always the optimal coding scheme k obtained in the previous step * After adjusting a round of coding scheme and scheduling strategy, according to the coding scheme k * and scheduling strategy r * The average computing time of the computing task is recorded as T new , when the time T after adjusting the plan new Less than T old When k is reached, the next round of adjustment is required. The initialization coding scheme and scheduling strategy at this time are the results of the previous round. The process of adjusting the coding scheme and task scheduling strategy will continue to iterate until the k * , r * The average computing time of the task obtained is T new Greater than the time T calculated in the previous round old In a specific embodiment of the present invention, the rate of computing tasks of a working node is μ=0.01, and the rate of canceling tasks is μ c =0.5, the task arrival rate in the system is λ=0.01, the number of working nodes is 5, according to Figure 3 The process shown obtains the coding scheme k=3 and the scheduling strategy r=4 after 3 iterations.

[0075] Then, according to the obtained coding scheme k=3, the matrix A is divided into three parts, and the (5, 3) MDS code is used to encode it into five matrix blocks, which are placed in different working nodes respectively, as shown in the following example: Figure 5 As shown in (a), the working node W i Store the encoded matrix When the matrix is ​​placed, the system can receive computing tasks. Tasks continuously arrive at the master node, and the master node randomly selects r = 4 working nodes to send tasks, such as Figure 5As shown in (a), task x1 is scheduled by the master node to worker nodes W1, W2, W3, and W4, and task x2 is scheduled by the master node to worker nodes W2, W3, W4, and W5. When the task queue of a worker node is not empty and the node is idle, it will get the task from the head of the task queue and execute the task. For example, when node W1 gets task x1, it will execute the calculation task. When the working node completes the local calculation, it sends the calculation results to the master node, such as Figure 5 As shown in (b), worker nodes W1, W3, and W4 complete the calculation of task x1 and send the calculation results. and To the master node. Based on the properties of MDS code, the master node can decode the result of Ax1 without waiting for the calculation result of node W2, which alleviates the problem of laggards in distributed computing. Since the master node has obtained the calculation result of Ax1, it will send a message to notify node W2 to cancel the execution of task x1. Figure 5 (b) and (c) show the process of canceling redundant task x1. First, after receiving the cancellation notification, node W2 will mark task x1 and mark task x1 as x′1. Then, when node W2 obtains task x′1, it will perform the cancellation operation and directly remove the task from the task queue. Figure 5 (d) shows the intermediate results of the distributed matrix calculation. The newly arrived task x4 is dispatched by the master node to the worker nodes W1, W2, W4 and W5. Node W1 completes the calculation of task x3 and sends the result to Sent to the master node, nodes W4 and W5 complete task x2 and send the result After the data is sent to the master node, the master node needs to wait for node W2 or W3 to complete task x2 before decoding the calculation result of Ax2. Similarly, it needs to wait for any two of nodes W3, W4 and W5 to complete task x3 before decoding the calculation result of Ax3.

[0076] Figure 6 The figure shows a schematic diagram of the adjustment coding scheme and task scheduling strategy when the task arrival rate changes. Figure 6 (a) is a diagram of the system operation status at a certain moment when the task arrival rate is 0.01. If the task arrival rate increases to 0.02 at this moment, and the task arrival rate change threshold set by the system is 0.01, then it is necessary to follow Figure 3 The process shown in the figure readjusts the coding scheme and task scheduling strategy. According to the specific parameters of the embodiment of the present invention, the adjusted coding scheme k=4 and the task scheduling strategy r=5 can be obtained. Therefore, it is necessary to re-encode the matrix A using the (4, 5) MDS code, as shown in FIG. Figure 6 As shown in (b), the matrix is ​​first divided into 4 blocks of the same size, namely A1, A2, A3 and A4, and then encoded into 5 blocks, namely The five encoded matrices are stored in different worker nodes. The master node now schedules task x5 according to the new scheduling strategy and dispatches task x5 to all worker nodes. The worker nodes can distinguish between the original tasks and the new tasks, such as distinguishing between tasks x4 and x5. For the original tasks, the worker nodes use the original saved matrices for calculations. For example, node W1 will perform the calculation task for task x4. For new tasks, the working node uses the newly saved matrix for calculation. For example, node W1 will perform the calculation task for task x5. The master node also needs to distinguish between the original task and the new task, and use different decoding methods to decode the overall task calculation results. When all the old tasks are completed, the working node can delete the original matrix to save storage space. For example, node W1 can be deleted after completing task x4.

[0077] It should be noted that the distributed computing method based on MDS coding and flexible scheduling strategy proposed in the present invention can be used not only for matrix calculation tasks, but also for various tasks with linear operation properties. For ordinary technicians in this field, without departing from the principles of the present invention, several modifications or equivalent substitutions can be made to the specific implementation methods of the present invention, all of which should be covered within the scope of protection of the claims of the present invention.

Claims

1. A distributed computing method based on MDS coding and flexible scheduling strategy, characterized in that: The method comprises the following steps: The average task computation time T in the system is obtained based on the computing power of a single working node and the task arrival rate in the system. The computing power includes the average time it takes for a working node to execute the entire task and the time required to cancel the task. The task arrival rate is the number of tasks arriving at the master node per unit time. The average task computation time T in the system is the time from when the task arrives at the master node to when the master node decodes the overall task result. The calculation formula is as follows: Where k is the encoding scheme, and its value represents the number of blocks of the model. r is the scheduling strategy, and its value represents the number of scheduled tasks to the working nodes. i is the harmonic series of order i, μ a is the average rate at which a single node processes tasks, ρ is the load of a single node, is a fixed coefficient; Iteratively search for a model encoding scheme and task scheduling strategy that minimizes the average computation time T of tasks in the system. The model encoding scheme is used to determine the number of blocks into which the task model is segmented, and the task scheduling strategy is used to determine which worker nodes the tasks are scheduled to execute. Execute the code and place the task model on the worker node; When the master node receives a task, it dispatches it to the worker node. The worker node obtains the task from the task queue, processes it, and returns the processing result to the master node. The master node collects the results from the worker nodes and decodes the overall task result.

2. The method according to claim 1, characterized in that H i The formula is expressed as μ a The calculation formula is μ is the rate at which the worker node computes tasks, μ c is the rate of task cancellation; the calculation formula of ρ is λ is the task arrival rate, n is the number of working nodes; The calculation formula is in is the number of combinations, which means the total number of choices of j numbers from r different numbers. It is also a fixed coefficient, and its calculation formula is as follows:

3. The method according to claim 1, characterized in that The iterative search for the model encoding scheme and task scheduling strategy that minimizes the average computation time T of tasks in the system includes: Set the encoding scheme to a random integer, denoted as k, and set the task scheduling strategy to a random integer, denoted as r, satisfying 1≤k≤r≤n, where n is the number of working nodes; The fixed coding scheme is k, and all task scheduling strategies r are traversed to find the scheduling strategy that minimizes the average computing time T of the task. The optimal scheduling strategy is r * ; Fixed task scheduling strategy is r * , traverse all encoding schemes k, find the encoding scheme that minimizes the average computing time T of the task, and record the optimal encoding scheme as k * ; Set the encoding scheme k=k * , repeat the above operations of finding the optimal scheduling strategy and the optimal encoding scheme until the obtained encoding scheme k and scheduling strategy r can no longer reduce the average computing time T of the task; Get the coding scheme k=k * , task scheduling strategy r=r * .

4. The method according to claim 1, wherein Executing the code and placing the task model on the worker node involves: Divide the task model into k parts according to the obtained encoding scheme k. If the task model cannot be divided exactly, add the specified element at the end of the model. For the k-partitioned model data, use the (n,k) MDS code to generate n new coded data blocks, where n is the number of working nodes and satisfies n ≥ k; The generated n encoded data blocks are saved on n working nodes respectively.

5. The method according to claim 1, wherein When the master node receives a task, it dispatches the task to the working node, including: When the master node receives a task from a user, it randomly selects r nodes from n working nodes according to the obtained scheduling strategy r to send the task, where r≤n; The worker node that receives the task uses the queue to save the task.

6. The method according to claim 1, characterized in that The working node obtains tasks from the task queue and processes the tasks including: When the task queue of a working node is not empty and the node is in an idle state, the node obtains tasks from the head of the queue; When the task obtained is an untagged task, the worker node will calculate the task with the model stored on itself and send the calculation result to the master node after the calculation is completed. The untagged task specifically refers to the task that needs to be calculated with the model on the worker node; When the acquired task is a marked task, the working node performs a cancel operation to remove the task from the task queue. The marked task refers to a redundant computing task that does not need to be calculated with the model.

7. The method according to claim 6, characterized in that The master node collects the results from the worker nodes and decodes the overall task results, including: The master node stores all the returned results from the worker nodes. When the number of results for a task reaches k, the master node decodes the overall calculation result of the task according to the properties of the MDS code and then sends the result to the user. For a task that receives k results, the master node will notify the rk worker nodes that have not completed the task to cancel the execution of the task. After receiving the notification, the worker node will find the task from the task queue and mark it.

8. The method according to claim 1, characterized in that The method further includes: adjusting the model encoding scheme and the task scheduling strategy in real time according to the task arrival rate, so as to make the execution of new and old tasks compatible.

9. The method according to claim 8, characterized in that The real-time adjustment of the model encoding scheme and task scheduling strategy based on the task arrival rate to ensure compatibility with the execution of new and old tasks includes: Set a threshold for the change in task arrival rate. When the change in task arrival rate exceeds the threshold, redesign the model encoding scheme and task scheduling strategy. Encode the task model using the new encoding scheme and save it to all working nodes; The master node uses the new scheduling policy to schedule newly arrived tasks to the worker nodes, and these tasks are calculated with the new model in the worker nodes. For the unfinished tasks in the queue, they are calculated with the old model in the worker nodes. The master node distinguishes between new and old task results and decodes the overall task result using different decoding methods; When all old tasks in the task queue of a working node are completed, the working node deletes the old task model data.

10. A distributed computing system based on MDS coding and flexible scheduling strategy, wherein the working cluster consists of a master node and multiple worker nodes. The master node is responsible for scheduling tasks and collecting and integrating task results, while the worker nodes are responsible for calculating subtask results and returning them to the master node. The system is characterized by: The task model in the working node is obtained by encoding according to a predetermined encoding scheme. When the master node receives the task, it schedules the task to the working node according to a predetermined scheduling strategy. The encoding scheme and scheduling strategy are obtained according to the following method: The average task computation time T in the system is obtained based on the computing power of a single working node and the task arrival rate in the system. The computing power includes the average time it takes for a working node to execute the entire task and the time required to cancel the task. The task arrival rate is the number of tasks arriving at the master node per unit time. The average task computation time T in the system is the time from when the task arrives at the master node to when the master node decodes the overall task result. The calculation formula is as follows: Where k is the encoding scheme, and its value represents the number of blocks of the model. r is the scheduling strategy, and its value represents the number of scheduled tasks to the working nodes. i is the harmonic series of order i, μ a is the average rate at which a single node processes tasks, ρ is the load of a single node, is a fixed coefficient; Iteratively search for a model encoding scheme and a task scheduling strategy that minimize the average computation time T of tasks in the system. The model encoding scheme is used to determine the number of blocks into which the task model is divided, and the task scheduling strategy is used to determine which work nodes the tasks are scheduled to execute.

Citation Information

Patent Citations

  • Cloud computing task scheduling method based on response time optimization

    CN103841208A

  • Distributed computing method based on priority coding

    CN111858721A