Task scheduling method, scheduling device, electronic equipment, storage medium and product
By building a global computing graph in a heterogeneous cluster, the mapping relationship between the operator and the accelerator is determined, and the problems of low accelerator resource utilization and poor task execution efficiency are solved, and efficient utilization and flexible scheduling of accelerator resources are achieved.
Patent Information
- Application Number
- CN202510892603.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-06-30
AI Technical Summary
In the prior art, the accelerator resource utilization rate of heterogeneous clusters is low, the task execution efficiency is poor, the task execution cannot be effectively coordinated, and the scheduling strategy cannot adapt to the asymmetry and computing needs of ML applications.
By establishing a mapping relationship between the operator and the accelerator in each task in a heterogeneous cluster, a global computing graph is constructed, a scheduling strategy is generated uniformly, the operator is sent to the corresponding accelerator for execution, and the task execution is performed in combination with the data dependency relationship.
It improves the accelerator resource utilization rate and task execution efficiency, enhances scheduling flexibility, adapts to the asymmetry and computing needs of ML applications, and reduces scheduling delay.
Smart Images

Figure CN120407127A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of task scheduling, and particularly to a task scheduling method, a scheduling device, an electronic device, a storage medium, and a product. Background Art
[0002] A heterogeneous cluster refers to a cluster system composed of devices with different architectures or operating systems. Specifically, a heterogeneous cluster contains multiple different types of nodes, which may run different operating systems (such as Linux and Windows) or use different hardware configurations. This diversity enables the heterogeneous cluster to better adapt to different workloads and requirements, improving the overall computing efficiency and flexibility.
[0003] A heterogeneous cluster includes multiple types of accelerators. During the task scheduling process, the scheduler allocates corresponding accelerator resources for a node task according to the clearly declared resource types and quantities received, for the execution of the task. However, this task scheduling method reduces the utilization rate of accelerator resources and is not conducive to the execution efficiency of tasks. Summary of the Invention
[0004] This application provides a task scheduling method, a scheduling device, an electronic device, a storage medium, and a product to at least solve the problem of low utilization rate of accelerator resources in related technologies.
[0005] This application provides a task scheduling method applied to a heterogeneous cluster, which includes multiple accelerators. The method includes: receiving at least one task to be executed, where each task includes multiple operators and data dependency relationships between the multiple operators; determining a first mapping relationship between the multiple operators in each task and the multiple accelerators, and establishing a global computation graph based on the first mapping relationship corresponding to each task; sending each operator to the corresponding accelerator according to the global computation graph, and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators.
[0006] This application also provides a task scheduling device applied to a heterogeneous cluster. The device includes: a receiving module, configured to receive at least one task to be executed, where each task includes multiple operators and data dependency relationships between the multiple operators; a building module, configured to determine a first mapping relationship between the multiple operators in each task and the multiple accelerators, and establish a global computation graph based on the first mapping relationship corresponding to each task; a scheduling module, configured to send each operator to the corresponding accelerator according to the global computation graph, and execute at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators.
[0007] The present application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above task scheduling methods when executing the computer program.
[0008] The present application also provides a non-volatile computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above task scheduling methods when executed by a processor.
[0009] The present application also provides a computer program product including a computer program, which implements the steps of any of the above task scheduling methods when executed by a processor.
[0010] Through the present application, by establishing the first mapping relationship between each operator in the task and the accelerator to determine the global computation graph, uniformly generating a scheduling policy, sending the operator to the corresponding accelerator during the task execution process, and completing the assigned operator based on the operator library of each accelerator, thereby executing the task, the type matching limitation between the task and the accelerator resources is broken, and the technical problems of low utilization rate of accelerator resources and poor task execution efficiency in the related art can be solved, achieving the technical effects of improving the utilization rate of accelerator resources and scheduling flexibility and enhancing the task execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 It is a flowchart of a task scheduling method provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the architecture of the task scheduling method in a specific embodiment of the present application; Figure 3 It is a schematic diagram of the mapping from an operator to an accelerator in a specific embodiment of the present application; Figure 4 It is a schematic diagram of the task scheduling method in a specific embodiment of the present application; Figure 5 It is a schematic diagram of the generation of a scheduling policy in a specific embodiment of the present application; Figure 6 It is a schematic diagram of fault tolerance processing in a specific embodiment of the present application; Figure 7 It is a flowchart of the task scheduling method in a specific embodiment of the present application; Figure 8Execution schematic diagram of the global computational graph of a specific embodiment of the present application; Figure 9 Connection schematic diagram of a task scheduling device provided by an embodiment of the present application; Figure 10 Block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0014] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0016] In scenarios such as large companies, multi-tenant clusters are often constructed for multiple users to share, aiming to centralize computing resources and improve resource utilization. In a multi-tenant cluster, each computing node may use different operating systems and different hardware configurations to meet different workloads and requirements, which is a heterogeneous system.
[0017] In addition, with the progress of machine learning (ML) technologies (such as mixture of experts (MoE), multi-modal models), the requirements of ML applications for storage and computing resources have a strong asymmetry, which makes the multi-tenant cluster system more heterogeneous. Among them, the asymmetry of ML applications is reflected in many aspects. For example, due to the computational sparsity of the MoE model structure, each inference may involve different control flows, resulting in different accelerators being busy; in a multi-modal model, some inferences involve text, while others involve video, making the granularity of inferences vary greatly; due to the characteristics of the business itself, the load fluctuates greatly at different times.
[0018] In addition, with the emergence of new accelerators such as NPU (Neural network Processing Unit) and TPU (Tensor Processing Unit), the introduction of interconnect technologies such as RDMA (Remote Direct Memory Access) and NVLink (NVIDIA NVLink, a bus and its communication protocol developed and launched by NVIDIA), and operation and maintenance factors such as hardware upgrades, the multi-tenant cluster includes various models of accelerators, making the cluster system more heterogeneous.
[0019] In the related art, the job scheduler allocates accelerator resources according to the received task declaration, and executes the task through the accelerator resources until the task is completed, and the accelerator resources are released for the execution of the next task. However, due to the asymmetry of ML applications, allocating a fixed number of accelerator resources for ML applications will result in low utilization of accelerators in the entire cluster. For example, a single task may not be able to fully utilize the accelerator resources allocated to it, so multiple tasks need to share the accelerator to improve resource utilization; for another example, when the business load intensity decreases, the computing resources allocated to the current task cannot be reallocated to other tasks, resulting in idle waste of resources. Therefore, the task scheduling method adopted in the related art cannot meet the asymmetry and changing computing requirements of ML applications.
[0020] During the operation of the job scheduler, the user submits the containerized application to the container orchestration tool (such as Kubernetes, which is a portable and extensible open-source platform mainly used to manage containerized workloads and services, supporting declarative configuration and automation) for deployment. Therefore, the job scheduler usually has three main limitations: First, the job scheduler generally uses the gang scheduling strategy, which means that the resources declared by the task will be allocated all at once only when all these resources are idle. These resources are then occupied by the task until the task is completed, and then released for other tasks to use; Second, the job scheduler hands over the accelerator to the task for use, which means that how to utilize the accelerator cluster resources depends on the implementation of the task itself; Third, the scheduler does not understand the internal details of the task.
[0021] Specifically, the job scheduler allocates accelerator resources according to the explicitly declared resource types and quantities of the received jobs, which determines that the accelerator resources allocated to jobs cannot be easily changed during task scheduling. The parallel strategies of most current ML applications are based on 3D parallelism, and 3D parallelism only supports homogeneous resources, which means that there is no highly heterogeneous scheduling strategy within the scheduling space of a single job. Therefore, in existing practices, it is not feasible to allocate heterogeneous accelerator resources to tasks.
[0022] In addition, the job scheduler lacks an understanding of the inside of the tasks, and it is difficult to schedule tasks onto heterogeneous accelerator resources. For example, when multiple tasks share a GPU, due to the lack of understanding of the details of the tasks, the scheduler cannot manage the computational and network interference between tasks, which may affect high-priority tasks and may even cause an OOM (Out Of Memory) error.
[0023] Based on the above analysis, it can be seen that the task scheduling methods in the related technologies have the following problems: (a) It is impossible to enable highly heterogeneous accelerators to execute tasks collaboratively.
[0024] (b) It is impossible to achieve fast elasticity, that is, to quickly adjust the accelerator resources allocated to a single task.
[0025] (c) It is impossible to estimate the feasibility and superiority of the scheduling strategy on heterogeneous resources.
[0026] (d) It is impossible to estimate the network and computational interference between tasks.
[0027] To solve at least one of the above technical problems, the present application proposes a task scheduling method, which is applied to a heterogeneous group including multiple accelerators. When receiving at least one task to be executed, a first mapping relationship between the operators in each task and the accelerators is established, thereby constructing a global computational graph. The mapping relationship is displayed in a graphical manner through the global computational graph, and a unified scheduling strategy is generated. During the task processing, each operator is separately sent to the corresponding accelerator based on the global computational graph, so as to execute the corresponding operator calculation through the corresponding accelerator, and complete the data transmission between the accelerators in combination with the data dependency relationship between multiple operators, and complete the task execution. Thus, this method comprehensively depicts the global state of all tasks and accelerator resource usage in the heterogeneous cluster through the Global Computational Graph (GCG), integrates the operators in all tasks submitted by users to construct a unified large-scale directed computational graph, and can realize global scheduling and resource management across tasks, thereby improving the resource utilization efficiency of accelerators in a multi-tenant cluster, reducing scheduling latency, and enhancing the running elasticity and system performance of tasks such as Graph Neural Network (GNN) and Large Language Model (LLM).
[0028] The task scheduling method of the embodiment of the present application will be described in detail below with reference to the accompanying drawings. This task scheduling method is executed by a scheduler in a heterogeneous cluster.
[0029] As Figure 1 shown, the task scheduling method of the embodiment of the present application, which is applied to a heterogeneous cluster, may include: S1. Receive at least one task to be executed, where each task includes multiple operators and the data dependency relationship between multiple operators; Specifically, as Figure 2 shown, the cluster watchdog obtains the online / offline status of each accelerator in the heterogeneous cluster and sends it to the scheduler. The scheduler identifies multiple accelerators in the online state based on the status of each accelerator in the system, so as to perform subsequent task scheduling applications based on the multiple accelerators. At the same time, the scheduler receives at least one task to be executed (such as Figure 2 Task 1, Task 2, Task M, etc. in the figure). Each task is embodied in the form of a data stream and includes multiple operators and the data dependency relationship between multiple operators. The operators in each task can be divided according to the operation symbols, and the data dependency relationship between the operators can be determined according to the calculation logic of the task.
[0030] S2. Determine the first mapping relationship between the multiple operators in each task and the multiple accelerators, and establish a global computational graph based on the first mapping relationship corresponding to each task; Specifically, each accelerator has its own operator library, which can meet the computing requirements of any operator. During the establishment of the first mapping relationship, the operators can be evenly distributed according to the number of operators in each task and the number of accelerators, so as to establish the first mapping relationship corresponding to each task; or a preset matching relationship between operators and accelerators can be determined in advance, and then after receiving a task, the mapping relationship between each operator in the task and the accelerator can be determined by looking up a table to obtain the corresponding first mapping relationship; or a distribution strategy can be set in advance, and the matching between operators and accelerators can be performed based on the preset distribution strategy, so as to establish the corresponding first mapping relationship. For example, according to the computing logic of the operators in each computing subtask, the accelerators can be matched in sequence from front to back to establish the mapping relationship between each operator and the accelerator. In Figure 3 it, the mapping relationship between each operator and the corresponding accelerator is represented by a dotted arrow.
[0031] After determining the first mapping relationship between multiple operators and multiple accelerators in each task, a global computation graph is established according to the first mapping relationship corresponding to each task, and the mapping of the operators of all tasks to the accelerators is displayed through the global computation graph to describe the mapping situation of all tasks. Thus, through the global computation graph, the computing subtasks in all tasks are merged into a large computation graph, as Figure 3 shown, so as to uniformly generate a scheduling strategy.
[0032] Furthermore, each task can be divided into multiple computing subtasks, and the first mapping relationship is constructed in units of computing subtasks. Through task module division, the construction efficiency of the global computation graph is improved. Combining Figure 3 as shown, Task 1 includes two computing subtasks, and there is a data dependency relationship between the two computing subtasks, that is, the computing result of one computing subtask is used for the computing of the other computing subtask. Then when the scheduler receives Task 1, it can identify the two computing subtasks in Task 1 and the data dependency relationship between the two computing subtasks, as well as the data dependency relationship between multiple operators in each computing subtask, and respectively determine the first mapping relationship between multiple operators and multiple accelerators in the two computing subtasks, improving the construction speed of the global computation graph. Task 2 only includes one computing subtask. Then when the scheduler receives Task 2, it can identify one computing subtask in Task 2, as well as the data dependency relationship between multiple operators in the computing subtask, and establish the first mapping relationship between multiple operators and multiple accelerators in Task 2.
[0033] S3. Send each operator to the corresponding accelerator according to the global computation graph, and execute at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between multiple operators.
[0034] That is to say, based on the first mapping relationship corresponding to each task in the global computation graph, each operator is separately sent to the corresponding accelerator, and thus added to the execution queue of the corresponding accelerator, so that the accelerator can execute the received operator computation tasks in sequence according to the queue order or the preset priority through its own operator library, obtain the operands of each operator in combination with the data dependency relationship, and execute the corresponding operator computation when the operands of the operator are all calculated.
[0035] Combined with Figure 4 As shown, a task includes seven operators respectively represented as Among them, the first mapping relationship between this computation subtask and multiple accelerators is as shown in the figure, that is, operator is assigned to accelerator a, operators , and are assigned to accelerator b, operators and are assigned to accelerator c, and operator is assigned to accelerator d. During the task execution process, based on the global computation graph, operator is sent to the queue of accelerator a, operators , and are sent to the queue of accelerator b, operators and are sent to the queue of accelerator c, and operator is sent to the queue of accelerator d, and operator computations are performed based on the corresponding queues in sequence. Based on the data dependency relationship, it can be known that the computation result of operator is the operand of operator . Therefore, after accelerator a completes the computation of operator based on its own operator library, it sends the computation result of operator to accelerator c for the computation of operator . In addition, the operands of operator include not only the computation result of operator but also the computation result of operator . Therefore, operator starts to compute after receiving the computation results of operator and operator . And so on until all operator computations are completed.
[0036] This embodiment establishes a mapping from operators to accelerators, which means that the accelerators in the heterogeneous cluster expose a unified operator execution interface to the global computation graph, making it possible for various heterogeneous accelerators to cooperate in task execution. Based on this embodiment, it is not necessary to compile the kernel function into binary code that can run on all accelerators. Instead, the assigned operator calculations are completed based on the operator library available for each accelerator. That is, when an operator is scheduled to the corresponding accelerator, the accelerator will run its own operator library to complete the operator calculation.
[0037] Thus, by establishing a unified mapping mechanism from operators to accelerators, this embodiment supports heterogeneous accelerators (such as GPUs, FPGAs (Field Programmable Gate Arrays), and TPUs) in exposing a unified operator execution interface, rather than forcing all operators to adapt to a common kernel code. This method allows heterogeneous accelerators to use their respective native operator libraries to execute tasks, eliminating the performance overhead of cross-platform kernel conversion, thereby enabling multiple heterogeneous devices to participate in the parallel computing of tasks collaboratively, effectively improving the overall resource utilization rate and scheduling flexibility. At the same time, when multiple tasks are received, the accelerator types and quantities allocated to each task can be quickly adjusted through operator matching, improving the task scheduling flexibility.
[0038] In some embodiments of the present application, receiving at least one task to be executed includes: receiving a directed acyclic graph of at least one task, where the nodes of the directed acyclic graph are used to represent the corresponding operators, and the edges of the directed acyclic graph are used to represent the data dependency relationships between the operators.
[0039] Specifically, a directed acyclic graph (DAG) is a data structure defined in graph theory, consisting of vertices and edges. Each edge has a clear direction, and the entire graph is acyclic, that is, there is no path in the graph that can start from a point, pass through a series of edges, and then return to that point, as shown in Figure 2 and Figure 3 shown. Receiving the directed acyclic graph of the task improves the construction efficiency of the first mapping relationship and the task scheduling accuracy.
[0040] In addition to the mapping relationship between operators and accelerators, the global computation graph may also include the directed acyclic graph of each received task, as shown in Figure 3 shown. Thus, the directed acyclic graph of each task and the mapping of each operator to the accelerator are displayed through the global computation graph to describe the mapping situation of all tasks. Then, the computational subtasks in all tasks are merged into a large computation graph through the global computation graph, and a scheduling strategy is generated uniformly for subsequent task scheduling applications.
[0041] In some embodiments of the present application, determining a first mapping relationship between multiple operators in each task and multiple accelerators includes: determining multiple task scheduling policies corresponding to the task based on a preset scheduling generation method; determining the scheduling performance of each task scheduling policy according to the preset calculation duration of each operator in the corresponding accelerator and the preset communication duration between multiple accelerators; determining a target task scheduling policy according to the scheduling performance of each task scheduling policy, and generating a first mapping relationship between multiple operators in the task and multiple accelerators based on the target task scheduling policy.
[0042] Specifically, after sending at least one task to the scheduler, first generate several task scheduling policies using different types and quantities of accelerator resources based on a preset scheduling generation method. Each task scheduling policy includes the accelerator allocation corresponding to each operator and the data transfer scheduling between accelerators. For Figure 4 example, the task includes seven operators, respectively represented by . Figure 4 The corresponding task scheduling policy is: calculate operator through accelerator a, and send the calculation result of operator to accelerator d; calculate operators , , and through accelerator b. First, calculate operator , and send the calculation result of operator to accelerator d. At the same time, calculate operator according to the result of operator , and send the calculation result of operator to accelerator c, and save the result of operator ; calculate operator through accelerator d according to the calculation results of operators and , and send the calculation result of operator to accelerator b; accelerator b calculates operator according to the results of operators and ; accelerator c calculates operators and respectively according to the calculation result of operator , thus completing the calculation of the task.
[0043] Then, determine the total execution duration of each task scheduling policy according to the preset calculation duration of the operator in the corresponding accelerator and the preset communication duration between multiple accelerators in each task scheduling policy, so as to determine the scheduling performance of each task scheduling policy, and evaluate the feasibility and execution stability of each task scheduling policy.
[0044] Among them, the pre-designed calculation duration of the operator in the corresponding accelerator can be determined based on historical calculation data or can be a preset duration, without specific limitation. Taking Figure 4 the operator in assigned to accelerator a as an example, the historical calculation duration of the operator in accelerator a can be found through historical execution data, and the mean value of the historical calculation duration is calculated as the pre-designed calculation duration of the operator in accelerator a; if there is no situation where accelerator a calculates the operator in the historical data, at this time, the historical calculation duration of the operator can be matched according to the type of accelerator a with that of a same-type accelerator as the pre-designed calculation duration of the operator in accelerator a; or a mapping table between the operator, accelerator type, and calculation duration can be established in advance, and the corresponding preset execution duration can be determined by looking up the table during the actual application process.
[0045] The preset communication duration between multiple accelerators includes the preset communication duration between each group of accelerators, which can be determined through historical data, can be a preset duration, or can be estimated based on the bandwidth between each group of accelerators, without specific limitation. Taking Figure 4 accelerator a in needing to send the calculation result of the operator to accelerator d as an example, the corresponding preset communication duration between the accelerators includes the historical data transmission duration when accelerator a sends the calculation result of the operator to accelerator d. The historical communication duration when accelerator a sends the calculation result of the operator to accelerator d can be queried through historical data, and the corresponding preset communication duration is obtained by taking the mean value; or it can be predicted based on the bandwidth between accelerator a and accelerator d and the estimated data size of the operator to obtain the corresponding preset communication duration; or a communication duration mapping table between each group of accelerators can be established in advance, and the corresponding preset communication duration can be determined by looking up the table.
[0046] Taking the scheduling performance of the task scheduling strategy as the total execution duration as an example, the pre-designed calculation duration and preset communication duration are selected according to each task scheduling strategy, and the total execution duration is calculated in combination with the scheduling strategy, so as to determine the scheduling performance of each task scheduling strategy. The task scheduling strategy with the shortest pre-execution total duration is selected as the target task scheduling strategy for this task, and thus the first mapping relationship corresponding to the task is determined according to the target task scheduling strategy, which is used to construct the global computational graph.
[0047] This embodiment improves the task scheduling performance by evaluating the performance of multiple task scheduling strategies, ensuring the construction quality of the global computation graph and the task scheduling effect.
[0048] Furthermore, in addition to evaluating each task scheduling strategy separately after generating multiple task scheduling strategies corresponding to each task to determine the target scheduling strategy, synchronous screening can also be performed during the scheduling strategy generation process to obtain the final target scheduling strategy. For example, Figure 5 as shown, after receiving a task, according to the topological order of the directed acyclic graph of the task, first map the first node of the directed acyclic graph, that is, the operator, to an accelerator as the root node of the entire scheduling decision tree; on this basis, map the second node of the directed acyclic graph to different accelerators to become the first-layer nodes of the scheduling decision tree. During the process of generating nodes, call a predictor to predict the quality of this scheduling decision; if it is found during the generation of the decision tree that a certain accelerator cannot execute the operator or the execution effect is not good, then abandon all subsequent scheduling decisions of this node to form pruning. And so on to generate the final scheduling decision tree as the target scheduling strategy.
[0049] In some embodiments of the present application, when multiple accelerators are in the task execution state, it further includes: determining the current global computation graph that the multiple accelerators are executing; determining the idle start time of each accelerator according to the current execution state of the current global computation graph, the pre-designed computation duration of each operator in the current global computation graph in the corresponding accelerator, and the pre-set communication duration between the multiple accelerators; predicting the execution end time of each task scheduling strategy according to the pre-designed computation duration of each operator in the task in the corresponding accelerator, the pre-set communication duration between the multiple accelerators, and the idle start time of each accelerator to determine the scheduling performance of each task scheduling strategy.
[0050] Specifically, when multiple accelerators in the heterogeneous cluster are in the idle state, if a task to be executed is received, the multiple accelerators can perform calculation operations at the first time of receiving the operator. Therefore, the scheduling performance can be evaluated only according to the total execution duration of each task scheduling strategy, and the task scheduling strategy with the shortest total execution duration is used as the target task scheduling strategy to construct the global computation graph.
[0051] However, when multiple accelerators in the heterogeneous cluster are in the task execution state, to further improve the performance evaluation accuracy, the idle start time of each accelerator is predicted.
[0052] Combined with Figure 2As shown, the execution time of each operator in the current global computation graph on the corresponding accelerator can be predicted based on a prediction function (such as a Kernel function). Specifically, each operator is executed one by one on the corresponding accelerator in a preset order. At the same time, based on a preset cluster transmission management algorithm, the execution order of all operators on the accelerator can be inferred, and the time when each operator finishes execution can be predicted. Moreover, the bandwidth between the two accelerators is also predictable, so as to predict the communication transmission duration between the two accelerators. In the current global computation graph, the real-time mapping from operators to accelerators can be reflected to infer the ready operators in the out-of-order queue on each accelerator. And based on the execution priority of the operators, it can be inferred which ready operators are being executed. A ready operator is an operator that can enter the execution state. For example, in Figure 4 there are no operands for operator and operator . Therefore, when operator and operator are added to the queue of the corresponding accelerator, they are ready operators. While for other operators, there are operands. When the operands of the corresponding operators are all calculated, the corresponding operators switch to the ready state and become ready operators. Since each operator on the accelerator is executed one by one, and due to the predictability of the operators and transmissions, the execution completion time of the currently running operators on the accelerator can be accurately predicted. Thus, it can be determined which operator will finish execution next, and then it is known which operator will be made ready by this completed operator. By predicting the execution of operators one by one in this way, the execution completion time of all operators can be predicted iteratively, and thus the idle start time of each accelerator can be predicted.
[0053] In the process of obtaining the scheduling performance of each task scheduling strategy, the calculation end time of each task scheduling strategy is predicted based on the preset calculation duration of each operator in the corresponding accelerator, the preset communication duration between multiple accelerators, and the idle start time of each accelerator, so as to evaluate the scheduling performance of each task scheduling strategy. Then, the task scheduling strategy with the earliest calculation end time is selected as the target scheduling strategy, and the global computation graph is constructed accordingly.
[0054] The above-mentioned preset calculation duration of the operator in the corresponding accelerator and the predicted communication duration between multiple accelerators can be predicted through historical data. Since when the task is executed for the first time, the above information of the task is lacking, the cluster prediction is particularly inaccurate. However, as the task is executed iteratively time after time, the prediction accuracy will gradually increase, thus improving the scheduling quality. Additionally, in the absence of historical data, it can also be determined by looking up a preset table.
[0055] When multiple accelerators receive a new task during the execution of tasks, this embodiment evaluates each task scheduling policy in combination with the predicted idle start time of each accelerator, so as to exclude task scheduling decisions with particularly late operator completion times. As a result, the target task scheduling decision schedules task computing to the idle computing power area in the cluster, that is, the accelerator with an earlier idle start time, which not only improves the utilization rate of the cluster's computing power but also ensures that the execution time of task computing is not too late, further improving the quality of task scheduling.
[0056] The generation process of this target scheduling policy weakens the specific model of resources and can generate scheduling decisions for heterogeneous resources to perform collaborative computing. The scheduling decision tree generated based on the pre-scheduling algorithm will extend in any direction to the accelerator cluster due to the scale of the directed acyclic graph, so it can be elastic. In addition, the prediction unit in the scheduler will consider both the execution time of the operator in the corresponding accelerator and the tensor transmission time, that is, the transmission time of the calculation result between accelerators, during the prediction process, so as to evaluate the feasibility and quality of the scheduling decision and improve the final task scheduling quality.
[0057] In some embodiments of the present application, determining the idle start time of each accelerator according to the current execution state of the current global computation graph, the pre-designed computation duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between multiple accelerators includes: determining the execution completion time of each operator in the current global computation graph according to the current execution state of the current global computation graph, the preset communication duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between multiple accelerators; predicting the idle start time of each accelerator according to the execution completion time of each operator in the current global computation graph.
[0058] Specifically, assume that when a new task is received, multiple accelerators are executing Figure 4 the global computation graph shown, where the operators and have been executed, the operators and are being executed, and the operators , and are waiting to be executed. Determine the pre-designed computation duration T1 of the operator in accelerator a, the pre-designed computation duration T5 of the operator in accelerator c, the pre-designed computation duration T2 of the operator in accelerator d, the pre-designed computation duration T6 of the operator in accelerator c, the pre-designed computation duration T7 of the operator [[ID=--34]] in accelerator b, and the transmission of the operator from accelerator b to accelerator d based on historical data or a preset mapping tableThe preset transmission duration T of the calculation result bd Accelerator a transmits an operator to accelerator d The preset transmission duration T of the calculation result ad Accelerator d transmits an operator to accelerator b The preset transmission duration T of the calculation result bd .
[0059] Since the operators and have been executed, and the calculation result of the operator has been transmitted to accelerator c, the calculation end time t3 of the operator , the calculation end time t4 of the operator , and the time t50 when accelerator c receives the calculation result of the operator can be determined according to the actual feedback. At the same time, the calculation start time t10 of the operator is determined. Then, the calculation end time t1 of the operator is t1 = t10 + T1, and the idle start time of accelerator a is determined to be t1; the calculation end time t2 of the operator is t2 = t3 + T bd + T2, and the idle start time of accelerator d is determined to be t2; the calculation end time t5 of the operator is t5 = t50 + T5, and the calculation end time t6 of the operator is t6 = t5 + T6, and the idle start time of accelerator c is determined to be t6; the calculation end time t7 of the operator is t7 = t2 + T bd + T7, and the idle start time of accelerator b is determined to be t7.
[0060] This embodiment determines the calculation end time of each operator based on the preset calculation duration of the operator in the corresponding accelerator and the preset communication duration between multiple accelerators, so as to predict the idle start time of each accelerator. The preset calculation duration and the preset communication duration can be predicted through historical data, and as the task is iteratively executed, the prediction accuracy of the preset calculation duration and the preset communication duration is improved to improve the prediction accuracy of the idle start time of each accelerator, thereby improving the scheduling quality.
[0061] In some embodiments of the present application, the global calculation graph includes the state information of each operator in at least one task. In the process of separately sending each operator to the corresponding accelerator according to the global calculation graph and executing at least one task through one or more of multiple accelerators in combination with the data dependency relationship between multiple operators, it further includes: receiving operator state feedback information sent by one or more of multiple accelerators; updating the state information of each operator according to the operator state feedback information.
[0062] Specifically, the global computation graph is a data structure, which can be defined as a five-tuple, namely the global computation graph . Among them, is DAG (Directed Acyclic Graph , a directed acyclic graph) graph, that is, the data stream corresponding to the received task, is the node set of the DAG graph, used to represent a node, that is, an operator; is the edge set between nodes, used to represent the data dependency relationship between node and node . , representing the set of accelerators in the heterogeneous cluster, used to represent all online accelerators in the heterogeneous cluster. is the mapping from the operator (node) to the accelerator, representing that the operator is executed on a certain accelerator. is the current execution state of the operator (node), representing whether the operator has been executed completely at present. For example, at , it means that the operator has not been executed completely; at , it means that the operator has been executed completely.
[0063] After sending the operator to the corresponding accelerator according to the global computation graph, the state of each operator is updated to , the operator is calculated by the corresponding accelerator, and the calculation result is fed back in real time. When the state feedback information of the corresponding operator fed back by the accelerator is calculated to be completed, the state of the corresponding operator is updated to . In addition, the state can also be distinguished by the gray scale of the operator in the global computation graph, as shown in Figure 4 .
[0064] This embodiment updates the state of each operator in real time according to the operator state feedback information and displays it, which is convenient for tracking the task execution state in the system.
[0065] In some embodiments of the present application, updating the state information of each operator according to the operator state feedback information includes: when it is determined according to the operator state feedback information that the operator is in the execution completed state, deleting the operator and the associated information of the operator in the global computation graph.
[0066] That is to say, when it is determined that the operator has been executed completely, the operator is deleted from the global computation graph, and at the same time, the associated information of the operator in the global computation graph is deleted, such as the edges in the DAG graph and the mapping relationship between the operator and the accelerator, etc., to reduce memory occupancy.
[0067] In some embodiments of the present application, in the process of sending each operator to the corresponding accelerator according to the global computation graph, and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between multiple operators, the task scheduling method also includes: in response to the operator insertion instruction, identifying the target operator corresponding to the operator insertion instruction and the data dependency relationship between the target operator and the operators in the global computation graph; establishing a second mapping relationship between the target operator and the multiple accelerators based on the data dependency relationship between the target operator and the operators in the global computation graph, and updating the global computation graph based on the second mapping relationship.
[0068] Specifically, the operator insertion instruction may be triggered when a new task is received or other computing subtasks in a task are completed.
[0069] Assume that when it is determined that a target operator needs to be added to the global computation graph based on the operator insertion instruction, first insert the instruction based on the operator, and the node set in the global computation graph Add a node corresponding to the target operator ,Right now At the same time, determine the data dependency between the target operator and the operators in the global computation graph, and determine its predecessor node as , and to Add an edge about the node in A second mapping relationship between the target operator and the multiple accelerators is established based on the data dependency relationship between the target operator and the operators in the global computation graph, thereby updating the global computation graph so as to complete the calculation of the newly added target operator based on the updated global computation graph.
[0070] This embodiment responds to the operator addition instruction to update the global computation graph, thereby performing computation on the newly added operator based on the updated global computation graph to meet computational requirements.
[0071] In some embodiments of the present application, the global computation graph includes an accelerator set corresponding to multiple accelerators. In the process of sending each operator to the corresponding accelerator according to the global computation graph, and executing at least one task through one or more of the multiple accelerators in combination with the data dependencies between the multiple operators, it also includes: obtaining status information of each accelerator in the heterogeneous cluster; and updating the accelerator set according to the status information of each accelerator.
[0072] Specifically, combined Figure 2 As shown, the cluster watchdog monitors the accelerator status in the cluster system in real time, so as to (Signature_Accelerator_Online) function, (Signature_Accelerator_Offline) Informs the scheduler of the accelerator status information in the entire system, so as to update the accelerator set in the global computation graph in real time based on the accelerator status information, and accurately establish subsequent mapping relationships based on the updated accelerator set.
[0073] In some embodiments of the present application, updating the accelerator set according to the status information of each accelerator includes: when it is determined according to the status information of each accelerator that there is a newly online accelerator in the heterogeneous cluster, adding the newly online accelerator to the accelerator set; and / or when it is determined according to the status information of each accelerator that there is an offline accelerator in the accelerator set, re - establishing the third mapping relationship between the uncompleted operators corresponding to the offline accelerator and the online accelerators in the accelerator set, and then deleting the offline accelerator in the accelerator set.
[0074] That is to say, when the cluster watchdog informs the scheduler that there is a new accelerator online through the (Signature_Accelerator_Online) function, directly add the newly online accelerator to the accelerator set for establishing the mapping relationship for subsequent task scheduling. That is , making .
[0075] When the cluster watchdog passes through (Signature_Accelerator_Offline) function to inform the scheduler that an accelerator in the global computation graph goes offline, first cancel the operators in the offline accelerator, re - establish the mapping relationship between the uncompleted operators corresponding to the offline accelerator and other online accelerators in the accelerator set, then delete the offline accelerator from the accelerator set, and send the uncompleted operators corresponding to the offline accelerator to the corresponding online accelerators based on the reconstructed mapping relationship to continue operator calculation. That is, when receiving , first cancel the operators related to the offline accelerator , that is (The practical significance is to reschedule the operators related to this and build the mapping relationship with other online accelerators); then make .
[0076] This embodiment directly adds the newly online accelerator to the accelerator set in the global computation graph for establishing the mapping relationship of subsequent tasks. When there is an offline accelerator, it adopts a fault - tolerant task calculation strategy, reconstructs the operators corresponding to the offline accelerator and then deletes them. Only through local mapping and adjustment of the scheduling relationship, the accuracy of task scheduling is guaranteed.
[0077] In some embodiments of the present application, when it is determined that there is an offline accelerator in the accelerator set according to the current state of each accelerator, it further includes: obtaining the intermediate operator calculation results stored in the offline accelerator, and sending the intermediate operator calculation results to the corresponding online accelerator based on the data dependency relationship between multiple operators.
[0078] Specifically, each accelerator is equipped with an out-of-order execution queue. When the scheduler launches an operator to a certain accelerator, the operator will be added to the out-of-order execution queue of the corresponding accelerator. Specifically, an enqueue_op interface can be defined in each accelerator, and the scheduler can launch an operator to the accelerator by calling the enqueue_op interface.
[0079] During the operation of the accelerator, each accelerator constantly monitors its own queue and executes the ready operators one by one. When an operator is first added to the queue, all its operands are ; when all the operation data of an operator are calculated, that is, the operands are , then the operator is in the state. When an operator is executed, its calculation result will be broadcast to the operands of the corresponding operators in the entire queue that need this value, and then that operand becomes state. At the same time, the calculation results of the corresponding operators are stored in the HBM (High Bandwidth Memory) of the accelerator.
[0080] When it is determined that the accelerator is in an offline state, the scheduler reads the intermediate operator calculation results stored in the HBM of the offline accelerator and forwards them to the corresponding accelerator for subsequent operator calculations. Taking Figure 4 as an example, assume that accelerator b goes offline. At this time, the scheduler can read the calculation results of the operator stored in accelerator b and send them to accelerator d for the calculation of operator , and at the same time forward the calculation results of operator to accelerator c for the calculations of operator and .
[0081] This embodiment defines a fault tolerance mechanism. As shown in Figure 6 , the intermediate results retained in the HBM of the accelerator are used to restore the calculation progress, and there is no need to load the intermediate calculation results from remote storage, thereby accelerating the fault tolerance recovery speed. At the same time, this fault tolerance recovery mechanism only withdraws the affected operator calculations and accelerators, which only affects the local part of the accelerator and does not affect the overall cluster. Thus, the stability of the cluster computing power is improved, and the accelerator will not be stagnant for too long due to fault tolerance recovery, improving the computing power utilization rate.
[0082] In some embodiments of the present application, in the process of separately sending each operator to the corresponding accelerator according to the global computation graph and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between the multiple operators, the following is further included: in the case of an operator computation error, moving the operator with the computation error out of the original accelerator, establishing a fourth mapping relationship between the operator with the computation error and other accelerators among the multiple accelerators, and updating the global computation graph based on the fourth mapping relationship.
[0083] Specifically, in combination with Figure 6 As shown, during the operator computation process, the computation result of the operator is always retained in the HBM of the corresponding accelerator. When it is determined that there is an operator computation error, the scheduler withdraws the operator with the computation error from the original accelerator, and re - establishes the fourth mapping relationship between the operator with the computation error and other accelerators among the multiple accelerators to update the global computation graph, so as to re - allocate the operator with the computation error for recalculation of the operator with the computation error.
[0084] In this embodiment, when a computation error occurs, it triggers fault - tolerance processing and withdraws the affected operator, which only affects a local part of the cluster and does not affect the overall cluster, improving the stability of the cluster computing power and the computing power utilization rate.
[0085] In some embodiments of the present application, after updating the global computation graph based on the fourth mapping relationship, the following is further included: obtaining the operands of the operator with the computation error, where the operands are stored in the corresponding accelerator; sending the operator with the computation error and the operands to the corresponding accelerator based on the updated global computation graph.
[0086] That is to say, after moving the operator with the computation error out of the original accelerator, re - establish the fourth mapping relationship between this operator and other accelerators, as well as the computation scheduling path related to this operator, so as to update the global computation graph and perform task computation according to the global computation graph.
[0087] During the operator computation process, the accelerator always retains the computation result of the operator in the HBM of the corresponding accelerator. When it is determined that there is an operator computation error, the cluster watchdog notifies the scheduler through the sign_acc_out function. The scheduler withdraws the operator with the computation error from the original accelerator, reads the operands of this operator from the original accelerator or the accelerator used to compute this operator with the computation error, and sends the operator with the computation error and the operands to the corresponding accelerator for recalculation of this operator.
[0088] This embodiment only withdraws the affected operators and establishes the mapping relationship between the operators and other accelerators, which only affects the local part of the cluster. In a complex cluster with a low mean time between failures, this fault tolerance mechanism can improve the stability of the cluster computing power. At the same time, it performs calculation recovery based on the operands stored inside the corresponding accelerator, so that the accelerator will not be stalled for too long due to fault tolerance recovery, improving the computing power utilization rate.
[0089] In some embodiments of the present application, the task scheduling method further includes: obtaining the actual calculation duration of the operator, and determining that the operator has a calculation error when the actual calculation duration of the operator exceeds a preset duration; or determining that the operator has a calculation error when the accelerator where the operator is located has an abnormality; or in response to an operator exception instruction, determining that the operator corresponding to the operator exception instruction has a calculation error.
[0090] That is to say, when determining the actual calculation duration of the operator based on the operator status information fed back by the accelerator, when the actual calculation duration of the operator exceeds the preset duration, it is considered that the accelerator is not suitable for the calculation of the operator and a calculation error occurs. Or when it is determined according to the status information of the accelerator that the accelerator gives an offline or fault feedback, it is considered that the accelerator cannot complete the calculation of the internal operator, and it is determined that all the operators corresponding to the accelerator have calculation errors and they are all taken out for re-mapping calculation. Or when receiving an externally triggered operator exception instruction, it is considered that the operator corresponding to the operator exception instruction has a calculation error and re-allocate the calculation, which ensures the recognition effect of calculation errors and improves the task calculation quality.
[0091] In some embodiments of the present application, the task scheduling method further includes: identifying the communication willingness between at least one group of accelerators; generating a corresponding communication permission based on the communication willingness between at least one group of accelerators, so that at least one group of accelerators can communicate based on the communication permission.
[0092] When all tasks execute complex scheduling strategies in the cluster, if communication operators are allowed to be triggered freely, it may lead to an uncontrollable cluster pattern. Therefore, it is necessary to handle the tensor transmission between accelerators in the management of operator granularity.
[0093] To this end, in this embodiment, after a tensor (the result of an operator's calculation) is calculated, the accelerator first notifies the scheduler that the tensor is ready to be transmitted. The scheduler recognizes the accelerators' willingness to communicate, then generates a corresponding communication permission based on the communication willingness and sends it to the corresponding accelerator. The accelerator will not begin transmitting the tensor until it receives the communication permission, which means that the scheduler agrees to transmit the tensor. As a specific embodiment, three special operators can be defined for communication between accelerators: the want operator, the recv operator, and the send operator. The want operator is used to notify the master that the tensor is ready for transmission; the recv and send operators actually call communication library functions; these two operators must be executed in pairs. Because the transmission requires the scheduler's approval, the execution of recv and send requires the scheduler to issue a permission instruction, and the execution of these communication operators is subject to conditional restrictions.
[0094] That is to say, the execution of some operators must meet certain conditions. For example, the execution condition of the pair of send and receive communication operators given above is to receive the communication permission instruction from the scheduler, at which time the communication operator will enter In addition, the execution conditions of some operators can be triggered by operands. For example, when the operands of an operator are all calculated, the operator enters the ready state and starts execution.
[0095] In some embodiments of the present application, when the communication willingness between at least one group of accelerators is identified, it also includes: constructing a communication willingness graph based on the communication willingness between at least one group of accelerators; performing a communication scheduling control strategy on at least one group of accelerators based on the communication willingness graph, so as to generate communication permissions corresponding to each group of accelerators based on the communication scheduling control strategy.
[0096] Specifically, if accelerator a wants to transfer a tensor to accelerator b, the scheduler will receive a communication intent from a to b. When multiple tasks are stacked in a cluster, communication intent becomes quite complex. Under certain constraints (for example, the send and recv operators on a single accelerator cannot execute simultaneously), it is possible to activate as many communication intents as possible. A communication intent graph is used to represent the communication intent within the cluster. Communication is then managed based on a communication scheduling control strategy (such as maximum bipartite matching). In the communication intent graph, nodes represent accelerators, and edges represent the communication intent between a group of accelerators.
[0097] For example, the communication intention map is .in, With the above in The same is true for the nodes of the communication intention graph, namely the accelerator; is the edge of the communication intention graph, The operator in The calculation result of needs to be transmitted to , and this transmission has not ended yet; Refers to The ongoing transmission is being executed.
[0098] It is used to construct the communication willingness graph inside the scheduler through an interface, and the communication willingness in the accelerator cluster is tracked in real time. Every time the communication willingness graph is updated, the maximum bipartite matching problem is solved incrementally, and then the corresponding communication operator is triggered, so as to achieve real-time cluster transmission management.
[0099] For example, The transmission event is added to the communication willingness graph through this function, ; then the maximum bipartite matching problem is solved incrementally. If this transmission event should be started, then is made, that is, the communication permission is triggered. Refers to that the ongoing transmission event has been completed, that is, , ; if the completion of this transmission event will change the solution of the maximum bipartite matching, then the corresponding transmission event is started, that is to say, .
[0100] This embodiment enables the data flow in the accelerator cluster to be controllable, so as to maximize the utilization of network transmission capabilities.
[0101] As an embodiment of this application, as Figure 2 shown, the heterogeneous cluster system includes a scheduler and multiple accelerators. Both the scheduler and the accelerators expose a set of interfaces externally. The scheduler and the accelerators can call each other's interfaces through RPC (remote procedure call, calling a local function from a remote network). The cluster watchdog tells the scheduler the change of the whole cluster scale in real time. As Figure 7 shown, the task scheduling method executed by the scheduler may include the following steps: S101, Receive the accelerator cluster pattern feedback by the cluster watchdog to determine multiple accelerators available for task scheduling.
[0102] S102, Receive at least one task.
[0103] Specifically, receive the directed acyclic graph of the task. The calculation processes of all services in this heterogeneous cluster are described in the form of a directed acyclic graph, and then it is sent to the scheduler.
[0104] S103, After receiving the given task execution instruction, start this task.
[0105] S104. First, generate several scheduling policies that use different types and amounts of resources according to the received tasks. Then, use cluster prediction to evaluate the superiority of each scheduling policy one by one, and select the best target task scheduling policy.
[0106] S105. Generate a global computation graph based on the target task scheduling policy.
[0107] S106. Send the operators to the corresponding accelerators based on the global computation graph to complete the operator calculations by invoking the corresponding accelerators.
[0108] S107. When it is recognized that there is a communication intention between the accelerators, trigger communication permission so that the accelerators can perform data transmission based on the communication permission.
[0109] S108. Receive the feedback information indicating that the operator calculations and / or transmissions on the accelerators are completed, and update the global computation graph based on the feedback information. For example, mark the operator as completed.
[0110] Further combined with Figure 8 As shown, the task is assigned to Accelerator 1 and Accelerator 2. After starting the task execution, track all relevant times, including changes in the global computation graph, changes in the out-of-order queues of Accelerator 1 and Accelerator 2; and the process of interface calls in the entire system. Specifically as follows: Call the run_task function to add an operator to the global computation graph and generate the corresponding mapping relationship.
[0111] And The scheduler adds the operator to the out-of-order queues of Accelerator 1 and Accelerator 2 respectively based on the global computation graph. The newly added operator is non-ready.
[0112] Accelerator 1 calls to notify the scheduler that the operator has been executed.
[0113] Accelerator 1 triggers a communication willingness to the scheduler, and this communication willingness is to transfer the tensor a (operand) from Accelerator 1 to Accelerator 2.
[0114] The scheduler allows the tensor transmission and changes the recv operator (receive operator) on Accelerator 2 to the ready state.
[0115] The recv operator on Accelerator 2 starts to execute, and Accelerator 2 calls the left_trace interface of Accelerator 1 to change the send operator (send operator) on Accelerator 1 to the ready state.
[0116] The send operator on accelerator 1 starts to execute, docks with the recv operator being executed on accelerator 2, and starts the transmission. After the transmission is completed, the recv operator on accelerator 2 calls the report_recv_done function to tell the scheduler that the transmission is completed.
[0117] Accelerator 2 calls the report_operator_done function to tell the scheduler that operators y1 and y2 have been executed.
[0118] Accelerator 2 calls the report_operator_done function to tell the scheduler that operator w has been executed.
[0119] Based on the global computation graph, this embodiment combines the computational tasks of all jobs into a large computational graph, generates a unified scheduling strategy, and launches operators into the accelerator cluster, realizing a unified mapping mechanism from operators to accelerators. It supports heterogeneous accelerators to expose a unified operator execution interface externally, rather than forcing all operators to adapt to the general kernel code, eliminating the performance overhead of cross-platform kernel conversion. As a result, multiple heterogeneous devices can cooperate to participate in the parallel computing of tasks, effectively improving the overall resource utilization rate and scheduling flexibility.
[0120] At the same time, the global computation graph is used to track the working conditions of all accelerators in the cluster in real time. The resource distribution and load trend of the cluster in the future period can be predicted based on the current global computation graph state, so as to evaluate the cost and performance impact of different scheduling strategies. It supports generating multiple resource occupancy plans when submitting tasks and selecting the optimal plan with the help of a prediction model to achieve fast and flexible scheduling response capabilities.
[0121] Based on the above analysis, it can be seen that this application has the following technical advantages: (1) Global operator-level scheduling granularity, breaking through the limitations of the existing task-level scheduling; (2) Support for native operator execution on heterogeneous accelerators, eliminating cross-platform execution bottlenecks; (3) The cluster pattern is predictable and the scheduling decision can be optimized, with elastic scheduling capabilities; (4) Support for joint scheduling across tasks, effectively alleviating the problem of resource fragmentation.
[0122] This task scheduling method is applicable to complex computing scenarios that require highly parallel and heterogeneous resource scheduling, such as large-scale graph neural networks, large language models, machine learning training and inference. Examples are as follows: 1. Training: Extract the computational process of model training into a computational graph and then iteratively submit it to the computing nodes. The computing nodes can automatically and efficiently find the idle computing power in the cluster for training; automatically handle different capacity granularities during the training process; and automatically handle occasional computational error problems.
[0123] 2. Inference: Describe the model inference process in the form of a computational graph and then submit it to the computing nodes. The computing nodes can automatically adapt to the fluctuations in business pressure caused by different time periods; and can automatically achieve pipelined parallelism to fully utilize the computing power of the cluster.
[0124] 3. General computing: Describe the task loads in various fields in the form of a computational graph and then submit it to the computing nodes to achieve general computing of the cluster.
[0125] Through the description of the above implementation manners, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0126] The embodiments of the present application also provide a task scheduling device, which is applied to a heterogeneous cluster. As Figure 9 shown, the task scheduling device of the embodiments of the present application includes: a receiving module 10, a building module 20, and a scheduling module 30.
[0127] Among them, the receiving module 10 is used to receive at least one task to be executed, where each task includes multiple operators and the data dependency relationships between the multiple operators; the building module 20 is used to determine the first mapping relationship between the multiple operators in each task and the multiple accelerators, and establish a global computational graph based on the first mapping relationship corresponding to each task; the scheduling module 30 is used to send each operator to the corresponding accelerator according to the global computational graph, and execute at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators.
[0128] In some embodiments of the present application, the building module 20 determines the first mapping relationship between the multiple operators in each task and the multiple accelerators, specifically: determining multiple task scheduling strategies corresponding to the task based on a preset scheduling generation method; determining the scheduling performance of each task scheduling strategy according to the preset computing duration of each operator in the corresponding accelerator and the preset communication duration between the multiple accelerators; determining the target task scheduling strategy according to the scheduling performance of each task scheduling strategy, and generating the first mapping relationship between the multiple operators in the task and the multiple accelerators based on the target task scheduling strategy.
[0129] In some embodiments of the present application, when multiple accelerators are in the task execution state, the establishment module 20 is further configured to: determine the current global computation graph being executed by the multiple accelerators; determine the idle start time of each accelerator according to the current execution state of the current global computation graph, the preset computation duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the multiple accelerators; predict the execution end time of each task scheduling policy according to the preset computation duration of each operator in the task in the corresponding accelerator, the preset communication duration between the multiple accelerators, and the idle start time of each accelerator, so as to determine the scheduling performance of each task scheduling policy.
[0130] In some embodiments of the present application, the establishment module 20 determines the idle start time of each accelerator according to the current execution state of the current global computation graph, the preset computation duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the multiple accelerators. Specifically, it is configured to: determine the execution completion time of each operator in the current global computation graph according to the current execution state of the current global computation graph, the preset communication duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the multiple accelerators; predict the idle start time of each accelerator according to the execution completion time of each operator in the current global computation graph.
[0131] In some embodiments of the present application, the global computation graph includes the state information of each operator in at least one task. During the process of respectively sending each operator to the corresponding accelerator according to the global computation graph and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between the multiple operators, the establishment module 20 is further configured to: receive the operator state feedback information sent by one or more of the multiple accelerators; update the state information of each operator according to the operator state feedback information.
[0132] In some embodiments of the present application, the establishment module 20 updates the state information of each operator according to the operator state feedback information. Specifically, it is configured to: delete the operator and the associated information of the operator in the global computation graph when it is determined according to the operator state feedback information that the operator is in the execution completion state.
[0133] In some embodiments of the present application, during the process of separately sending each operator to the corresponding accelerator according to the global computation graph and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators, the establishment module 20 is further configured to: in response to an operator insertion instruction, identify the target operator corresponding to the operator insertion instruction and the data dependency relationships between the target operator and the operators in the global computation graph; establish a second mapping relationship between the target operator and the multiple accelerators based on the data dependency relationships between the target operator and the operators in the global computation graph, and update the global computation graph based on the second mapping relationship.
[0134] In some embodiments of the present application, the global computation graph includes an accelerator set corresponding to multiple accelerators. During the process of separately sending each operator to the corresponding accelerator according to the global computation graph and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators, the establishment module 20 is further configured to: obtain the status information of each accelerator in the heterogeneous cluster; update the accelerator set according to the status information of each accelerator.
[0135] In some embodiments of the present application, when the establishment module 20 updates the accelerator set according to the status information of each accelerator, it is specifically configured to: when it is determined according to the status information of each accelerator that there is a newly online accelerator in the heterogeneous cluster, add the newly online accelerator to the accelerator set; and / or when it is determined according to the status information of each accelerator that there is an offline accelerator in the accelerator set, re - establish a third mapping relationship between the uncompleted operators corresponding to the offline accelerator and the online accelerators in the accelerator set, and then delete the offline accelerator in the accelerator set.
[0136] In some embodiments of the present application, when it is determined according to the current status of each accelerator that there is an offline accelerator in the accelerator set, the scheduling module 30 is further configured to: obtain the intermediate operator calculation results stored in the offline accelerator, and send the intermediate operator calculation results to the corresponding online accelerator based on the data dependency relationships between the multiple operators.
[0137] In some embodiments of the present application, during the process of separately sending each operator to the corresponding accelerator according to the global computation graph and executing at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators, the scheduling module 30 is further configured to: when there is an operator calculation error, move the operator with the calculation error out of the original accelerator, establish a fourth mapping relationship between the operator with the calculation error and other accelerators in the multiple accelerators, and update the global computation graph based on the fourth mapping relationship.
[0138] In some embodiments of the present application, after the scheduling module 30 updates the global computation graph based on the fourth mapping relationship, it is further configured to: obtain the operands of the operator with a computation error, where the operands are stored in the corresponding accelerator; and send the operator with the computation error and the operands to the corresponding accelerator based on the updated global computation graph.
[0139] In some embodiments of the present application, the scheduling module 30 is further configured to: obtain the actual computation duration of the operator, and determine that the operator has a computation error when the actual computation duration of the operator exceeds a preset duration; or determine that the operator has a computation error when the accelerator where the operator is located has an exception; or determine that the operator corresponding to the operator exception instruction has a computation error in response to the operator exception instruction.
[0140] In some embodiments of the present application, the scheduling module 30 is further configured to: identify the communication willingness between at least one group of accelerators; and generate a corresponding communication permission based on the communication willingness between at least one group of accelerators, so that at least one group of accelerators can communicate based on the communication permission.
[0141] In some embodiments of the present application, when the scheduling module 30 identifies the communication willingness between at least one group of accelerators, it is further configured to: construct a communication willingness graph according to the communication willingness between at least one group of accelerators; and perform a communication scheduling control strategy on at least one group of accelerators based on the communication willingness graph, so as to generate a communication permission corresponding to each group of accelerators based on the communication scheduling control strategy.
[0142] In some embodiments of the present application, the receiving module 10 receives at least one task to be executed, and specifically is configured to: receive a directed acyclic graph of at least one task, where the nodes of the directed acyclic graph are used to represent the corresponding operators, and the edges of the directed acyclic graph are used to represent the data dependency relationship between the operators.
[0143] For the description of the features in the corresponding embodiments of the task scheduling device, reference can be made to the relevant description of the corresponding embodiments of the task scheduling method, which will not be elaborated here one by one.
[0144] Embodiments of the present application further provide an electronic device, as Figure shown, the electronic device 100 includes a memory 110 and a processor 120. A computer program is stored in the memory 110, and the processor 120 is configured to run the computer program to execute the steps in any of the above embodiments of the task scheduling method.
[0145] Embodiments of the present application further provide a non-volatile computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above embodiments of the task scheduling method when running.
[0146] In an exemplary embodiment, the above non-volatile computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs.
[0147] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the task scheduling method.
[0148] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the task scheduling method.
[0149] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0150] The above has introduced in detail a task scheduling method, a scheduling device, an electronic device, a storage medium, and a product provided by the present application. Specific examples are used in this article to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A task scheduling method, characterized in that, Applied to a heterogeneous cluster, the heterogeneous cluster including a plurality of accelerators, the method includes: Receiving at least one task to be executed, where each task includes a plurality of operators and data dependency relationships between the plurality of operators; Determining a first mapping relationship between the plurality of operators in each task and the plurality of accelerators, and establishing a global computation graph based on the first mapping relationship corresponding to each task; Sending each operator to the corresponding accelerator according to the global computation graph, and executing the at least one task through one or more of the plurality of accelerators in combination with the data dependency relationships between the plurality of operators.
2. The task scheduling method according to claim 1, wherein The determining a first mapping relationship between the plurality of operators in each task and the plurality of accelerators includes: Determining a plurality of task scheduling policies corresponding to the task based on a preset scheduling generation method; Determining the scheduling performance of each task scheduling policy according to the preset computation duration of each operator in the corresponding accelerator in the task and the preset communication duration between the plurality of accelerators; Determining a target task scheduling policy according to the scheduling performance of each task scheduling policy, and generating a first mapping relationship between the plurality of operators in the task and the plurality of accelerators based on the target task scheduling policy.
3. The task scheduling method according to claim 2, wherein In the case where the plurality of accelerators are in a task execution state, it further includes: Determining the current global computation graph being executed by the plurality of accelerators; Determining the idle start time of each accelerator according to the current execution state of the current global computation graph, the preset computation duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the plurality of accelerators; Predicting the execution end time of each task scheduling policy according to the preset computation duration of each operator in the task in the corresponding accelerator, the preset communication duration between the plurality of accelerators, and the idle start time of each accelerator, so as to determine the scheduling performance of each task scheduling policy.
4. The task scheduling method according to claim 3, wherein The determining the idle start time of each accelerator according to the current execution state of the current global computation graph, the preset computation duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the plurality of accelerators includes: Determining the execution completion time of each operator in the current global computation graph according to the current execution state of the current global computation graph, the preset communication duration of each operator in the current global computation graph in the corresponding accelerator, and the preset communication duration between the plurality of accelerators; Predicting the idle start time of each accelerator according to the execution completion time of each operator in the current global computation graph.
5. The task scheduling method according to claim 1, characterized in that The global computation graph includes the state information of each operator in the at least one task. In the process of sending each operator to the corresponding accelerator according to the global computation graph and executing the at least one task through one or more of the plurality of accelerators in combination with the data dependency relationships between the plurality of operators, it further includes: Receiving operator state feedback information sent by one or more of the plurality of accelerators; Updating the state information of each operator according to the operator state feedback information.
6. The task scheduling method according to claim 5, wherein, Updating the status information of each operator according to the operator status feedback information includes: When it is determined according to the operator status feedback information that the operator is in the execution completion state, deleting the operator and the associated information of the operator in the global computation graph.
7. The task scheduling method according to claim 1, wherein In the process of sending each operator to the corresponding accelerator according to the global computation graph and executing the at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between the multiple operators, it further includes: In response to an operator insertion instruction, identifying the target operator corresponding to the operator insertion instruction and the data dependency relationship between the target operator and the operators in the global computation graph; Establishing a second mapping relationship between the target operator and the multiple accelerators based on the data dependency relationship between the target operator and the operators in the global computation graph, and updating the global computation graph based on the second mapping relationship.
8. The task scheduling method according to claim 1, wherein The global computation graph includes an accelerator set corresponding to the multiple accelerators. In the process of sending each operator to the corresponding accelerator according to the global computation graph and executing the at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between the multiple operators, it further includes: Obtaining the status information of each accelerator in the heterogeneous cluster; Updating the accelerator set according to the status information of each accelerator.
9. The task scheduling method according to claim 8, wherein Updating the accelerator set according to the status information of each accelerator includes: When it is determined according to the status information of each accelerator that there is a newly online accelerator in the heterogeneous cluster, adding the newly online accelerator to the accelerator set; and / or When it is determined according to the status information of each accelerator that there is an offline accelerator in the accelerator set, re - establishing a third mapping relationship between the uncompleted operators corresponding to the offline accelerator and the online accelerators in the accelerator set, and then deleting the offline accelerator in the accelerator set.
10. The task scheduling method according to claim 9, characterized in that, When it is determined according to the current status of each accelerator that there is an offline accelerator in the accelerator set, it further includes: Obtaining the intermediate operator calculation results stored in the offline accelerator, and sending the intermediate operator calculation results to the corresponding online accelerator based on the data dependency relationship between the multiple operators.
11. The task scheduling method according to claim 1, wherein In the process of sending each operator to the corresponding accelerator according to the global computation graph and executing the at least one task through one or more of the multiple accelerators in combination with the data dependency relationship between the multiple operators, it further includes: When there is an operator calculation error, moving the operator with the calculation error out of the original accelerator, establishing a fourth mapping relationship between the operator with the calculation error and other accelerators in the multiple accelerators, and updating the global computation graph based on the fourth mapping relationship.
12. The task scheduling method according to claim 11, wherein After updating the global computation graph based on the fourth mapping relationship, it further includes: Obtaining the operands of the operator with the calculation error, where the operands are stored in the corresponding accelerator; Send the operator with the calculation error and the operand to the corresponding accelerator based on the updated global computational graph.
13. The task scheduling method according to claim 11, characterized in that, It further includes: Obtain the actual calculation duration of the operator, and determine that the operator has a calculation error when the actual calculation duration of the operator exceeds a preset duration; Or Determine that the operator has a calculation error when the accelerator where the operator is located has an abnormality; Or In response to an operator abnormality instruction, determine that the operator corresponding to the operator abnormality instruction has a calculation error.
14. The task scheduling method according to claim 1, wherein It further includes: Identify the communication willingness between at least one group of accelerators; Generate a corresponding communication permission based on the communication willingness between the at least one group of accelerators, so that the at least one group of accelerators can communicate based on the communication permission.
15. The task scheduling method according to claim 14, wherein When the communication willingness between at least one group of accelerators is identified, it further includes: Construct a communication willingness graph according to the communication willingness between the at least one group of accelerators; Perform a communication scheduling control strategy on the at least one group of accelerators based on the communication willingness graph, so as to generate a corresponding communication permission for each group of accelerators based on the communication scheduling control strategy.
16. The task scheduling method according to claim 1, wherein The receiving of at least one task to be executed includes: Receive a directed acyclic graph of at least one task, where the nodes of the directed acyclic graph are used to represent the corresponding operators, and the edges of the directed acyclic graph are used to represent the data dependency relationships between the operators.
17. A task scheduling device, characterized in that, Applied to a heterogeneous cluster, the heterogeneous cluster includes multiple accelerators, and the device includes: A receiving module, configured to receive at least one task to be executed, where each task includes multiple operators and the data dependency relationships between the multiple operators; A establishing module, configured to determine a first mapping relationship between the multiple operators in each task and the multiple accelerators, and establish a global computational graph based on the first mapping relationship corresponding to each task; A scheduling module, configured to send each operator to the corresponding accelerator according to the global computational graph, and execute the at least one task through one or more of the multiple accelerators in combination with the data dependency relationships between the multiple operators.
18. An electronic device, characterized in that, It includes: A memory, configured to store a computer program; A processor, configured to implement the steps of the task scheduling method according to any one of claims 1 to 16 when executing the computer program.
19. A non-volatile computer-readable storage medium, characterized in that, A non-volatile computer-readable storage medium stores a computer program, where the computer program implements the steps of the task scheduling method according to any one of claims 1 to 16 when executed by a processor.
20. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the task scheduling method according to any one of claims 1 to 16 when executed by a processor.
Citation Information
Patent Citations
Chiplet-based hardware acceleration method and hardware accelerator
CN116932226A
Neural network accelerator scheduling method and device, electronic equipment and storage medium
CN117010465A
Operation acceleration method and operation accelerator
CN117501278A
Multi-kernel neural vector retrieval hardware accelerator and scheduling method thereof
CN118502900A
Memory management method and device and AI computing equipment
CN119440785A