Task scheduling method and device

By combining task dispatch and task retrieval modes, this task scheduling method overcomes the limitations of single scheduling modes in existing technologies, achieving efficient, flexible, and high-performance task scheduling that is adaptable to different applications and platforms and supports user-defined strategies.

CN120803624APending Publication Date: 2025-10-17HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410444912.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing task scheduling model only supports one scheduling mode, which leads to room for optimization of task scheduling, inability to flexibly adapt between different applications and platforms, and lack of support for user-defined policies.

Method used

This paper provides a task scheduling method that combines the task delivery mode and the task pull mode. It realizes dynamic scheduling and pulling of tasks through the collaborative work of private task containers and public task containers, supports user-defined strategies, and adapts to the needs of different scenarios.

Benefits of technology

It improves task scheduling performance, enhances load balancing and scalability, adapts to different applications and platforms, supports user-defined strategies, and optimizes the flexibility and efficiency of task scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803624A_ABST
    Figure CN120803624A_ABST
Patent Text Reader

Abstract

The invention provides a task scheduling method and device. The method comprises the steps of obtaining a to-be-scheduled task set; scheduling the to-be-scheduled task to M private task containers and public task containers corresponding to M execution units of a local scheduling domain; the tasks in each private task container in the M private task containers are issued to the private task queue of the execution unit corresponding to each private task container, so that the tasks in the private task queues are executed by the execution units; and in response to a task pulling request of the first execution unit, selecting a task from the public task container and adding the task into a first private task container, the task pulling request being generated when the number of tasks in a private task queue of the first execution unit does not reach the standard, and the first private task container being a private task container corresponding to the first execution unit. The task scheduling method provided by the invention supports the task issuing mode and the task pulling mode at the same time, realizes advantage complementation of the two modes, and improves the task scheduling performance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a task scheduling method and device. BACKGROUND

[0002] The task-based programming model can automatically analyze the dependency relationship between tasks through a runtime, and execute tasks that have no dependency with other tasks (referred to as executable tasks), effectively alleviating the problem of idle resources caused by manual dependency management in the traditional fork-join programming model in a complex parallel environment, and improving the utilization of hardware resources (such as Figure 1

[0003] The existing task scheduling model often only supports one task scheduling mode, for example, the OMMPSS2 task scheduling model and the FFRT2.0 task scheduling model only support the task pull mode, and the Legion task scheduling model and the TaskFlow task scheduling model only support the task push mode, but the two modes have advantages and disadvantages, resulting in optimization space for the task scheduling of the current task scheduling model. SUMMARY

[0004] Embodiments of the present application provide a task scheduling method and device, which simultaneously supports the task push mode and the task pull mode, realizes the complementary advantages of the two modes, and improves the performance of task scheduling.

[0005] In a first aspect, the present application provides a task scheduling method applied to a runtime task scheduling system, the method comprising: obtaining a task set to be scheduled, the task set comprising N tasks, the N tasks all being executable tasks, and N being a positive integer; scheduling the N tasks to M private task containers and a public task container corresponding to M execution units of a local scheduling domain, M being a positive integer; pushing the tasks in each of the M private task containers to a private task queue of an execution unit corresponding to the private task container, so that the tasks in the private task queue are executed by the execution unit; and in response to a task pull request of a first execution unit, selecting a task from the public task container and adding the task to a first private task container, the task pull request being generated when the number of tasks in the private task queue of the first execution unit is less than or equal to a preset threshold, and the first private task container being a private task container corresponding to the first execution unit.

[0006] The task scheduling method provided by the present application simultaneously supports the task push mode and the task pull mode, realizes the complementary advantages of the two modes, and improves the performance of task scheduling, for example, the task scheduling method provided by the present application has the advantages of scalability and low competition of the task push mode, and the advantage of load balancing of the task pull mode.

[0007] ​In one possible implementation, a specific implementation of scheduling N tasks to M private task containers and a public task container corresponding to M executable units of a local scheduling domain is as follows: obtaining first task information, first hardware information, and first scheduling domain state information, the first task information indicating task type information of each task in the N tasks, the first hardware information indicating hardware information in the local scheduling domain, and the first scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; taking the first task information, the first hardware information, and the first scheduling domain state information as inputs of an execution unit selection strategy, and outputting a target execution unit corresponding to each task; and scheduling each task to a private container of the target execution unit corresponding to the task.

[0008] In this possible implementation, a user is provided with a custom strategy mounting interface to support mounting of a user-defined execution unit selection strategy, and to implement user-defined strategy override of execution unit selection in the task scheduling issuing stage.

[0009] In another possible implementation, for a task for which the execution unit selection strategy does not output a target execution unit, the task is scheduled to a public task container of the local scheduling domain.

[0010] In another possible implementation, the runtime task scheduling system includes multiple scheduling domains; scheduling N tasks to M private task containers and a public task container corresponding to M executable units of a local scheduling domain, and further including: obtaining second task information, second hardware information, and second scheduling domain state information, the second task information indicating task type information of each task in the N tasks, the second hardware information indicating hardware information of each scheduling domain in the multiple scheduling domains, and the second scheduling domain state information indicating state information of each hardware unit in each scheduling domain in the multiple scheduling domains; taking the second task information, the second hardware information, and the second scheduling domain state information as inputs of a scheduling domain selection strategy, and outputting a first target scheduling domain corresponding to each task; and scheduling each task to the target scheduling domain corresponding to the task.

[0011] For a multi-scheduling domain application scenario, such as a scenario with a remote scheduling domain, when a task is scheduled and issued, a target scheduling domain for task scheduling and issuing is first determined, for example, whether the task is scheduled and issued to a local scheduling domain or a remote scheduling domain, and then an execution unit selection strategy is used to select an execution unit for task issuing. The target scheduling domain for task scheduling and issuing also supports a user-defined scheduling domain selection strategy, and a suitable scheduling domain is selected for a task through the user-defined scheduling domain selection strategy.

[0012] In one example, the obtaining the second task information, the second hardware information and the second scheduling domain state information further comprises: obtaining affinity descriptions of the tasks and hardware information of the execution units in the plurality of scheduling domains; determining the prompt information corresponding to each of the tasks based on the affinity descriptions of the tasks, the hardware information of the execution units in the plurality of scheduling domains and an affinity description analysis strategy, the prompt information being used to prompt a target execution unit corresponding to the task; and the first task information further comprises the prompt information corresponding to each of the tasks.

[0013] The affinity description of the task is converted into the selection prompt of the execution unit by the affinity description analysis strategy, which assists the subsequent execution unit selection strategy to select a suitable execution unit for the task to be scheduled.

[0014] Optionally, the affinity description comprises an explicit affinity description or an implicit affinity description.

[0015] In another possible implementation, a specific implementation of the task in each of the M private task containers being dispatched to a private task queue of an execution unit corresponding to each of the private task containers comprises: obtaining information of each of the private task containers, third hardware information and third scheduling domain state information, the information of each of the private task containers indicating task information in each of the private task containers, the third hardware information indicating hardware information corresponding to the execution unit where each of the private task containers is located, and the third scheduling domain state information indicating state information of each of the hardware units in the local scheduling domain; taking the information of each of the private task containers, the third hardware information and the third scheduling domain state information as inputs of a first task dispatching strategy, and outputting a first task list, the first task list indicating tasks to be dispatched in the current time by each of the private task containers and an order of the tasks to be dispatched in the current time; and based on the first task list, dispatching the tasks in each of the private task containers to the private task queue of the execution unit corresponding to each of the private task containers, the execution order of the tasks in the private task queue of the execution unit being the same as the order of the tasks in the task list.

[0016] The order in which the tasks are executed by the execution units is controlled by the task dispatching strategy, for example, a task with a higher priority can be executed first, and a task with a lower priority can be executed later; a task that is more suitable for the current memory condition is executed first, for example, the available space of the current memory is 5 MB, the memory occupation of task 1 when executed is 4.9 MB, and the memory occupation of task 2 when executed is 5.6 MB, task 1 is arranged before task 2, and task 1 is executed first.

[0017] In another possible implementation, a specific implementation of the task in the first private task container being selected from the public task container in response to a task pulling request of the first execution unit comprises: determining that the number of tasks in the first private task container is insufficient; and selecting a task from the public task container and adding the task to the first private task container.

[0018] When the number of tasks in the private task container of the execution unit is insufficient, tasks are pulled from the public task container to the private container, so that the tasks of the execution unit are actively pulled.

[0019] In another possible implementation, a specific implementation of selecting tasks from the public task container to join the first private task container includes: obtaining public task container information and fourth hardware information, the public task container information indicating attribute information of each task in the public task container, and the fourth hardware information indicating hardware information of the first execution unit; taking the public task container information and the fourth hardware information as inputs of a task selection strategy, and outputting a target task and a target task quantity, the target task indicating a task selected from the public task container; and adding the target task to the first private task container.

[0020] According to the task information (for example, task attribute information) in the public task container and the hardware information of the execution unit, a certain number of tasks in the public task container are selected and transferred to the execution unit that initiates the task pulling operation. For example, the execution unit is a graphic processing unit (GPU), and the GPU is suitable for performing double-precision floating point operations (for example, matrix or vector operations), so corresponding tasks are selected from the public task container and added to the private task container of the GPU; for another example, the execution unit is a neural network processing unit (NPU), and the NPU is suitable for performing half-precision floating point operations, so corresponding tasks are selected from the public task container and added to the private task container of the NPU.

[0021] In another possible implementation, selecting tasks from the public task container to join the first private task container further includes: determining that the number of tasks in the first private task container is sufficient; and selecting tasks from the first private task container to join the private task queue of the first execution unit.

[0022] In other words, when the number of tasks in the execution unit in the private task queue of the execution unit is insufficient (for example, less than 16), tasks are first pulled from the private container of the execution unit to the private queue of the execution unit itself, and if the number of tasks pulled from the private container is insufficient (for example, 48 tasks need to be pulled, but only 32 tasks are pulled, and 16 tasks are still needed), the remaining number of tasks are pulled from the public task queue.

[0023] In another possible implementation, a specific implementation of selecting tasks from the first private task container to join the private task queue of the execution unit is as follows: obtaining first private task container information, fifth hardware information, and fourth scheduling domain state information, the first private task container information indicating task information in the private task container of the first execution unit, the fifth hardware information indicating hardware information of the first execution unit, and the fourth scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; taking the first private task container information, the fifth hardware information, and the fourth scheduling domain state information as inputs of a second task issuing strategy, and outputting a second task list, the second task list indicating tasks issued in the current time by the first private task container and an order of the tasks; and based on the second task list, issuing the tasks in the first private task container to the private task queue of the execution unit, the tasks in the private task queue of the execution unit being executed in the same order as the order of the tasks in the second task list.

[0024] In another possible implementation, the task scheduling method provided in this application further includes: determining that the number of tasks selected from the common task container is insufficient; obtaining second private task container information, the second private task container information indicating task information in other task containers in the M private task containers except the first private task container; taking the second private task container information as an input of a local task stealing strategy, and outputting a stealing target task and a first stealing number, the stealing target task indicating a task stolen from the other task containers, and the first stealing number indicating the number of tasks stolen from the other task containers; based on the stealing target task and the stealing number, performing a task stealing operation from the other task containers; and adding the tasks stolen from the other task containers to the first private task container.

[0025] The task scheduling method provided in this application also supports local task stealing, that is, when the number of tasks to be executed in the private queue of the execution unit is insufficient, and not enough tasks can be pulled from the private task container and the common task container, the local task stealing strategy can also be used to steal tasks from the private containers of other execution units in the local scheduling domain.

[0026] In another possible implementation, the task scheduling method provided in the present application further includes: determining that the number of tasks stolen by the local task stealing strategy is insufficient; obtaining sixth hardware information and fifth scheduling domain state information, the sixth hardware information indicating hardware information of scheduling domains other than the local scheduling domain in the plurality of scheduling domains, and the fifth scheduling domain state information indicating state information of each hardware unit in the other scheduling domains; taking the sixth hardware information and the fifth scheduling domain state information as inputs of a remote task stealing strategy, and outputting a stealing target and a second stealing number, the stealing target indicating a second target scheduling domain in the other scheduling domains that is to be subjected to a stealing operation, and the second stealing number indicating the number of tasks stolen from the stealing target scheduling domain; performing a task stealing operation from the second target scheduling domain based on the stealing target and the stealing number; and adding the tasks stolen from the second target scheduling domain to the first private task container.

[0027] The task scheduling method provided in the present application further supports a remote task stealing mechanism, so that when the number of tasks stolen from the local scheduling domain is still insufficient, tasks can be stolen from a remote scheduling domain, thereby ensuring that an execution unit with insufficient number of tasks to be executed can steal a sufficient number of tasks.

[0028] In another possible implementation, the task scheduling method provided in the present application further includes: obtaining information of an abnormal event, the abnormal event including one or more of the following: the number of tasks to be executed in the local scheduling domain reaches the upper limit of the task capacity, the target scheduling domain is unreachable, the number of tasks to be executed in the target scheduling domain reaches the upper limit of the task capacity, a sufficient number of tasks is not selected from the public task container and the local task stealing strategy is not enabled, a sufficient number of tasks is not stolen by the local task stealing strategy and the remote task stealing strategy is not enabled, and a sufficient number of tasks is not stolen by the remote task stealing strategy; taking the information of the abnormal event as an input of an abnormal event processing strategy, and outputting a processing strategy, the processing strategy indicating a processing operation for the abnormal event.

[0029] For abnormal situations that may occur in the task scheduling process, corresponding processing measures are set in advance to ensure the smooth progress of task scheduling.

[0030] In another possible implementation, one or more of the execution unit selection strategy, the scheduling domain selection strategy, the first task distribution strategy, the task selection strategy, the second task distribution strategy, the local task stealing strategy, the remote task stealing strategy, and the abnormal event processing strategy can be configured to be enabled or disabled. In this way, partial strategy operation is supported, and a user can disable steps that are not needed in the process, thereby reducing task scheduling overhead and improving task execution efficiency.

[0031] In another possible implementation, the task capacity of each private task container in the M private task containers can be configured to be 0; or the task capacity of the public task container can be configured to be 0.

[0032] That is, the user can switch among the three modes of supporting task push mode and task pull mode, supporting only task push mode and supporting only task pull mode by configuring to remove the private container or the public task container, adapting to different applications and platforms.

[0033] In a second aspect, the present application provides a task scheduling device, comprising an acquisition module, a scheduling module, a task push module and a task pull module, wherein the acquisition module is configured to acquire a task set to be scheduled, the task set comprising N tasks, the N tasks all being executable tasks, and N being a positive integer; the scheduling module is configured to schedule the N tasks to M private task containers and a public task container corresponding to M execution units of a local scheduling domain, M being a positive integer; the task push module is configured to push the tasks in each of the M private task containers to a private task queue of an execution unit corresponding to each private task container, so that the tasks in the private task queue are executed by the execution unit; and the task pull module is configured to, in response to a task pull request of a first execution unit, select a task from the public task container and add the task to a first private task container, the task pull request being generated when the number of tasks in the private task queue of the first execution unit is less than or equal to a preset threshold, and the first private task container being a private task container corresponding to the first execution unit.

[0034] In one possible implementation, the scheduling module is specifically configured to: acquire first task information, first hardware information and first scheduling domain state information, the first task information indicating task type information of each task in the N tasks, the first hardware information indicating hardware information in the local scheduling domain, and the first scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; take the first task information, the first hardware information and the first scheduling domain state information as inputs of an execution unit selection strategy, and output a target execution unit corresponding to each task; and schedule each task to a private container of the target execution unit corresponding to the task.

[0035] In another possible implementation, for a task for which the execution unit selection strategy does not output a target execution unit, the task is scheduled to a public task container of the local scheduling domain.

[0036] In another possible implementation, the runtime task scheduling system includes multiple scheduling domains; the scheduling module is further used to: obtain second task information, second hardware information and second scheduling domain status information, the second task information indicates the task type information of each task in the N tasks, the second hardware information indicates the hardware information of each scheduling domain in the multiple scheduling domains, and the second scheduling domain status information indicates the status information of each hardware unit in each scheduling domain in the multiple scheduling domains; use the second task information, the second hardware information and the second scheduling domain status information as inputs to the scheduling domain selection strategy, and output the first target scheduling domain corresponding to each task; and schedule each task to the target scheduling domain corresponding to each task.

[0037] In another possible implementation, the task scheduling device provided by the present application also includes an affinity parsing module, which is used to obtain the affinity description of each task and the hardware information of each execution unit in multiple scheduling domains; based on the affinity description of each task and the hardware information of each execution unit in multiple scheduling domains, as well as the affinity description parsing strategy, the prompt information corresponding to each task is determined, and the prompt information is used to prompt the target execution unit corresponding to each task; the first task information also includes the prompt information corresponding to each task.

[0038] Optionally, the affinity description includes an explicit affinity description or an implicit affinity description.

[0039] In another possible implementation, the task dispatching module is specifically used to: obtain information of each private task container, third hardware information and third scheduling domain status information, the information of each private task container indicates the task information in each private task container, the third hardware information indicates the hardware information corresponding to the execution unit where each private task container is located, and the third scheduling domain status information indicates the status information of each hardware unit in the local scheduling domain; use the information of each private task container, the third hardware information and the third scheduling domain status information as input of the first task dispatching strategy, and output a first task list, the first task list indicates the tasks currently dispatched by each private task container and the order of the tasks currently dispatched; based on the first task list, dispatch the tasks in each private task container to the private task queue of the execution unit corresponding to each private task container, and the execution order of the tasks in the private task queue of the execution unit is the same as the task order in the task list.

[0040] In another possible implementation, the task pulling module is specifically used to: determine that the number of tasks in the first private task container is insufficient; select tasks from the public task container and add them to the first private task container;

[0041] In another possible implementation, a specific implementation of selecting a task from the public task container to join the first private task container includes: obtaining public task container information and fourth hardware information, the public task container information indicating attribute information of each task in the public task container, and the fourth hardware information indicating hardware information of the first execution unit; taking the public task container information and the fourth hardware information as inputs of a task selection strategy, and outputting a target task and a target task quantity, the target task indicating a task selected from the public task container; and adding the target task to the first private task container.

[0042] In another possible implementation, the task pulling module is further configured to: determine that a task quantity in the first private task container is sufficient before the task is selected from the public task container to join the first private task container; and select a task from the first private task container to join a private task queue of the execution unit.

[0043] In another possible implementation, a specific implementation of selecting a task from the first private task container to join the private task queue of the execution unit includes: obtaining first private task container information, fifth hardware information, and fourth scheduling domain state information, the first private task container information indicating task information in the private task container of the first execution unit, the fifth hardware information indicating hardware information of the first execution unit, and the fourth scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; taking the first private task container information, the fifth hardware information, and the fourth scheduling domain state information as inputs of a second task issuing strategy, and outputting a second task list, the second task list indicating tasks issued in a current time by the first private task container and an order of the tasks issued in the current time; and issuing the tasks in the first private task container to the private task queue of the execution unit based on the second task list, the execution order of the tasks in the private task queue of the execution unit being the same as the order of the tasks in the second task list.

[0044] In another possible implementation, the task scheduling apparatus provided in the present application further includes a task stealing module, which is configured to determine that a quantity of tasks selected from the public task container is insufficient; obtain second private task container information, the second private task container information indicating task information in other task containers of the M private task containers except the first private task container; take the second private task container information as an input of a local task stealing strategy, and output a stealing target task and a first stealing quantity, the stealing target task indicating a task stolen from the other task containers, and the first stealing quantity indicating a quantity of tasks stolen from the other task containers; perform a task stealing operation from the other task containers based on the stealing target task and the stealing quantity; and add the task stolen from the other task containers to the first private task container.

[0045] In another possible implementation, the task stealing module is also used to, when it is determined that the number of tasks stolen through the local task stealing strategy is insufficient; obtain sixth hardware information and fifth scheduling domain status information, the sixth hardware information indicates hardware information of other scheduling domains except the local scheduling domain in multiple scheduling domains, and the fifth scheduling domain status information indicates status information of each hardware unit in other scheduling domains; use the sixth hardware information and the fifth scheduling domain status information as input of the remote task stealing strategy, output a stealing target and a second stealing quantity, the stealing target indicates a second target scheduling domain in other scheduling domains on which the stealing operation will be performed, and the second stealing quantity indicates the number of tasks stolen from the stealing target scheduling domain; based on the stealing target and the stealing quantity, perform a task stealing operation from the second target scheduling domain; and add the tasks stolen from the second target scheduling domain to the first private task container.

[0046] In another possible implementation, the task scheduling device provided by the present application also includes an abnormal event processing module, which is specifically used to obtain information about abnormal events, and the abnormal events include one or more of the following: the tasks to be executed in the local scheduling domain reach the upper limit of the task capacity, the target scheduling domain cannot be delivered, the tasks to be executed in the target scheduling domain reach the upper limit of the task capacity, not enough tasks are selected from the public task container and the local task stealing strategy is not turned on, not enough tasks are stolen through the local task stealing strategy and the remote task stealing strategy is not turned on, and not enough tasks are stolen through the remote task stealing strategy; the information about the abnormal event is used as the input of the abnormal event processing strategy, and the processing strategy is output, and the processing strategy indicates the processing operation for the abnormal event.

[0047] In another possible implementation, one or more of the above-mentioned execution unit selection strategy, scheduling domain selection strategy, first task dispatching strategy, task selection strategy, second task dispatching strategy, local task stealing strategy, remote task stealing strategy and abnormal event handling strategy can be configured to be turned on or off.

[0048] In another possible implementation, the task capacity of each of the M private task containers may be configured as 0; or the task capacity of the public task container may be configured as 0.

[0049] In a third aspect, an embodiment of the present application provides a computing device comprising a memory and a processor, wherein the memory stores instructions, and when the instructions are executed by the processor, the method described in the first aspect is implemented.

[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the first aspect is implemented.

[0051] In a fifth aspect, the embodiments of the present application further provide a computer program or a computer program product, which comprises instructions, when the instructions are executed, causing a computer to execute the method of the first aspect.

[0052] In a sixth aspect, the embodiments of the present application further provide a chip, comprising at least one processor and a communication interface, wherein the processor is configured to execute the method of the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0053] Figure 1 A comparison diagram of a task coding model and a traditional fork-join programming model;

[0054] Figure 2 A comparison diagram of advantages and disadvantages of an existing task scheduling model is shown;

[0055] Figure 3 A service code affinity prompt diagram is shown;

[0056] Figure 4 A system architecture diagram of a runtime task scheduling system using the task scheduling method provided by the embodiments of the present application is shown;

[0057] Figure 5 A system architecture diagram of another runtime task scheduling system using the task scheduling method provided by the embodiments of the present application is shown;

[0058] Figure 6 A deployment mode diagram of a runtime system provided by the embodiments of the present application as a running instance deployed in a cluster is shown;

[0059] Figure 7 A scheduling mechanism diagram of the task scheduling method provided by the embodiments of the present application is shown;

[0060] Figure 8 A flow diagram of the task scheduling method provided by the embodiments of the present application is shown;

[0061] Figure 9 A diagram of task dependency relationships between a plurality of to-be-processed tasks after dependency relationship analysis by a runtime analysis unit is shown;

[0062] Figure 10 A diagram of executable tasks without task dependency relationships analyzed by the runtime analysis unit is shown;

[0063] Figure 11 A diagram of scheduling executable tasks by a scheduling unit is shown;

[0064] Figure 12A schematic diagram showing a task dependency relationship in a set of executable tasks after an update;

[0065] Figure 13 A schematic diagram showing an implicit affinity description and a mapping relationship between the affinity description and hardware;

[0066] Figure 14 A schematic diagram showing an explicit affinity description and a mapping relationship between the affinity description and hardware;

[0067] Figure 15 A schematic diagram of a task scheduling mechanism of a task scheduling system after a task capacity of a private task container is configured to 0;

[0068] Figure 16 A schematic diagram of a task scheduling mechanism of a task scheduling system after a task capacity of a public task container is configured to 0;

[0069] Figure 17 A schematic diagram of an implementation process of a task issuing stage of a task scheduling method provided by an embodiment of the present application;

[0070] Figure 18 A schematic diagram of an implementation process of a task pulling stage of a task scheduling method provided by an embodiment of the present application;

[0071] Figure 19 A structural schematic diagram of a task scheduling apparatus provided by an embodiment of the present application;

[0072] Figure 20 A structural schematic diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0073] The term “and / or” mentioned in the present document is a description of the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that there are three cases of A alone, A and B together, and B alone. The symbol “ / ” in the present document represents an or relationship of the associated objects, for example, A / B represents A or B.

[0074] The terms “first” and “second” and the like in the description and claims in the present document are used to distinguish different objects, and are not used to describe a specific order of the objects. For example, the first task list and the second task list are used to distinguish different task lists, and are not used to describe a specific order of the task lists.

[0075] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0076] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple execution units refer to two or more execution units, etc.; multiple scheduling domains refer to two or more scheduling domains, etc.

[0077] To facilitate understanding of the solutions of the embodiments of the present application, the technical terms involved in this document are first explained below.

[0078] Runtime task scheduling is the process of selecting appropriate execution units and assigning available tasks to them for execution. This process directly determines the load on each execution unit and has a major impact on the execution time of parallel programs.

[0079] Figure 2 The following diagram shows the advantages and disadvantages of existing task scheduling models. Figure 2 As shown, the task scheduling of the existing task scheduling model mainly includes two modes and three characteristics. The two modes include the execution unit task pull mode (also referred to as the task pull mode or the reverse pull mode) and the scheduling unit task delivery mode (also referred to as the task delivery mode or the forward delivery mode). The two task scheduling modes have their own advantages and disadvantages.

[0080] For example, in terms of load balancing, the execution unit task pull mode offers better load balancing than the scheduling unit task delivery mode. This is because the execution unit task pull mode only causes load imbalance when the number of pending tasks is less than the number of execution unit tasks. However, the scheduling unit task delivery mode is unable to accurately estimate the runtime of various tasks, which can easily lead to load imbalance among execution units. This requires the introduction of features such as task stealing to compensate for this.

[0081] In terms of scalability, the scalability of the scheduling unit task issuing mode is higher than that of the execution unit task pulling mode. In the execution unit task pulling mode, each execution unit will have a problem of competing for access to a public task container. For example, in a scheduling domain, there are multiple execution units, but there is usually only one public task container. In the task pulling mode, when the number of tasks to be executed in the private task container of the execution unit is insufficient, the execution unit will pull tasks from the public task container to its own private task queue. However, the public task container only allows one execution unit to perform task pulling operations at the same time, which inevitably leads to competition for access to the public task container by multiple execution units. When the number of execution units competing for access to the public task container reaches a certain number, the competition access delay will become the main bottleneck, and the task pulling mode is not suitable for execution in a cross-node cluster environment, and the cost of remote access to the public task container is too high. The scheduling unit task issuing mode accesses its own private task container during task scheduling, and there is no competition for access to the public task container, so there is no competition access delay problem, and the scalability is greatly increased.

[0082] The three features include business code affinity hints, execution unit task stealing, and custom scheduling strategy mounting. Among them, the business code affinity hint feature provides a business code to directly specify or hint at suitable execution unit capabilities. In the case where the business programming personnel has an understanding of the load and parallelism of the program, the business code affinity hint is used to assist the runtime in allocating hardware resources. Exemplarily, Figure 3 A business code affinity hint schematic diagram is shown. As Figure 3 shown, the executable tasks to be scheduled in the scheduling domain include N tasks, and the execution units include M. The N tasks to be scheduled are task 1, task 2, …, and task N, and the execution units can be central processing units (CPUs). The M execution units are CPU 0, CPU 1, …, and CPU M. The business code affinity hint gives the following hints: task 1 is recommended to be scheduled for execution on CPU 0, task 2 is recommended to be scheduled for execution on CPU 1, and task N is recommended to be scheduled for execution on CPU M.

[0083] The execution unit task stealing feature refers to a feature or mechanism in which an execution unit steals tasks from the private task containers of other execution units when the number of tasks in its own private task container is below a threshold (for example, the number of tasks to be executed is less than 16). The implementation complexity of this feature is N*(N-1), and the implementation complexity of the task pulling of the execution unit task pulling mode is N, where N represents the number of execution units in the scheduling domain. It can be seen that the complexity of the execution unit task stealing feature is significantly higher than that of the execution unit task pulling mode, and it needs to be used carefully to balance the cost and benefits.

[0084] Custom scheduling policy mounting refers to that a runtime task scheduling framework supports user-defined policies such as execution unit selection and a user-defined runtime scheduling process.

[0085] By Figure 2 It can be known that, in the existing task scheduling platforms, an OMPSS2 platform (referred to as an OMPSS2 platform if the OMPSS2 task scheduling model is adopted) and an FFRT2.0 platform only support an execution unit task pull mode, and neither of them supports the three features; a Legion platform supports a scheduling unit task push mode, and supports the three features; a TaskFlow platform supports the scheduling unit task push mode, supports a business code affinity hint, and supports an execution unit task stealing feature, but does not support a custom scheduling policy mounting feature.

[0086] In summary, the current task scheduling scheme has the following problems:

[0087] 1. The execution unit pull mode and the scheduling unit task push mode each have advantages and disadvantages, and no scheduling model currently supports both modes to complement each other and adapt to different applications and platforms.

[0088] 2. Only the Legion model supports user-defined scheduling policy mounting, but the policy only covers the task scheduling unit push and task stealing functions.

[0089] 3. The Legion does not support partial policy running, and a task must go through the whole process, which is not friendly to small task scenarios, and the task scheduling unit may occupy a large proportion of processing time.

[0090] 4. For special situations such as inability to push and inability to obtain executable tasks caused by total capacity threshold control, only one mechanism can be used for processing, and user processing flexibility cannot be provided.

[0091] Therefore, the embodiments of the present application provide a task scheduling method, which simultaneously supports a scheduling unit task push mode and an execution unit task pull mode, pushes a to-be-scheduled task to each execution unit and a public task container in a scheduling domain through the scheduling unit task push mode, and when the number of to-be-executed tasks in a private queue of an execution unit is insufficient, the execution unit pulls a task from the public task container to a private task container of the execution unit through a task pulling operation, so that the task scheduling method has the advantages of both the scheduling unit task push mode and the execution unit task pull mode, and the performance of task scheduling is improved, for example, high load balancing and high scalability of task scheduling.

[0092] The specific implementation scheme of the task scheduling method provided by the embodiments of the present application is described in detail below with reference to the drawings.

[0093] The task scheduling method provided in the embodiment of the present application can be applied to a task scheduling scenario in a single scheduling domain, for example, to a task scheduling scenario only for the local machine.

[0094] Figure 4 The system architecture diagram of a runtime task scheduling system using the task scheduling method provided by the embodiment of the present application is shown. Figure 4 As shown, the runtime task scheduling system (hereinafter referred to as the runtime system or task scheduling system for the convenience of description) includes a submission unit, a runtime resolution (dependency manager) unit, a scheduling unit and an execution unit (worker), wherein the submission unit is responsible for submitting the task to the runtime resolution unit of the runtime system, the runtime resolution unit is responsible for resolving the dependency of the task, and sending the task to the scheduling unit when the task has no input dependency. The scheduling unit selects a suitable execution unit from the execution units of the local scheduling domain to execute the task, and the execution unit is used to execute the assigned task, and after the task is completed (taskdone), it sends a task completion signal to the runtime resolution unit so that the runtime resolution unit updates the dependency of other tasks, and the task execution process is completed.

[0095] The task scheduling method provided in the embodiments of the present application can also be applied to task scheduling scenarios in multiple scheduling domains, for example, it can be applied to task scheduling scenarios in distributed clusters.

[0096] Figure 5 FIG2 shows a system architecture diagram of another runtime task scheduling system using the task scheduling method provided in the embodiment of the present application. Figure 5 As shown, the runtime system includes a submission unit, a runtime parsing unit, a scheduling unit, and an execution unit. The submission unit is responsible for submitting tasks to the runtime parsing unit of the runtime system. The runtime parsing unit is responsible for parsing the dependencies of tasks and sending tasks to the scheduling unit when the task has no input dependencies. The scheduling unit first decides whether the task is to be executed in the local scheduling domain or in the remote scheduling domain. If it is decided to be executed in the local scheduling domain, one or more execution units are selected from the execution units of the local scheduling domain to execute the task. The execution unit is used to execute the assigned task and, after the task is executed, sends a signal indicating that the task is completed to the runtime parsing unit so that the runtime parsing unit updates the dependencies of other tasks, and the task execution process is completed. If it is decided that the task is to be executed in the remote scheduling domain, the task is sent to the scheduling unit of the remote scheduling domain. The scheduling unit in the remote scheduling domain selects a suitable execution unit from the execution units of the remote scheduling domain to execute the task. The execution unit executes the task and, after the task is executed, sends a signal indicating that the task is completed to the runtime parsing unit of the local scheduling domain to update the dependencies of other tasks.

[0097] The task scheduling method provided by the embodiments of the present application mainly improves the scheduling unit in the runtime task scheduling system, so that the runtime task scheduling system can support both task push mode and task pull mode, ensure that the task scheduling has the advantages of both modes, and improve the performance of task scheduling.

[0098] The runtime system using the task scheduling method provided by the embodiments of the present application can be applied in a cluster environment, and a runtime instance is deployed in a node of the cluster to implement task scheduling in the cluster environment.

[0099] Figure 6 A deployment manner of deploying the runtime system provided by the embodiments of the present application as a running instance in a cluster is shown. Each node in the cluster can deploy one or more runtime instances, and each runtime instance corresponds to a scheduling domain, that is, each runtime instance is responsible for task scheduling in a scheduling domain. As shown in Figure 6 The cluster includes node 0, node 1 and node 2, runtime instance 0 is deployed in node 0, runtime instance 1 is deployed in node 1, runtime instance 2 and runtime instance 3 are deployed in node 2. Runtime instance 1 is responsible for task scheduling in the entire node 0, runtime instance 1 is responsible for task scheduling in the entire node 1, and multiple execution units in node 2 are divided into two scheduling domains, including a first scheduling domain and a second scheduling domain (for example, node 2 has 32 execution units, of which 16 are divided into the first scheduling domain, and the other 16 are divided into the second scheduling domain), runtime instance 2 is responsible for task scheduling in the first scheduling domain, and runtime instance 3 is responsible for task scheduling in the second scheduling domain. The runtime instance implements the task scheduling logic provided by the embodiments of the present application and is responsible for task scheduling across scheduling domains and local scheduling domains.

[0100] It is easy to understand that the scheduling domain refers to a set of hardware resources specified by a user and managed by a running management, which can include but is not limited to a specified number of CPU cores, CPUs, NPUs, GPUs, etc. It can also be referred to as a set of hardware resources related to task execution, such as CPU cores, CPUs, NPUs, GPUs, etc. For example, Figure 6 In the above, node 0 has a runtime instance 0 deployed therein, and the runtime instance 0 is responsible for managing a set of hardware resources related to task execution in node 0. The set of hardware resources related to task execution in node 0 can be referred to as a scheduling domain. For example, node 0 has 16 CPUs, and the scheduling domain includes the 16 CPUs in node 0. The runtime instance can schedule an executable task to one or more CPUs in the 16 CPUs for execution.

[0101] It is worth noting that, Figure 6 Only one possible example of deploying the runtime instance provided by the embodiments of the present application in a cluster environment is shown,Figure 6 The number of clusters and the deployment method of runtime instances in the embodiment of the present application do not constitute a limitation. For example, the cluster may include fewer or more nodes, and one or more runtime instances may be deployed in a node, such as deploying two runtime instances in node 1 or deploying one runtime instance in node 2.

[0102] Figure 7 A schematic diagram of a scheduling mechanism of the task scheduling method provided in the embodiment of the present application is shown. Figure 7 As shown, the scheduling unit dispatches multiple executable tasks to be scheduled to the private task containers and public task containers of multiple execution units. Multiple execution units respectively execute the tasks in their own private task containers. When the number of tasks in the private task container of an execution unit is less than a preset threshold (for example, less than 16), the execution unit pulls a certain number of tasks from the public task container to its own private task container to ensure that it has enough tasks to execute. If the execution unit still fails to pull a sufficient number of tasks from the public task container, it starts the task stealing operation, steals tasks from the private task containers of other execution units, and adds them to its own private task container. In this way, both the task dispatching mode and the task pulling mode are supported, with both the scalability of the task dispatching mode and the load balancing of the task pulling mode. It provides task scheduling performance.

[0103] Figure 8 The following is a flow chart showing a task scheduling method provided by an embodiment of the present application. Figure 4 or Figure 5 The runtime task scheduling system shown in the figure can be deployed on any device, equipment, platform or equipment cluster with computing capabilities to schedule tasks on it. It supports both task delivery mode and task pull mode, so that task scheduling has the advantages of both modes and provides task scheduling performance. Figure 8 As shown, the task scheduling method provided in the embodiment of the present application includes steps S801 to S807.

[0104] In step S801 , the upper layer application generates a task to be processed.

[0105] The upper-layer application may include one or more applications, and during operation, the one or more applications may request the hardware platform to provide corresponding data processing services, such as services for reading and writing data and services for performing operations on business data, so as to realize the specific functions of the upper-layer application.

[0106] In actual implementation, the upper-layer application can generate a task to be processed. The task can be a single-threaded task, that is, a thread is executed to process the task. Alternatively, the task can be a multi-threaded task, that is, the task is processed by executing multiple threads. In actual application scenarios, multiple threads can be concurrently executed to speed up the processing of the task.

[0107] In actual scenarios, the upper-layer application has multiple applications, and the multiple applications generate a large number of tasks to be processed. In order to speed up the processing of the tasks, the hardware platform usually has multiple execution units that execute the assigned tasks in parallel, thereby reducing the time delay of the upper-layer application for obtaining a feedback result of the task processing.

[0108] Reasonable scheduling of the multiple tasks to be processed generated by the upper-layer application can ensure load balancing of the multiple execution units in the hardware platform and adaptability of the execution units to the execution tasks (for example, the tasks are assigned to the execution units suitable for executing the tasks according to the attributes of the tasks), thereby effectively improving the execution efficiency of the tasks and increasing the overall performance of the system. Therefore, a runtime task scheduling system is arranged between the upper-layer application and the hardware platform, which is responsible for efficiently and reasonably scheduling the tasks to be processed generated by the upper-layer application to the multiple execution units of the hardware platform for execution, thereby ensuring the efficiency of the task execution.

[0109] After the upper-layer application generates a task to be processed, the task is provided to the task scheduling system, so that the task scheduling system schedules the task to a suitable execution unit in the hardware platform for processing.

[0110] It is easy to understand that the hardware platform provides hardware resources required for implementing the task processing, for example, including memory, execution units, and the like, and can implement the processing operation of the task. The execution unit can be understood as a hardware unit that specifically implements the processing task. For example, the execution unit can be a processor, including but not limited to a CPU, a GPU, an NPU, a field-programmable gate array (FPGA), and the like, which can complete processing operations (for example, arithmetic operations and read-write operations). When the processor is a multi-core processor, the execution unit can also be a core in the processor. For a cloud scenario, the execution unit can also be a node, a container, and the like, which can implement the task processing.

[0111] The task scheduling system can be implemented through software, for example, through at least one of a virtual machine, a container, and a computing engine. Alternatively, the task scheduling device can be implemented through the physical device of a processor, wherein the processor can be a CPU, and any processor or any combination thereof, such as an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a complex programmable logical device (CPLD), an FPGA, a generic array logic (GAL), a system on chip (SoC), a software-defined infrastructure (SDI) chip, an artificial intelligence (AI) chip, etc. Alternatively, the task scheduling device can be implemented through software plus hardware, such as the task scheduling device can collect relevant information of the hardware platform at runtime through hardware, and implement task scheduling through software.

[0112] In step S802 , a task set to be scheduled is obtained, where the task set includes N tasks.

[0113] After receiving multiple pending tasks from the upper-level application, the task scheduling system submits the multiple pending tasks to the runtime parsing unit through the submission unit. The runtime parsing unit parses the dependencies between the received multiple pending tasks and then sends the tasks without dependencies to the scheduling unit. The tasks without dependencies constitute the set of tasks to be scheduled.

[0114] For example, the business program code of the submitting unit expresses the task to be submitted through a specific application programming interface (API), and the implementation code is as follows:

[0115]

[0116] It is easy to understand that a task without dependencies means that the input of the task does not depend on the output of other tasks. Tasks without dependencies can be automatically parsed out by the runtime parsing unit. Tasks without dependencies are executable tasks, that is, tasks to be scheduled.

[0117] For example, the plurality of to-be-processed tasks of the upper application received by the task scheduling system include task 1, task 2, task 3, task 4, task 5, task 6, task 7 and task 8. After the analysis of the runtime analysis unit, the inputs of task 1, task 3, task 5 and task 8 do not depend on the outputs of other tasks, and task 1, task 3, task 5 and task 8 are tasks without dependency relationship and can be used as a to-be-scheduled task set for once scheduling. The input of task 2 depends on the output of task 1, the input of task 4 depends on the output of task 3, the input of task 6 depends on the output of task 5, and the input of task 7 depends on the output of task 8.

[0118] Figure 9 A schematic diagram of the task dependency relationship between a plurality of to-be-processed tasks after the dependency relationship analysis by the runtime analysis unit is shown.

[0119] As shown in Figure 9 The runtime analysis unit analyzes the dependency of the tasks submitted by the submission unit and labels the dependency relationship between the tasks by using, for example, an execution flow diagram.

[0120] The runtime system identifies the tasks that do not depend on other tasks in the dependency relationship description, for example, after the completion of the startup of the runtime system, Figure 10 The tasks 1, 3, 5 and 8 in the middle box are tasks without dependency relationship and can be scheduled and executed. The runtime analysis unit sends the task set that can be scheduled and executed this time, i.e., the tasks 1, 3, 5 and 8 to the scheduling unit, so that the scheduling unit schedules the tasks 1, 3, 5 and 8.

[0121] In step S803, the N tasks are scheduled to the M private task containers and the public task container corresponding to the M execution units of the local scheduling domain.

[0122] The scheduling unit schedules the executable task set received, for example, the tasks 1, 3, 5 and 8, and schedules them to the private task containers and the public task container of the M execution units of the local scheduling domain.

[0123] Optionally, the runtime system can provide a user-defined policy mounting point (also referred to as a mounting interface) for the user, support the mounting of the user-defined scheduling-related policy, and enable the user to mount the user-defined scheduling policy, for example, an execution unit selection policy, through the policy mounting point. The execution unit selection policy is used to select a suitable execution unit in the local scheduling domain for the executable task to execute the task.

[0124] The input of the execution unit selection strategy can be the task information of the executable task received by the scheduling unit, the hardware information of the hardware platform and the local scheduling domain status information, and the output is the execution unit corresponding to the executable task. That is to say, the execution unit selection strategy is used to select a suitable execution unit for the executable task to execute the executable task.

[0125] In one example, task information can be information about the type of executable task; hardware platform hardware information can be some hardware description information on the hardware platform where the task is currently located, such as the number of CPUs the hardware platform has, the number of cores the CPU has, the memory size of the hardware platform, and other hardware description information about the hardware platform; local scheduling domain status information can be status information of each hardware unit in the local scheduling domain, such as the status information of each execution unit in the local scheduling domain and network latency information, and other hardware status information in the local scheduling domain that affects task execution. For example, if the local scheduling domain has 8 CPUs, the local scheduling domain status information is the occupancy rate of these 8 CPUs. It can be seen from this that the hardware platform hardware information is static information that does not change, while the local scheduling domain status information is dynamic information that changes as the program runs. For example, as the number of executed tasks increases, the CPU occupancy rate in the scheduling domain will increase.

[0126] Take the executable task 1 as an example. The hardware platform where task 1 is currently located is Figure 6 On node 0 shown in the figure, the runtime system obtains that the task type of task 1 is double-precision floating-point arithmetic; the hardware information of node 0 indicates that it has 16 executable units and 8GB of memory; and the local scheduling domain status indicates the occupancy rate of each of the 16 executable units. The runtime system uses the obtained task type information of task 1, the hardware platform information, and the local scheduling domain status information as input to the execution unit selection policy. The execution unit selection policy outputs the execution unit ID selected for task 1 based on this input information. For example, if the execution unit ID output by the execution unit selection policy is execution unit 1, the scheduling unit schedules task 1 to the private task container of execution unit 1 based on the execution unit ID output by the execution unit selection policy.

[0127] Of course, the input information of the execution unit selection strategy described above is a possible example and does not constitute a limitation on the embodiments of the present application. Users can also define other input information as needed. For example, the input of the execution unit selection strategy is defined to include task attribute information and local scheduling domain status information.

[0128] Of course, it is not ruled out that the execution unit selection strategy may also output an incorrect execution unit ID. For example, there are 16 executable units in the local scheduling domain, execution unit 1 to execution unit 16, but the execution unit ID output by the execution unit selection strategy is execution unit 17, that is, the execution unit ID output by the execution unit selection strategy does not exist in this scheduling domain, so the task will be scheduled to the public task container.

[0129] Therefore, when the execution unit selection strategy does not output a valid execution unit ID (for example, the output execution unit ID does not exist in this scheduling domain or no execution unit ID is output or an error code is output), the task is scheduled to the public task container.

[0130] The tasks in the executable task set each perform operations similar to those described above for Task 1. That is, the task information of Task 1, Task 3, Task 5, and Task 8, the hardware information of the hardware platform, and the local scheduling domain state information are used as inputs to the execution unit selection policy, and the execution unit corresponding to Task 1, Task 3, Task 5, and Task 8 are output, respectively. Alternatively, all tasks in the executable task set simultaneously perform operations similar to those described above for Task 1. That is, the task information of Task 1, Task 3, Task 5, and Task 8, the hardware information of the hardware platform, and the local scheduling domain state information are used as inputs to the execution unit selection policy. The output of the execution unit selection policy includes the execution unit corresponding to Task 1, Task 3, Task 5, and Task 8. For example, the execution unit corresponding to Task 1 and Task 3 is Execution Unit 1, and the execution unit corresponding to Task 5 and Task 8 is Execution Unit 2. Based on the output of the execution unit selection policy, the scheduling unit schedules Task 1 and Task 3 to the private container of Execution Unit 1, and schedules Task 5 and Task 8 to the private task container of Execution Unit 2.

[0131] Figure 11 FIG. 1 shows a schematic diagram of a scheduling unit scheduling executable tasks. Figure 11 As shown, the scheduling unit schedules tasks 1 and 3 to the private container of execution unit 1 and schedules tasks 5 and 8 to the private task container of execution unit 2 according to the output result of the execution unit selection strategy.

[0132] Execution unit 1 and execution unit 2 execute the tasks in their own private containers separately, executing task 1, task 3, task 5 and task 8. After execution, they return an execution completion signal to the task scheduling system, so that the runtime parsing unit of the task scheduling system updates the task dependencies.

[0133] Figure 12 FIG. 4 shows a schematic diagram of the updated task dependency relationship in the executable task set. Figure 12As shown, after the execution of task 1, task 3, task 5 and task 8 is completed, the dependency relationship of task 2, task 4, task 6 and task 7 is released, and task 2, task 4, task 6 and task 7 are executable tasks, which constitute a new executable task set, and the scheduling unit continues to schedule the new executable task set.

[0134] In another example, for a task scheduling scenario of multiple scheduling domains, the task scheduling system further provides a policy mounting point of a scheduling domain selection policy, supports a user to mount a custom scheduling domain selection policy to the task scheduling system, and before selecting an execution unit for an executable task, the scheduling domain selection policy is used to select a scheduling domain for the executable task, to determine whether to execute in a local scheduling domain or a remote scheduling domain.

[0135] The input of the scheduling domain selection policy can be task information of an executable task received by a scheduling unit, hardware information of a hardware platform corresponding to each scheduling domain, and state information of each scheduling domain, and the output is a target scheduling domain corresponding to the executable task, that is, a suitable target scheduling domain is selected for the executable task, and the executable task is executed in the target scheduling domain.

[0136] In one example, the task information can be type information of the executable task; the hardware information of the hardware platform corresponding to each scheduling domain can be some hardware description information on a hardware platform corresponding to each scheduling domain in multiple scheduling domains, for example, how many CPUs a hardware platform corresponding to each scheduling domain has, how many cores a CPU has, how large the memory of the hardware platform is, and other hardware description information of the hardware platform; and the state information of each scheduling domain can be state information of each hardware unit in each scheduling domain, such as state information and network delay information of each execution unit in each scheduling domain and other hardware state information affecting task execution in each scheduling domain.

[0137] Taking the executable task as task 1 for example, task 1 is currently in Figure 6 As shown in the cluster environment, it includes scheduling domain 0 corresponding to node 0, scheduling domain 1 corresponding to node 1, two scheduling domains 2 and 3 corresponding to node 2, and therefore, task 1 can be executed in scheduling domain 0, scheduling domain 1, scheduling domain 2 and scheduling domain 3, in other words, the task scheduling system can schedule task 1 to any of scheduling domain 0, scheduling domain 1, scheduling domain 2 and scheduling domain 3.

[0138] The runtime system obtains that the task type of the task 1 is a floating point operation of double precision; the hardware information of each scheduling domain is as follows: the hardware information of the scheduling domain 0 is that there are 16 executable units and the memory is 8G, the hardware information of the scheduling domain 1 is that there are 32 executable units and the memory is 16G, the hardware information of the scheduling domain 2 is that there are 32 executable units and the memory is 8G, and the hardware information of the scheduling domain 3 is that there are 64 executable units and the memory is 32G; the state of each scheduling domain is as follows: the occupancy rate of each executable unit in each scheduling domain. The runtime system takes the task type information of the task 1, the hardware information of each scheduling domain and the state information of each scheduling domain as the input of the scheduling domain selection strategy, and the scheduling domain selection strategy outputs the ID of the target scheduling domain corresponding to the task 1 according to the input information. For example, if the target scheduling domain ID output by the scheduling domain selection strategy is the scheduling domain 0, the scheduling unit schedules the task 1 to the scheduling domain 0 according to the scheduling domain ID output by the scheduling domain selection strategy, that is, the task 1 is executed in the local scheduling domain.

[0139] In another example, in order to assist the scheduling domain selection strategy to better select a suitable scheduling domain and / or the execution unit selection strategy to select a suitable execution unit, the task scheduling system further provides an affinity description analysis strategy mounting point, supports a user to mount a self-defined affinity description analysis strategy to the task scheduling system, and determines the affinity hint of the executable task by using the affinity description analysis strategy before selecting the target scheduling domain and the execution unit for the executable task.

[0140] The input of the affinity description analysis strategy can be the affinity description of the executable task and the hardware information of the hardware platform, and the output is the mapping relationship of the affinity description to the hardware resource.

[0141] For example, the task scheduling system obtains the affinity description of the executable task and the hardware information of each execution unit in the plurality of scheduling domains; takes the affinity description of the executable task and the hardware information of each execution unit in the plurality of scheduling domains as the input of the affinity description analysis strategy, and outputs the hint information corresponding to the executable task, which is used to prompt the assigned execution unit of the executable task.

[0142] The hint information corresponding to the executable task can be taken as the input of the scheduling domain selection strategy to assist the scheduling domain selection strategy to select a suitable scheduling domain for the executable task, and / or the hint information corresponding to the executable task can be taken as the input of the execution unit selection strategy to assist the execution unit selection strategy to select a suitable execution unit for the executable task.

[0143] Optionally, the affinity description can be an implicit affinity description or an explicit affinity description.

[0144] Figure 13A diagram showing an implicit affinity description and a mapping relationship between the affinity description and hardware is shown.

[0145] Figure 14 A diagram showing an explicit affinity description and a mapping relationship between the affinity description and hardware is shown.

[0146] In step S804, the tasks in each of the M private task containers are dispatched to the private task queue of the execution unit corresponding to each private task container.

[0147] After the scheduling unit schedules the tasks in the executable task set to the private task containers of the execution unit through the above steps, the executable tasks in the private task containers are dispatched to the private task queue of the execution unit, so that the executable tasks are executed by the execution unit.

[0148] The executable tasks in the private containers can be dispatched to the private queue of the executable unit through a certain task dispatching strategy. The dispatching order determines the ordering of the executable tasks in the private queue, and the execution unit executes the executable tasks in the private queue according to the task ordering.

[0149] The task dispatching strategy can be a default task dispatching strategy provided by the task scheduling system, for example, the default dispatching strategy is to sort and dispatch the executable tasks in the private task container according to the priority of the tasks, or to sort and dispatch the executable tasks in the private task container according to the real-time requirement level of the tasks, or to directly randomly select a certain number of tasks for dispatching, etc. The task dispatching strategy can also be a user-defined task dispatching strategy. That is, the task scheduling system provides a mounting point for the user-defined task dispatching strategy for the user, supports the mounting of the user-defined task dispatching strategy, and realizes the dispatching of the executable tasks in the private task container through the user-defined task dispatching strategy.

[0150] For example, the input of the user-defined task dispatching strategy can be the information of the private task container, the hardware information of the hardware platform and the local scheduling domain state information, and the output is a task list including the tasks dispatched from the private task container and the task ordering.

[0151] Specifically, the information of the private task container can be task information in the private task container, for example, the private task container 1 of the execution unit 1 has tasks 1, 2, 3, 4, 5 and 6, and the information of the private task container is related description information of executable tasks in the container, such as priority of the task, real-time requirement of the task and type of the task, and the hardware information of the hardware platform can be hardware information of the execution unit, such as memory information corresponding to the execution unit, type of the execution unit, computing power of the execution unit and other hardware description information, and the local scheduling domain state information can be state information of each hardware unit in the local scheduling domain, such as available memory in the local scheduling domain and execution unit occupancy.

[0152] After the task scheduling system obtains the information of the private task container, the hardware information of the hardware platform and the local scheduling domain state information, it takes them as inputs of the task distribution strategy, and the task distribution strategy outputs a task list, for example, the output task list is shown in the following table:

[0153] Task 2 Task 1 Task 4 Task 5 Task 6

[0154] The above task list indicates that the executable tasks in the private task container 1 in this distribution are tasks 1, 2, 4, 5 and 6, and the order is: task 2, task 1, task 4, task 5 and task 6, and the order of the private task queue of the execution unit 1 is also task 2, task 1, task 4, task 5 and task 6. The execution unit performs tasks in the order of task 2, task 1, task 4, task 5 and task 6.

[0155] Through the task distribution strategy, the order of tasks executed by the execution unit is controlled, for example, tasks with higher priority can be executed first, and tasks with lower priority can be executed later; tasks more suitable for the current memory situation are executed first, for example, the available space of the current memory is 5MB, the memory occupancy of task 2 is 4.9MB, the memory occupancy of task 1 is 5.6MB, and the priority of task 2 is higher than that of task 1, so task 2 is arranged before task 1 and task 2 is executed first.

[0156] In step S805, the execution unit executes tasks according to the order of tasks in the private task queue.

[0157] The execution unit executes executable tasks in order according to the order of tasks in the respective private task queue.

[0158] In step S806, when the number of tasks in the private task queue of the execution unit is insufficient, a task pulling request is generated.

[0159] When the number of tasks in the private task queue of the execution unit is insufficient, for example, the number of tasks in the private task queue of the execution unit is less than or equal to a set threshold value (for example, 16), a task pulling request is generated, and the task pulling request is sent to the scheduling unit to enable the scheduling unit to pull tasks to the private task queue of the execution unit.

[0160] For example, the number of tasks in the private task queue of the execution unit 1 in an execution state is 16, which is equal to the set threshold value 16, and the execution unit 1 generates a task pulling request, which is sent to the scheduling unit to enable the scheduling unit to pull a certain number of tasks from the private task queue or the public task queue.

[0161] Optionally, the task pulling request carries the number of tasks to be pulled, for example, the number of tasks in the private task queue of the execution unit 1 in an execution state is 16, and the capacity of the private task queue is 64, so the task pulling request carries the number of tasks to be pulled, which is 48. That is, the number of tasks requested to be pulled by the scheduling unit is 48.

[0162] In step S807, in response to the task pulling request of the execution unit, tasks are selected from the public task container and added to the private task container of the execution unit.

[0163] Taking the task pulling of the execution unit 1 as an example, the scheduling unit responds to the task pulling request, first performs a task pulling operation from the private task container of the execution unit 1, and if a sufficient number of tasks can be pulled, the execution unit 1 continues to execute the tasks in its private task queue, and if the number of tasks pulled is insufficient, the scheduling unit continues to pull tasks from the public task container.

[0164] For example, the number of tasks in the private task queue of the execution unit 1 in an execution state is 16, which reaches the set threshold value 16, and a task pulling request is generated, which carries the number of tasks to be pulled, which is 48. The execution unit 1 sends the task pulling request to the scheduling unit, the scheduling unit first performs a task pulling operation from the private task container of the execution unit 1, and the number of tasks in the private task container is 24, so the 24 tasks are pulled to the private task queue of the execution unit 1, and then a task pulling operation is performed from the public task container to pull the remaining 24 tasks.

[0165] In one example, the task selected from the public task container can be added to the private task container of the execution unit 1 by a task selection strategy. The task selection strategy can be a default task selection strategy provided by the task scheduling system, for example, the default task selection strategy is to pull tasks from the public task container according to task attributes. The task selection strategy can also be a user-defined task selection strategy. That is, the task scheduling system provides a mounting point for the user-defined task selection strategy, supports the mounting of the user-defined task selection strategy, and realizes the task pulling operation from the public task container by the user-defined task selection strategy.

[0166] For example, the input of the user-defined task pulling strategy can be the information of the public task container and the hardware information of the hardware platform, and the output can be the target task pulled from the public task container and the number of the target task.

[0167] Optionally, the information of the public task container includes the attribute information of each task in the public task container, and the hardware information of the hardware platform includes the hardware information of the execution unit 1. For example, the scheduling unit obtains the attribute information of each task in the public task container and the hardware information of the execution unit 1, then takes the attribute information of each task in the public task container and the hardware information of the execution unit 1 as the input of the task pulling strategy, outputs the target task pulled from the public task container and the number of the target task, and then the scheduling unit performs the task pulling operation from the public task container according to the output of the task pulling strategy, pulls the target task in the public task container to the private task container of the execution unit 1, and then sends the task in the private task container to the private task queue of the execution unit 1 by the task sending strategy, so as to avoid the idle of the execution unit 1 and improve the efficiency of task execution.

[0168] That is, when certain conditions are met (for example, the number of tasks to be executed of the execution unit is insufficient), the execution unit needs to call the function interface provided by the reverse task pulling process of the scheduling unit to perform the reverse task pulling operation to pull tasks for execution, and the code implementation is as follows:

[0169]

[0170] In order to ensure the execution efficiency of the task, the hardware platform often has a large number of execution units, for example, 128 execution units. The 128 execution units pull tasks from the common task container. It is very likely that the tasks in the common task container are pulled empty or the number of tasks remaining is very small, which causes the execution unit to pull insufficient number of tasks from the common task container. For example, execution unit 2 needs to pull 48 tasks from the common task container, but there are only 12 tasks in the common task container, which causes the execution unit 2 to pull insufficient number of tasks from the common task container. In order to avoid this situation, the task scheduling method provided by the embodiment of the application also provides a local task stealing mechanism. When insufficient number of tasks is pulled from the common task container, the task stealing mechanism is started to steal tasks from the private task container of other execution units in the local scheduling domain.

[0171] For example, the scheduling unit can perform the task stealing operation from the private task container of other execution units in the local scheduling domain through the local task stealing strategy provided by the task scheduling system, or perform the local task stealing operation through the user-defined local task stealing strategy. That is, the task scheduling system provides a mounting point for the user-defined local task stealing strategy, supports the mounting of the user-defined local task stealing strategy, and realizes the task stealing operation on the private task container of other execution units in the local scheduling domain through the user-defined local task stealing strategy.

[0172] Optionally, the input of the user-defined local task stealing strategy can be the information of the private task container of other execution units in the local scheduling domain, and the output is the target task to be stolen and the local stealing number. For example, the local scheduling domain includes execution unit 1, execution unit 2, execution unit 3 and execution unit 4. The number of tasks in the private task queue of execution unit 1 is insufficient, and the number of tasks is still insufficient after the task pulling operation is performed. Then the local task stealing operation is continued, the task information in the private task container of execution unit 2, the task information in the private task container of execution unit 3 and the task information in the private task container of execution unit 4 are obtained, and then they are taken as the input of the local task stealing strategy. The target task to be stolen and the local stealing number are output, for example, the target task to be stolen in the private task container of execution unit 2 is task 1 to task 10, the target task to be stolen in the private task container of execution unit 3 is task 11 to task 20, the target task to be stolen in the private task container of execution unit 4 is task 10 to task 37, and the stealing number is 48. Then the scheduling unit selects the target task from the private task container of execution unit 2, execution unit 3 and execution unit 4 according to the output result of the local task stealing strategy and adds the target task to the private task container of execution unit 1, and then distributes the tasks in the private task container to the private task queue of execution unit 1 according to the task distribution strategy.

[0173] Considering that the number of tasks to be executed in the private task queue of the execution unit can still be insufficient after the local task stealing operation is performed from the private task container of other execution units, the embodiment of the present application further provides a remote task stealing mechanism. That is, when the number of tasks stolen through the local task stealing strategy is still insufficient, the remote task stealing mechanism is started to steal tasks from the remote scheduling domain and add them to the execution unit.

[0174] The scheduling unit can steal executable tasks from the remote scheduling domain through the remote task stealing strategy provided by the task scheduling system, or perform a task stealing operation from the remote scheduling domain through a user-defined remote task stealing strategy. That is, the task scheduling system provides a mounting point for the user-defined remote task stealing strategy to the user, supports mounting of the user-defined remote task stealing strategy, and implements the task stealing operation on the remote scheduling domain through the user-defined remote task stealing strategy.

[0175] The input of the user-defined remote task stealing strategy can be hardware information of the remote scheduling domain and state information of the remote scheduling domain, and the output is a stealing target and a stealing number. The stealing target indicates a remote target scheduling domain of the remote task stealing operation and the number of tasks stolen from each remote target scheduling domain. For example, a plurality of scheduling domains responsible for task scheduling of the task scheduling system include scheduling domain 1, scheduling domain 2, scheduling domain 3, and scheduling domain 4. The number of tasks to be executed in the private task queue of execution unit 1 in scheduling domain 1 is insufficient, and after the task pulling operation and the local task stealing operation are performed, there are still 24 tasks missing. Then the remote task stealing operation is continued. Scheduling domain 2, scheduling domain 3, and scheduling domain 4 are remote scheduling domains. The scheduling unit obtains the hardware information and scheduling domain state information of scheduling domain 2, the hardware information and scheduling domain state information of scheduling domain 3, and the hardware information and scheduling domain state information of scheduling domain 4, which are taken as the input of the remote task stealing strategy. The output stealing target is scheduling domain 2, and the stealing number is 10. The output stealing target is scheduling domain 4, and the stealing number is 14. Then the scheduling unit sends a remote task stealing request to scheduling domain 2, which carries the stealing task number 10. The scheduling unit sends a remote task stealing request to scheduling domain 4, which carries the stealing task number 14. Scheduling domain 2 and scheduling domain 4 respectively respond to the remote task stealing request, select the requested number of tasks from their respective public task container or private task container, and return them to the scheduling unit of scheduling domain 1. The scheduling unit receives the 10 tasks returned by scheduling domain 2 and the 14 tasks returned by scheduling domain 4, adds them to the private task container of execution unit 1, and then issues the tasks in the private task container to the private task queue of execution unit 1 according to the task issuing strategy.

[0176] The task scheduling method provided in the application supports a remote task stealing mechanism, so that when the number of tasks stolen from the local is still insufficient, tasks can be stolen from the remote scheduling domain, and the execution unit with insufficient number of tasks to be executed can further steal a sufficient number of tasks.

[0177] In some other embodiments, the task scheduling method provided in the application further includes: obtaining information of an abnormal event, the abnormal event including one or more of the following: the number of tasks to be executed in the local scheduling domain reaches the upper limit of the task capacity, the target scheduling domain is unreachable, the number of tasks to be executed in the target scheduling domain reaches the upper limit of the task capacity, a sufficient number of tasks is not selected from the public task container and the local task stealing strategy is not enabled, a sufficient number of tasks is not stolen by the local task stealing strategy and the remote task stealing strategy is not enabled, and a sufficient number of tasks is not stolen by the remote task stealing strategy; taking the information of the abnormal event as an input of an abnormal event processing strategy (also referred to as an abnormal state processing strategy), and outputting a processing strategy, the processing strategy indicating a processing operation for the abnormal event.

[0178] For example, when the task scheduling system receives a task to be processed generated by an upper application, the number of tasks to be executed in the local scheduling domain where the upper application is located reaches the upper limit of the task capacity, for example, the upper limit of the number of tasks to be executed in the local scheduling domain is 74 million, and the number of tasks to be executed in the local scheduling domain is 74 million, a first system error code (the first system error code indicates an abnormal event that the number of tasks to be executed in the local scheduling domain reaches the upper limit of the task capacity) is sent to the abnormal event processing strategy, the abnormal event processing strategy outputs a processing strategy, for example, scheduling to the remote scheduling domain for execution.

[0179] For another example, when the scheduling domain selected by the scheduling domain selection strategy is the remote scheduling domain, but the task scheduling to the remote scheduling domain fails, a second system error code (the second system error code indicates that the target scheduling domain is unreachable) is sent to the abnormal event processing strategy, and the abnormal event processing strategy outputs a processing strategy, for example, circular waiting or releasing the thread occupation of the task.

[0180] For another example, when the scheduling domain selected by the scheduling domain selection strategy is the remote scheduling domain, but the number of tasks in the remote scheduling domain is full, a third system error code (the third system error code indicates that the number of tasks in the remote scheduling domain is full) is sent to the abnormal event processing strategy, and the abnormal event processing strategy outputs a processing strategy, for example, temporarily suspending the task.

[0181] For another example, when the number of tasks in the private task queue of the execution unit is insufficient, and the sufficient number of tasks is not selected from the public task container and the local task stealing strategy is not enabled, a fourth system error code (indicating that the number of tasks in the execution unit is insufficient, the sufficient number of tasks is not selected from the public task container, and the local task stealing strategy is not enabled) is sent to the exception event processing strategy, and the exception event processing strategy outputs a processing strategy, such as enabling the local stealing strategy.

[0182] For another example, when the number of tasks in the private task queue of the execution unit is insufficient, and the sufficient number of tasks is not stolen by the local task stealing strategy and the remote task stealing strategy is not enabled, a fifth system error code (indicating that the number of tasks in the execution unit is insufficient, the sufficient number of tasks is not selected from the public task container, the sufficient number of tasks is not stolen by the local task stealing strategy, and the remote task stealing strategy is not enabled) is sent to the exception event processing strategy, and the exception event processing strategy outputs a processing strategy, such as enabling the remote stealing strategy.

[0183] For another example, when the number of tasks in the private task queue of the execution unit is insufficient, and the sufficient number of tasks is not stolen by the remote task stealing strategy, a sixth system error code (indicating that the number of tasks in the execution unit is insufficient, the sufficient number of tasks is not selected from the public task container, the sufficient number of tasks is not stolen by the local task stealing strategy, and the sufficient number of tasks is not stolen by the remote task stealing strategy) is sent to the exception event processing strategy, and the exception event processing strategy outputs a processing strategy, such as issuing a task shortage reminder in the execution unit, or the execution unit entering sleep after completing the task.

[0184] The task scheduling method provided by the embodiments of the present application sets corresponding processing measures in advance for possible abnormal situations in the task scheduling process, and ensures the smooth progress of the task scheduling.

[0185] In one example, the exception event processing strategy can be provided by the task scheduling system, or can be a user-defined exception event processing strategy. That is, the task scheduling system provides a mounting point for the user-defined exception event processing strategy to the user, supports the mounting of the user-defined exception event processing strategy, and realizes the processing of the abnormal event in the task scheduling system through the user-defined exception event processing strategy.

[0186] In another possible implementation, one or more of the execution unit selection strategy, the scheduling domain selection strategy, the first task issuing strategy, the task selection strategy, the second task issuing strategy, the local task stealing strategy, the remote task stealing strategy, and the exception event processing strategy can be configured to be enabled or disabled. In this way, partial strategy running is supported, the user can disable unnecessary steps in the process, reduce the task scheduling overhead, and improve the task execution efficiency.

[0187] In some other embodiments, the task scheduling system supports a user to configure the task capacity size of the private task container of the execution unit and the task capacity size of the public task container, so that the user can switch the task scheduling system between supporting the task push and pull mode, supporting only the task push mode and supporting only the task pull mode by changing the task capacity size of the private task container and the task capacity size of the public task container.

[0188] For example, the user switches the task scheduling system to support only the task pull mode by configuring the task capacity of the private task container to 0, i.e. removing the private task container of the execution unit.

[0189] Figure 15 A schematic diagram of the task scheduling mechanism of the task scheduling system after the task capacity of the private task container is configured to 0 is shown. As shown in the figure, the scheduling unit directly schedules the executable tasks to be scheduled to the public task container, and each execution unit in the scheduling domain performs a task pull operation from the public task container as needed, and the execution unit executes the pulled task to ensure load balancing of each execution unit in the scheduling domain. Figure 15

[0190] For another example, the user switches the task scheduling system to support only the task push mode by configuring the task capacity of the public task container to 0, i.e. removing the public task container.

[0191] Figure 16 A schematic diagram of the task scheduling mechanism of the task scheduling system after the task capacity of the public task container is configured to 0 is shown. As shown in the figure, the scheduling unit schedules the executable tasks to be scheduled to the private task container of each execution unit, and each execution unit in the scheduling domain executes the task in the private task container thereof. When the number of tasks in the private task container of the execution unit is insufficient, a task stealing operation can be performed to steal tasks from the private task container of other execution units to avoid idle of the execution unit. Figure 16 The task scheduling method provided by the embodiments of the present application supports a user to change the configuration of the private task container and the configuration of the public task container to switch the task scheduling system between supporting the task push and pull mode, supporting only the task push mode and supporting only the task pull mode, so that the task scheduling system can adapt to different platforms or applications according to actual needs.

[0192]

[0193] ​​The policy mounting point provided by the task scheduling system implemented in the application includes one or more of the execution unit selection policy, the scheduling domain selection policy, the first task issuing policy, the task selection policy, the second task issuing policy, the local task stealing policy, the remote task stealing policy, and the abnormal event processing policy mentioned above. The input and output of each policy mounting point of the task scheduling system are defined to conform to the user-defined function or other feasible code defined by the task scheduling system. The input is provided by the task scheduling system, and the output is processed by the customized policy on the input data. The output data has an impact on the decision of scheduling, for example, the execution unit selection policy outputs the execution unit ID, and the task scheduling system issues the task to the corresponding execution unit.

[0194] The code implementation of the policy mounting point provided by the task scheduling system to implement the policy mounting is introduced below.

[0195] 1. Policy input and output (input output, IO) setting: the input and output of each policy are defined, and the code implementation is as follows:

[0196]

[0197]

[0198] 2. Policy function implementation: in combination with the defined IO, the function code of the run part is implemented as follows:

[0199]

[0200] 3. Policy mounting: the customized policy function or other form of function code is registered to the runtime task scheduling system, and multiple implementations of the same policy can be flexibly replaced, and the implementation code is as follows:

[0201]

[0202]

[0203] 4. Policy execution: the runtime task scheduling system will automatically execute the mounted policy in the task scheduling process (such as task issuing, task pulling, and task stealing, etc.), and the code implementation is as follows:

[0204]

[0205]

[0206] The task scheduling system of the application embodiment supports the mounting of user-defined policies, and can realize flexible optimization of task scheduling for different businesses to ensure the execution efficiency of various applications.

[0207] A specific implementation of the task scheduling method provided by the embodiments of the present application in actual application is introduced below. A multi-scheduling domain scenario is taken as an example for illustration.

[0208] Figure 17 An implementation flowchart of a task issuing stage of the task scheduling method provided by the embodiments of the present application is shown. As shown in Figure 17 the task issuing process includes the following steps:

[0209] Step S11: A locally deployed runtime task scheduling system receives an executable task. A scheduling unit first determines whether the number of tasks to be executed reaches a system capacity upper limit. For example, a local scheduling domain has 64 execution units, each of which can accommodate 1 million tasks in a private task container and 10 million tasks in a public task container. Therefore, the system capacity upper limit of the local scheduling domain is 10 million+64*1 million=74 million. If the number of tasks to be executed in the local scheduling domain reaches 74 million, it is determined that the system capacity upper limit is reached, otherwise, the system capacity upper limit is not reached.

[0210] If the scheduling unit determines that the number of executable tasks in the local scheduling domain reaches the system capacity upper limit, a user-defined exception handling strategy is called for exception handling. If not, the next step is entered.

[0211] Step S12: It is determined whether a user-defined affinity description analysis strategy (affinityMappingPlolicy) is activated. If activated, the affinity description (affinityMap) of the task is analyzed according to the strategy, and an affinity prompt is generated to assist the scheduling decision of steps S13 / S14.

[0212] Step S13: It is determined whether a scheduling domain selection strategy is activated by the user. If not, it is directly entered into step S14. If activated, the strategy is run to determine whether the task is sent to other scheduling domains according to the output of the strategy. The task sent to other scheduling domains is not repeated in this step by default (to avoid forming a loop by repeatedly transmitting the same task). If the remote transmission task cannot be delivered or the number of remote tasks reaches a threshold value (i.e., the number of tasks to be executed in the remote scheduling domain reaches an upper limit value, for example, 74 million), an exception handling strategy is called.

[0213] Step S14: An execution unit for executing the task is selected according to a user-defined execution unit selection strategy, and the task is stored in the private task container of the execution unit. If the execution unit selection strategy does not output a valid execution unit ID or the capacity of the private task container of the execution unit is configured as 0, the task is directly added to the public task container.

[0214] Step S15: According to the user-defined task distribution strategy, the tasks are sorted in the private task container of the execution unit or a certain number of tasks are directly selected and distributed to the private task queue of the execution unit.

[0215] Step S16: The execution unit executes the executable tasks in the private task queue thereof.

[0216] Of course, in some other embodiments, the various task scheduling related strategies in the remote scheduling domain can also be configured to be started, for example, one or more of the affinity description analysis strategy, the scheduling domain selection strategy, and the exception state processing strategy in the remote scheduling domain can be configured to be in the started state, and the tasks distributed from other scheduling domains continue to be processed by the user-defined strategy.

[0217] Figure 18 An implementation flowchart of the task pulling phase of the task scheduling method provided by the embodiment of the application is shown. As shown in Figure 18 The task distribution process includes the following steps:

[0218] Step S21: When the number of tasks waiting to be executed by an execution unit is less than a threshold value (for example, 16), the scheduling unit checks whether there are enough tasks in the local private task container of the execution unit, and if yes, directly distributes the tasks to the private task queue of the execution unit to enter a waiting execution state, and exits the flowchart. If not or the number of tasks is insufficient, step S22 is entered.

[0219] Step S22: According to the user-defined local task selection strategy, a certain number of tasks are selected from the local shared task container and added to the private task container of the execution unit that initiates the pulling flowchart. When enough tasks can be obtained, the flowchart ends. When not, whether the task stealing function is started is determined, and the next operation is performed, for example, if not started, the exception processing strategy is entered, and if started, step S23 is entered.

[0220] Step S23: According to the user-defined local task stealing strategy, a certain number of tasks are selected from the private task container of other execution units in the local execution unit and added to the private task container of the execution unit that initiates the pulling flowchart. When enough tasks can be obtained, the flowchart ends. When not, whether the remote task stealing function is started is determined, and the next operation is performed, for example, if not started, the exception processing strategy is entered, and if started, step S24 is entered.

[0221] Step S24: According to the user-defined remote task stealing strategy, one or more are selected from other scheduling domains, a remote task stealing request is sent to the selected scheduling domain, and the returned tasks are waited for, for example, if enough tasks are obtained, the tasks are added to the private task container of the initiator (i.e., the execution unit). If there are not enough tasks (the number of tasks waiting to be executed by the execution unit is less than the threshold value), the exception processing strategy is entered.

[0222] Step S25: The tasks are sorted in the private task container of the execution unit, or a certain number of tasks are directly selected and issued to the private task queue of the execution unit.

[0223] Step S26: The execution unit executes the executable tasks in the private task queue thereof.

[0224] The task scheduling method provided in the present application simultaneously supports the task issuing mode, the task pulling mode and the task stealing mechanism, realizes the complementary advantages of the two modes, improves the performance of task scheduling, and can activate one or more of them by configuring the size of the private task container, the size of the public task container and the opening or closing of the task stealing strategy, thereby guaranteeing the execution efficiency of various applications. In addition, a strategy mounting point is provided externally, the mounting of the user-defined strategy is supported, the scheduling scheme can be flexibly optimized for different businesses, and by supporting the operation of part of the mechanisms or strategies, the user can close the steps that are not needed in the process, reduce the scheduling overhead, and improve the task execution efficiency.

[0225] Based on the same idea as the foregoing embodiment of the task scheduling method, the present embodiment further provides a task scheduling device 1900, which can be deployed in a scheduling unit in a runtime task scheduling system, and realizes the task scheduling method provided in the present embodiment, so as to simultaneously support the task issuing mode and the task pulling mode, realize the complementary advantages of the two modes, and improve the performance of task scheduling. The task scheduling device 1900 includes units or modules for realizing each step in the task scheduling method shown in the present embodiment. Figures 8-18 The task scheduling device 1900 includes units or modules for realizing each step in the task scheduling method shown in the present embodiment.

[0226] Figure 19 A structural diagram of a task scheduling device provided in the present embodiment. As shown in the figure, the task scheduling device 1900 includes a task receiving unit 1910, a task sorting unit 1920, a task issuing unit 1930, a task pulling unit 1940, a task stealing unit 1950, and a strategy mounting unit 1960. Figure 19As shown, the task scheduling apparatus 1900 at least includes an obtaining module 1901, a scheduling module 1902, a task issuing module 1903, and a task pulling module 1904. The obtaining module 1901 is configured to obtain a task set to be scheduled, the task set including N tasks, the N tasks all being executable tasks, and N being a positive integer. The scheduling module 1902 is configured to schedule the N tasks to M private task containers and a public task container corresponding to M execution units of a local scheduling domain, and M being a positive integer. The task issuing module 1903 is configured to issue a task in each of the M private task containers to a private task queue of an execution unit corresponding to each private task container, so that the task in the private task queue is executed by the execution unit. The task pulling module 1904 is configured to, in response to a task pulling request of a first execution unit, select a task from the public task container and add the task to a first private task container, the task pulling request being generated when a number of tasks in a private task queue of the first execution unit is less than or equal to a preset threshold, and the first private task container being a private task container corresponding to the first execution unit.

[0227] In one possible implementation, the scheduling module 1902 is specifically configured to: obtain first task information, first hardware information, and first scheduling domain state information, the first task information indicating task type information of each task in the N tasks, the first hardware information indicating hardware information in the local scheduling domain, and the first scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; take the first task information, the first hardware information, and the first scheduling domain state information as inputs of an execution unit selection strategy, and output a target execution unit corresponding to each task; and schedule each task to a private container of the target execution unit corresponding to the task.

[0228] In another possible implementation, for a task for which the execution unit selection strategy does not output a target execution unit, the task is scheduled to a public task container of the local scheduling domain.

[0229] In another possible implementation, the runtime task scheduling system includes a plurality of scheduling domains. The scheduling module 1902 is further configured to: obtain second task information, second hardware information, and second scheduling domain state information, the second task information indicating task type information of each task in the N tasks, the second hardware information indicating hardware information of each scheduling domain in the plurality of scheduling domains, and the second scheduling domain state information indicating state information of each hardware unit in each scheduling domain in the plurality of scheduling domains; take the second task information, the second hardware information, and the second scheduling domain state information as inputs of a scheduling domain selection strategy, and output a first target scheduling domain corresponding to each task; and schedule each task to the target scheduling domain corresponding to the task.

[0230] In another possible implementation, the task scheduling apparatus 1900 provided by the present application further includes an affinity analysis module 1905, configured to acquire affinity descriptions of respective tasks and hardware information of respective execution units in the plurality of scheduling domains; determine prompt information corresponding to the respective tasks based on the affinity descriptions of the respective tasks and the hardware information of the respective execution units in the plurality of scheduling domains, and an affinity description analysis strategy, the prompt information being used to prompt target execution units corresponding to the respective tasks; and the first task information further includes the prompt information corresponding to the respective tasks.

[0231] Optionally, the affinity description includes an explicit affinity description or an implicit affinity description.

[0232] In another possible implementation, the task issuing module 1903 is specifically configured to: acquire information of each private task container, third hardware information and third scheduling domain state information, the information of each private task container indicating task information in each private task container, the third hardware information indicating hardware information corresponding to an execution unit where each private task container is located, and the third scheduling domain state information indicating state information of respective hardware units in the local scheduling domain; take the information of each private task container, the third hardware information and the third scheduling domain state information as inputs of a first task issuing strategy, and output a first task list, the first task list indicating tasks issued in a current time by each private task container and an order of the tasks issued in the current time; and based on the first task list, issue the tasks in each private task container to a private task queue of an execution unit corresponding to each private task container, the execution order of the tasks in the private task queue of the execution unit being the same as the order of the tasks in the task list.

[0233] In another possible implementation, the task pulling module 1904 is specifically configured to: determine that the number of tasks in the first private task container is insufficient; and select tasks from the public task container to join the first private task container.

[0234] In another possible implementation, a specific implementation of selecting tasks from the public task container to join the first private task container includes: acquiring information of the public task container and fourth hardware information, the information of the public task container indicating attribute information of respective tasks in the public task container, and the fourth hardware information indicating hardware information of the first execution unit; taking the information of the public task container and the fourth hardware information as inputs of a task selection strategy, and outputting a target task and a number of the target task, the target task indicating the tasks selected from the public task container; and joining the target task to the first private task container.

[0235] In another possible implementation, the task pulling module 1904 is further configured to: determine that the number of tasks in the first private task container is sufficient before selecting a task from the public task container to join the first private task container; and select a task from the first private task container to join the private task queue of the first execution unit.

[0236] In another possible implementation, a specific implementation of selecting a task from the first private task container to join the private task queue of the execution unit includes: obtaining first private task container information, fifth hardware information, and fourth scheduling domain state information, the first private task container information indicating task information in the private task container of the first execution unit, the fifth hardware information indicating hardware information of the first execution unit, and the fourth scheduling domain state information indicating state information of each hardware unit in the local scheduling domain; taking the first private task container information, the fifth hardware information, and the fourth scheduling domain state information as inputs of a second task issuing strategy, and outputting a second task list, the second task list indicating tasks issued in the current time by the first private task container and an order of the tasks; and issuing the tasks in the first private task container to the private task queue of the execution unit based on the second task list, the execution order of the tasks in the private task queue of the execution unit being the same as the order of the tasks in the second task list.

[0237] In another possible implementation, the task scheduling apparatus 1900 provided in this application further includes a task stealing module 1906 configured to: determine that the number of tasks selected from the public task container is insufficient; obtain second private task container information, the second private task container information indicating task information in other task containers of the M private task containers except the first private task container; take the second private task container information as an input of a local task stealing strategy, and output a target task to be stolen and a first stolen quantity, the target task to be stolen indicating a task to be stolen from the other task containers, and the first stolen quantity indicating the number of tasks to be stolen from the other task containers; perform a task stealing operation from the other task containers based on the target task to be stolen and the stolen quantity; and add the tasks stolen from the other task containers to the first private task container.

[0238] In another possible implementation, the task stealing module 1906 is further configured to: when it is determined that the number of tasks stolen by the local task stealing strategy is insufficient, acquire sixth hardware information and fifth scheduling domain state information, the sixth hardware information indicating hardware information of scheduling domains other than the local scheduling domain in the plurality of scheduling domains, and the fifth scheduling domain state information indicating state information of each hardware unit in the other scheduling domains; take the sixth hardware information and the fifth scheduling domain state information as inputs of a remote task stealing strategy, and output a stealing target and a second stealing number, the stealing target indicating a second target scheduling domain in the other scheduling domains that is to be subjected to a task stealing operation, and the second stealing number indicating a number of tasks stolen from the stealing target scheduling domain; perform the task stealing operation from the second target scheduling domain based on the stealing target and the stealing number; and add the tasks stolen from the second target scheduling domain to the first private task container.

[0239] In another possible implementation, the task scheduling apparatus 1900 provided in the present application further includes an exception event processing module 1907, which is specifically configured to acquire information of an exception event, the exception event including one or more of the following: a task capacity upper limit of a task to be executed in the local scheduling domain is reached, a target scheduling domain is unreachable, a task capacity upper limit of a task to be executed in the target scheduling domain is reached, a sufficient number of tasks is not selected from the common task container and the local task stealing strategy is not enabled, a sufficient number of tasks is not stolen by the local task stealing strategy and the remote task stealing strategy is not enabled, and a sufficient number of tasks is not stolen by the remote task stealing strategy; take the information of the exception event as an input of an exception event processing strategy, and output a processing strategy, the processing strategy indicating a processing operation for the exception event.

[0240] In another possible implementation, one or more of the execution unit selection strategy, the scheduling domain selection strategy, the first task distribution strategy, the task selection strategy, the second task distribution strategy, the local task stealing strategy, the remote task stealing strategy, and the exception event processing strategy can be configured to be enabled or disabled.

[0241] In another possible implementation, the task capacity of each private task container in the M private task containers can be configured to be 0, or the task capacity of the common task container can be configured to be 0.

[0242] The task scheduling apparatus 1900 according to the embodiments of the present application can correspond to performing the methods described in the embodiments of the present application, and the above and other operations and / or functions of each module in the task scheduling apparatus 1900 are respectively for implementing the corresponding procedures of each method in the embodiments of the present application, which will not be described herein again for simplicity. Figures 8-18

[0243] ​The present application also provides a computing device comprising at least one processor, a memory, and a communication interface, wherein the processor is configured to execute Figures 8-18 The method described.

[0244] Figure 20 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.

[0245] like Figure 20 As shown, the computing device 2000 includes at least one processor 2001, a memory 2002 and a communication interface 2003. The processor 2001, the memory 2002 and the communication interface 2003 are communicatively connected, and the communication connection can be achieved through a wired manner (such as a bus) or a wireless manner. The communication interface 2003 is used to send and / or receive data sent by other devices; the memory 2002 stores computer instructions, and the processor 2001 executes the computer instructions and executes the recommended method in the aforementioned method embodiment to achieve simultaneous support for the task issuing mode and the task pulling mode, realize the complementary advantages of the two modes, and improve the performance of task scheduling.

[0246] It should be understood that in the embodiment of the present application, the processor 2001 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0247] The memory 2002 may include a read-only memory and a random access memory, and provides instructions and data to the processor 2001. The memory 2002 may also include a nonvolatile random access memory.

[0248] The memory 2002 can be a volatile memory or a nonvolatile memory, or can include both volatile and nonvolatile memory. The nonvolatile memory can be a read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically EPROM (EEPROM), or flash memory, among others. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous dynamic RAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM), among others.

[0249] It should be understood that the computing device 2000 according to the embodiments of the present application can execute the method shown in the embodiments of the present application, and detailed description of the method implemented is described above. For brevity, it will not be repeated here. Figures 8-18 It should be understood that the computing device 2000 according to the embodiments of the present application can execute the method shown in the embodiments of the present application, and detailed description of the method implemented is described above. For brevity, it will not be repeated here.

[0250] The embodiments of the present application provide a computer readable storage medium, which stores a computer program, when the computer program is executed by a processor, the above-mentioned method is implemented.

[0251] The embodiments of the present application provide a chip, which includes at least one processor and an interface, the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the above-mentioned method.

[0252] The embodiments of the present application provide a computer program or computer program product, which includes instructions, when the instructions are executed, the computer executes the above-mentioned method.

[0253] Those skilled in the art should further understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in a general manner. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0254] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be implemented in hardware, software executed by a processor, or a combination of both. The software modules can be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art.

[0255] The above detailed description of the specific implementation is further detailed for the purpose of the present application, technical solutions and beneficial effects. It should be understood that the above description is only a specific implementation of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A task scheduling method, characterized in that: Applied to a runtime task scheduling system, the method includes: Obtain a task set to be scheduled, wherein the task set includes N tasks, and the N tasks are all executable tasks, where N is a positive integer; Scheduling the N tasks to M private task containers and public task containers corresponding to M execution units of the local scheduling domain, where M is a positive integer; Sending the tasks in each of the M private task containers to the private task queue of the execution unit corresponding to each private task container, so that the tasks in the private task queue are executed by the execution unit; In response to a task pull request from the first execution unit, a task is selected from the public task container and added to the first private task container. The task pull request is generated by the first execution unit when the number of tasks in its private task queue is less than or equal to a preset threshold. The first private task container is the private task container corresponding to the first execution unit.

2. The method according to claim 1, characterized in that The step of scheduling the N tasks to the M private task containers and the public task containers corresponding to the M executable units of the local scheduling domain includes: Acquire first task information, first hardware information, and first scheduling domain status information, where the first task information indicates task type information of each of the N tasks, the first hardware information indicates hardware information in the local scheduling domain, and the first scheduling domain status information indicates status information of each hardware unit in the local scheduling domain; Using the first task information, the first hardware information, and the first scheduling domain state information as inputs of an execution unit selection strategy, and outputting a target execution unit corresponding to each task; Each of the tasks is dispatched to a private container of a target execution unit corresponding to each of the tasks.

3. The method according to claim 2, characterized in that For tasks of the target execution unit that are not output by the execution unit selection strategy, the tasks are scheduled to the public task container of the local scheduling domain.

4. The method according to claim 2 or 3, characterized in that The runtime task scheduling system includes a plurality of scheduling domains; The step of scheduling the N tasks to the M private task containers and the public task containers corresponding to the M execution units of the local scheduling domain further includes: Obtaining second task information, second hardware information, and second scheduling domain status information, where the second task information indicates task type information of each of the N tasks, the second hardware information indicates hardware information of each of the multiple scheduling domains, and the second scheduling domain status information indicates status information of each hardware unit in each of the multiple scheduling domains; using the second task information, the second hardware information, and the second scheduling domain state information as inputs of a scheduling domain selection strategy, and outputting a first target scheduling domain corresponding to each task; Schedule each of the tasks to the target scheduling domain corresponding to each of the tasks.

5. The method according to claim 4, characterized in that The obtaining of the second task information, the second hardware information and the second scheduling domain state information further includes: Obtaining affinity descriptions of the respective tasks and hardware information of the respective execution units in the plurality of scheduling domains; Determining prompt information corresponding to each task based on the affinity description of each task and the hardware information of each execution unit in the multiple scheduling domains, and an affinity description parsing strategy, wherein the prompt information is used to prompt a target execution unit corresponding to each task; The first task information also includes prompt information corresponding to each task.

6. The method according to claim 5, characterized in that The affinity description includes an explicit affinity description or an implicit affinity description.

7. The method according to any one of claims 1 to 6, characterized in that The sending of the tasks in each of the M private task containers to the private task queue of the execution unit corresponding to each private task container includes: Obtain information of each private task container, third hardware information, and third scheduling domain status information, wherein the information of each private task container indicates task information in each private task container, the third hardware information indicates hardware information corresponding to the execution unit where each private task container is located, and the third scheduling domain status information indicates status information of each hardware unit in the local scheduling domain; Using the information of each private task container, the third hardware information, and the third scheduling domain state information as inputs of a first task issuing strategy, outputting a first task list, the first task list indicating tasks currently issued by each private task container and an ordering of the tasks currently issued; Based on the first task list, the tasks in each private task container are sent to the private task queue of the execution unit corresponding to each private task container, and the execution order of the tasks in the private task queue of the execution unit is the same as the task order in the task list.

8. The method according to any one of claims 1 to 7, characterized in that The step of responding to the task pull request of the first execution unit and selecting a task from the public task container and adding the task to the first private task container includes: Determining that the number of tasks in the first private task container is insufficient; Select tasks from the public task container and add them to the first private task container.

9. The method according to claim 8, characterized in that The selecting a task from the public task container and adding it to the first private task container includes: Acquire information of the common task container and fourth hardware information, wherein the information of the common task container indicates attribute information of each task in the common task container, and the fourth hardware information indicates hardware information of the first execution unit; Using the information of the common task container and the fourth hardware information as inputs of a task selection strategy, outputting a target task and the number of the target tasks, wherein the target task indicates a task selected from the common task container; Add the target task to the first private task container.

10. The method according to claim 8 or 9, characterized in that The selecting of tasks from the public task container and adding them to the first private task container may also include: Determining that the number of tasks in the first private task container is sufficient; A task is selected from the first private task container and added to the private task queue of the first execution unit.

11. The method according to claim 10, characterized in that The selecting a task from the first private task container and adding it to the private task queue of the execution unit includes: Obtain information of the first private task container, fifth hardware information, and fourth scheduling domain status information, where the information of the first private task container indicates task information in the private task container of the first execution unit, the fifth hardware information indicates hardware information of the first execution unit, and the fourth scheduling domain status information indicates status information of each hardware unit in the local scheduling domain; Using the information of the first private task container, the fifth hardware information, and the fourth scheduling domain state information as inputs of a second task issuing strategy, outputting a second task list, the second task list indicating tasks currently issued by the first private task container and an ordering of the tasks currently issued; Based on the second task list, the tasks in the first private task container are sent to the private task queue of the execution unit, and the execution order of the tasks in the private task queue of the execution unit is the same as the task order in the second task list.

12. The method according to any one of claims 4 to 11, characterized in that: Also includes: Determining that the number of tasks selected from the common task container is insufficient; Acquire information of a second private task container, where the information of the second private task container indicates task information in other task containers among the M private task containers except the first private task container; Using the information of the second private task container as input of the local task stealing strategy, outputting a stealing target task and a first stealing quantity, wherein the stealing target task indicates the task stolen from the other task container, and the first stealing quantity indicates the quantity of tasks stolen from the other task container; Based on the theft target task and the theft quantity, performing a task stealing operation from the other task container; The tasks stolen from the other task containers are added to the first private task container.

13. The method according to claim 12, characterized in that Also includes: determining that the number of tasks stolen by the local task stealing strategy is insufficient; Acquire sixth hardware information and fifth scheduling domain status information, wherein the sixth hardware information indicates hardware information of other scheduling domains among the multiple scheduling domains except the local scheduling domain, and the fifth scheduling domain status information indicates status information of each hardware unit in the other scheduling domains; Using the sixth hardware information and the fifth scheduling domain state information as inputs of a remote task stealing strategy, outputting a stealing target and a second stealing quantity, wherein the stealing target indicates a second target scheduling domain in the other scheduling domains on which a stealing operation is to be performed, and the second stealing quantity indicates the number of tasks to be stolen from the stealing target scheduling domain; Based on the stealing target and the stealing quantity, performing a task stealing operation from the second target scheduling domain; The tasks stolen from the second target scheduling domain are added to the first private task container.

14. The method according to any one of claims 4 to 13, characterized in that Also includes: Acquire information about abnormal events, the abnormal events including one or more of: tasks to be executed in the local scheduling domain reaching an upper limit of task capacity; the target scheduling domain being undeliverable; tasks to be executed in the target scheduling domain reaching an upper limit of task capacity; insufficient number of tasks being selected from the public task container and the local task stealing policy not being enabled; insufficient number of tasks being stolen through the local task stealing policy and the remote task stealing policy not being enabled; and insufficient number of tasks being stolen through the remote task stealing policy; The information of the abnormal event is used as input of an abnormal event processing strategy, and a processing strategy is output, wherein the processing strategy indicates a processing operation for the abnormal event.

15. The method according to claim 14, characterized in that One or more of the execution unit selection strategy, the scheduling domain selection strategy, the first task issuance strategy, the task selection strategy, the second task issuance strategy, the local task stealing strategy, the remote task stealing strategy and the abnormal event handling strategy can be configured to be turned on or off.

16. The method according to any one of claims 1 to 15, characterized in that The task capacity of each private task container in the M private task containers may be configured to be 0; Alternatively, the task capacity of the public task container may be configured to be 0.

17. A task scheduling device, characterized in that: include: An acquisition module is used to acquire a task set to be scheduled, wherein the task set includes N tasks, and the N tasks are all executable tasks, where N is a positive integer; a scheduling module, configured to schedule the N tasks to M private task containers and public task containers corresponding to M execution units of the local scheduling domain, where M is a positive integer; a task issuing module, configured to issue the tasks in each of the M private task containers to the private task queue of the execution unit corresponding to each private task container, so that the tasks in the private task queue are executed by the execution unit; A task pulling module is used to respond to a task pulling request from the first execution unit, select tasks from the public task container and add them to the first private task container, wherein the task pulling request is generated by the first execution unit when the number of tasks in its private task queue is less than or equal to a preset threshold, and the first private task container is the private task container corresponding to the first execution unit.

18. A computing device comprising a memory and a processor, characterized in that: Instructions are stored in the memory, and when the instructions are executed by the processor, the method according to any one of claims 1 to 16 is implemented.

19. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.

20. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 16 is implemented.