Task scheduling method and apparatus, and computing system
By dividing partitions in the parallel computing system and setting up partition schedulers, task scheduling is performed based on task criticality, the problem of unbalanced task scheduling is solved, and resource utilization and scheduling efficiency are improved.
Patent Information
- Application Number
- PCT/CN2024/099950
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2024-06-18
- Publication Date
- 2025-05-08
AI Technical Summary
In parallel computing systems with high complexity and large scale, unbalanced task scheduling leads to low resource utilization.
By dividing multiple computing units of the computing system into multiple partitions, each partition is set up with a partition scheduler, fine-grained in-partition task scheduling based on the task criticality, and cross-partition task scheduling is performed when necessary, to achieve balanced task scheduling.
It improves the resource utilization rate of the computing system, realizes balanced scheduling of tasks, and reduces the overhead of task scheduling.
Smart Images

Figure CN2024099950_08052025_PF_FP_ABST
Abstract
Description
Task scheduling method, device and computing system
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on October 31, 2023, with application number 202311440475.8 and application name “A task scheduling method, device and computing system”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a task scheduling method, device, and computing system. Background Art
[0003] In highly complex and large-scale parallel computing systems, such as multi-core and multi-processor server systems, cluster computing systems in data centers, high-performance computing (HPC) supercomputing systems, and cloud computing, the rational scheduling of tasks faces many challenges.
[0004] Among them, an important challenge is that when the scheduler in the computing system performs parallel task scheduling, there is an imbalance in task scheduling in the computing system, which leads to low resource utilization of the computing system.
[0005] Summary of the Invention
[0006] The present application provides a task scheduling method, device, and computing system, which can improve the resource utilization of the computing system.
[0007] This application adopts the following technical solutions:
[0008] In a first aspect, the present application provides a task scheduling method for scheduling tasks in a computing system, wherein the computing system includes multiple computing units, the multiple computing units are divided into multiple partitions, each of the multiple partitions corresponds to a partition scheduler, and each partition includes multiple task queues with different priorities set based on task criticality, wherein the task criticality is determined based on the latency tolerance of the computing unit running the task to the task. The method includes: determining a partition criticality for each of the multiple partitions, the partition criticality being used to indicate the latency tolerance of the computing unit in the partition to the scheduled task; when it is determined that cross-partition task scheduling needs to be initiated based on the partition criticality of each of the multiple partitions, a second partition scheduler moves tasks from a first task queue of a first partition of the multiple partitions to a task queue of a second partition of the multiple partitions, so that the computing unit in the second partition processes the task; the second partition scheduler is the partition scheduler corresponding to the second partition; and the first partition scheduler moves tasks from a second task queue of the second partition of the multiple partitions to a task queue of the first partition, so that the computing unit in the first partition processes the task, the first partition scheduler is the partition scheduler corresponding to the first partition; the priority of the first task queue is different from the priority of the second task queue.
[0009] In this application, whether cross-partition task scheduling is required can be determined based on the partition criticality. If cross-partition task scheduling is required, task scheduling is performed between two partitions, which can balance the tasks of the computing system and improve the resource utilization of the computing system.
[0010] Furthermore, multiple computing units of the computing system are partitioned, and tasks are scheduled in parallel across multiple partitions. A partition scheduler is set for each partition to perform autonomous task scheduling within the partition, thereby achieving fine-grained task scheduling. For different partitions, cross-partition task scheduling can be performed when necessary to achieve coarse-grained task scheduling. It can be seen that the task scheduling method provided in the embodiments of the present application can take into account both fine-grained and coarse-grained task scheduling, and has a good scheduling effect.
[0011] In one possible implementation, the aforementioned multiple computing units include processing cores of one or more processors; wherein the one or more processors include a heterogeneous first processor and / or a second processor.
[0012] Optionally, the first processor is a general-purpose processor, such as a CPU, and the second processor is a processor heterogeneous to the first processor, such as a GPU, an FPGA, or an ASIC.
[0013] In one possible implementation, the partition criticality is the proportion of the number of processing cores whose core criticality values are greater than a first threshold value among the core criticalities of multiple processing cores in a partition; the core criticality is determined based on the latency tolerance of the processing cores for tasks running on the processing cores. In the present application, the latency tolerance of the processing and verification tasks is related to the time slices (time slices are the resource granularity of the processing cores) consumed by the tasks during the running of the tasks and / or the waiting time of the processing cores. The lower the latency tolerance of the processing and verification tasks, the higher the task criticality; the higher the latency tolerance of the processing and verification tasks, the lower the task criticality.
[0014] In one possible implementation, the task scheduling method provided in an embodiment of the present application also includes: calculating the criticality difference of multiple partitions based on the partition criticality of multiple partitions, and the criticality difference is used to indicate the difference between the multiple partition criticalities of multiple partitions. The criticality difference of multiple partitions can measure whether the task scheduling in the entire computing system is balanced; when the criticality difference is greater than a second threshold, determining that cross-partition task scheduling needs to be initiated; when the criticality difference is less than or equal to the second threshold, each partition scheduler schedules the tasks within each partition.
[0015] In this application, partitioning the computing system and performing autonomous task scheduling within the partitions can improve the resource utilization of the processing cores within the partitions; when the criticality of each partition varies greatly, cross-partition task scheduling can improve the resource utilization of the computing system.
[0016] In one possible implementation, multiple partitions of a computing system correspond to a federal scheduler, and the federal scheduler is used to manage multiple partition schedulers; the federal scheduler can obtain the partition criticality of each of the multiple partitions from the partition scheduler, and determine a first partition and a second partition from the multiple partitions when determining that cross-partition task scheduling needs to be initiated based on the partition criticality of each of the multiple partitions; wherein the first partition is the partition corresponding to the partition criticality with the largest value among the multiple partition criticalities, and the second partition is the partition corresponding to the partition criticality with the smallest value among the multiple partition criticalities; and send bilateral scheduling instructions to the first partition scheduler and the second partition scheduler respectively; the bilateral scheduling instructions are used to instruct cross-partition task scheduling for the first partition and the second partition.
[0017] In one possible implementation, the priority of a task queue increases with increasing task criticality, and different numbers of time slices are allocated to queues of different priorities; the task criticality increases with decreasing latency tolerance for processing and checking tasks; the first task queue is the queue with the highest priority among the task queues of the first partition, and the second task is the queue with the lowest priority among the task queues of the second partition.
[0018] In the process of cross-partition task scheduling, a partition with high partition criticality obtains tasks from the low-priority queue of a partition with low partition criticality, which can reduce the criticality of the high partition; a partition with low partition criticality obtains tasks from the high-priority queue of a partition with high partition criticality, which can increase the criticality of the low partition, thereby reducing the difference between the partition criticalities of the two partitions and achieving balanced task scheduling between the two partitions.
[0019] In one possible implementation, each partition scheduler schedules tasks within each partition, including: for any one of the multiple tasks in each partition, while the computing unit is executing the task, the partition scheduler updates the queue to which the task belongs, so that the processing core executes the task according to the priority of the queue to which the task belongs.
[0020] In one possible implementation, the above-mentioned updating of the queue to which the task belongs includes: updating the task criticality of the task; and updating the queue to which the task belongs based on the task criticality of the task; wherein, when the task criticality of the task increases, the partition scheduler moves the task from the current queue to a queue with a priority higher than the priority of the current queue; when the task criticality of the task decreases, the partition scheduler moves the task from the current queue to a queue with a priority lower than the priority of the current queue.
[0021] In this application, the process of task scheduling within a partition is a dynamic scheduling process based on task criticality. The partition scheduler of each partition performs autonomous task scheduling on the partition, and schedules based on the task criticality of the tasks within the partition. The scheduling granularity is finer, so that the task scheduling within the partition is balanced, which can improve the resource utilization of the processing core within the partition; and for each task, the overhead required to obtain the task criticality is relatively small, which can significantly reduce the scheduling overhead required for task scheduling.
[0022] Furthermore, compared with the static scheduling method, the embodiment of the present application updates the queue based on the criticality of the task, does not rely on prior knowledge, and does not require any profiling performance collection data, thus greatly simplifying development and engineering optimization.
[0023] In one possible implementation, the task scheduling method provided by an embodiment of the present application also includes: each partition scheduler obtains the task dependency relationship between multiple tasks within the partition; and determines the initial task criticality of the multiple tasks respectively according to the task dependency depth indicated by the task dependency relationship; and according to the initial task criticality of the multiple tasks, determines the initial queues for the multiple tasks respectively from the multiple queues of different priorities of the partition.
[0024] In this application, the deeper the dependency depth of a task, the more other tasks need to rely on the data of the task, and the higher the criticality of the task.
[0025] In a second aspect, an embodiment of the present application provides a computing system, which includes a partition scheduler and may also include a federal scheduler. The partition scheduler and the federal scheduler include various modules for implementing the method described in the first aspect and one of its possible implementation methods, such as an acquisition module, a determination module, a scheduling module, a computing module, a sending module, etc.
[0026] The partition scheduler and federated scheduler in the computing system described above have the functionality to implement the behaviors described in the method examples of the first aspect and any one of its possible implementations. These functions can be implemented via hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the aforementioned functions.
[0027] In a third aspect, the present application provides a computing device comprising a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by at least one processor, the computing device executes the method of the first aspect and any one of its possible implementations.
[0028] In a fourth aspect, the present application provides a computer-readable storage medium storing computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0029] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions are run on a computer, the method of the first aspect and any one of its possible implementations is executed.
[0030] In a sixth aspect, the present application provides a chip system, comprising: a processor for calling and running a computer program from a memory, so that a storage device equipped with the chip system executes the method of the first aspect and any one of its possible implementations.
[0031] It should be understood that the beneficial effects achieved by the technical solutions of the second to sixth aspects of this application and the corresponding possible implementation methods can be referred to the technical effects of the first aspect and its corresponding possible implementation methods mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a schematic diagram of a task dependency diagram provided in an embodiment of the present application;
[0033] FIG2 is a schematic diagram of an architecture of a computing system provided in an embodiment of the present application;
[0034] FIG3 is a schematic diagram of the relationship between partitions and schedulers in a computing system provided in an embodiment of the present application;
[0035] FIG4 is a second schematic diagram of the architecture of a computing system provided in an embodiment of the present application;
[0036] FIG5 is a schematic diagram of a scheduler deployment in a computing system according to an embodiment of the present application;
[0037] FIG6 is a schematic diagram of a scheduler deployment in a computing system according to an embodiment of the present application;
[0038] FIG7 is a schematic diagram of a scheduler deployment in a computing system according to an embodiment of the present application;
[0039] FIG8 is a schematic diagram of a scheduler deployment in a computing system according to an embodiment of the present application;
[0040] FIG9 is a schematic diagram of a scheduler deployment in a computing system according to an embodiment of the present application;
[0041] FIG10 is a flowchart of a task scheduling method according to an embodiment of the present application;
[0042] FIG11 is a second flow chart of a task scheduling method provided in an embodiment of the present application;
[0043] FIG12 is a schematic diagram of a task queue provided in an embodiment of the present application;
[0044] FIG13 is a third flow chart of a task scheduling method provided in an embodiment of the present application;
[0045] FIG14 is a fourth flow chart of a task scheduling method provided in an embodiment of the present application;
[0046] FIG15 is a schematic diagram of a partition criticality reporting method provided in an embodiment of the present application;
[0047] FIG16 is a schematic diagram of a task scheduling process provided in an embodiment of the present application;
[0048] FIG17 is a schematic diagram of the structure of a computing system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone.
[0050] In the description and claims of the embodiments of this application, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order of objects. For example, the terms "first partition" and "second partition" are used to distinguish different partitions, rather than to describe a specific order of partitions.
[0051] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0052] In the description of the embodiments of the present application, unless otherwise specified, “plurality” means two or more than two, and “plurality” may also be described as “at least two”.
[0053] The embodiments of the present application provide a task scheduling method, device and system, which mainly relate to a computing system with high computational complexity and large scale, for a computing architecture with multiple cores, multiple processors or multiple computing nodes. The provided task scheduling method is used to flexibly schedule tasks in the computing system so that the overall load of the computing system is balanced, that is, tasks are evenly scheduled, thereby improving the resource utilization of the computing system.
[0054] First, some technical terms involved in the task scheduling method, device and system provided in the embodiments of the present application are explained.
[0055] 1. Task Scheduling
[0056] In computing systems, the essence of a task is the access, calculation, analysis, and processing of data. In actual processing, data may have sequential dependencies, which in turn creates dependencies between tasks. For example, executing Task A requires executing Task B first, as the processing of Task A depends on the results of Task B.
[0057] To ensure efficient processing of large-scale data, tasks must be executed in parallel, utilizing as many computing resources as possible, especially in multi-core computing systems. However, because storage resources such as memory and cache are generally and significantly less than computing units, the tasks scheduled by the scheduler may differ depending on whether they access memory or perform computations. Task scheduling is based on the dependencies between tasks and their differences, allowing tasks to be scheduled and executed according to a specific strategy.
[0058] It should be understood that in the embodiments of the present application, tasks may include I / O-intensive tasks and computation-intensive tasks. Optionally, I / O-intensive tasks include memory access tasks or tasks that interact with users, and computation-intensive tasks may include tasks in the fields of cloud computing, AI computing, etc.
[0059] 2. Task criticality
[0060] For each task, the task criticality can be defined. The task criticality is determined based on the latency tolerance of the processing core running the task. The latency tolerance of the processing core to the task is related to the time slices consumed by the task during the task running process and / or the waiting time of the processing core. The time slice is a unit of measurement for the length of time a processor is assigned to execute a task, and a task consumes one or more time slices.
[0061] It is understandable that the waiting time of the processing core or the time slice consumed by the task is related to the operations involved in the task running process, such as memory access operations and / or computing operations. For example, when there is an operation to initiate a memory access request during the task running process, it indicates that the data used for the processing core to perform calculations is not ready, and the data needs to be obtained through a memory access operation, and then the processing core performs computing operations based on the obtained data. It can be seen that when the task executes the memory access operation, the latency of the task increases, the processing core is not currently performing computing operations, and no longer consumes time slices. The processing core is in an idle or suspended state, and the processing core needs to wait for the memory access operation to end before it can perform computing operations. It can be seen that the waiting time of the processing core is long, so it is considered that the processing core has a low latency tolerance for the task. It is understandable that tasks that consume fewer time slices or cause the processing core to wait for a longer time can be called high-latency tasks.
[0062] In contrast, when the task is executing a computational operation, the processing core is currently busy and consumes fewer time slices. Furthermore, the fact that the current operation is a computational operation indicates that the data required for the processing core to perform the calculation is ready, eliminating the need for memory access. The processing core does not need to wait, or the waiting time is short. Therefore, the task processing core has a higher latency tolerance for the task. It can be understood that tasks that consume more time slices or shorten the processing core's waiting time are considered low-latency tasks.
[0063] It should be understood that the relationship between task criticality and the latency tolerance of processing and verification tasks is: the lower the latency tolerance of the processing and verification task, the higher the task criticality; the higher the latency tolerance of the processing and verification task, the lower the task criticality. In other words, higher task criticality indicates that the processing core is more idle, while lower task criticality indicates that the processing core is more busy. In some cases, it can also be understood that memory access tasks have higher task criticality, while compute tasks have lower task criticality.
[0064] In one implementation, the criticality of the task CL It is defined by the following bisection formula:
[0065] Task criticality includes high task criticality (C L =1) and low task criticality (C L =0), low delay tolerance corresponds to high task criticality, and high delay tolerance corresponds to low task criticality.
[0066] Optionally, the mission criticality C L It can also be defined in other ways, for example, the task criticality C L More values can be included, and the delay tolerance can be divided into multiple levels, with different levels corresponding to different task criticalities.
[0067] 3. Nuclear criticality
[0068] Nuclear criticality (C core ) is determined based on the latency tolerance of the tasks running on the processing cores. Referring to the definition of task criticality above, at a given moment, core criticality and task criticality are equal, but over a period of time, they may differ. This is because the processing core running a task may change. For example, if a task is currently running on processing core 1 and is next running on processing core 2, the core criticality of processing core 1 will no longer be equal to the task criticality of the task.
[0069] 4. Partition criticality
[0070] First, the concept of partitioning in the embodiments of the present application is introduced. Partitioning is the division of a computing unit into blocks. A computing system includes multiple computing units, and the multiple computing units are divided into multiple partitions. A partition includes one or more computing units.
[0071] The multiple computing units of the above-mentioned computing system may include processing cores of one or more processors. The one or more processors in the embodiment of the present application may include homogeneous processors or heterogeneous processors.
[0072] Each partition in the computing system corresponds to a partition criticality. The partition criticality is the proportion of processing cores whose core criticality values are greater than a first threshold among the core criticality values of multiple processing cores in a partition. That is, the partition criticality is the proportion of processing cores whose core criticality values are higher (the core criticality values are greater than the first threshold) among the multiple processing cores in a partition.
[0073] Referring to the above introduction to the dichotomous formula of task criticality, it is assumed that the nuclear criticality is also defined by the dichotomous formula, that is, the nuclear criticality includes the high nuclear criticality (C core =1) and low nuclear criticality (Ccore =0), then the partition criticality C partition It can be determined by the following formula:
[0074] Where n represents the total number of processing cores in a partition, C core_i represents the core criticality of the i-th processing core.
[0075] It should be understood that partition criticality indicates the latency tolerance of computing units within a partition for scheduled tasks. The higher the partition criticality, the lower the latency tolerance of computing units within the partition for scheduled tasks; the lower the partition criticality, the higher the latency tolerance of computing units within the partition for scheduled tasks.
[0076] 5. Task dependency graph (TDG)
[0077] A task dependency graph is a representation form that can represent the relationship between tasks. The task dependency path can be determined through the task dependency graph, and the task dependency path can indicate the dependency relationship between tasks.
[0078] Typically, when tasks are generated, a task dependency graph can be generated. However, it should be understood that not all computing systems have a task dependency graph. Referring to FIG1 , the task dependency graph shown in FIG1 includes tasks T0-T13, and the task dependency graph includes three task dependency paths, namely:
[0079] Path 1: T1→T3→T6→T11→T13
[0080] Path 2: T1→T4→T7→T8→T10→T12→T13
[0081] Path 3: T1→T5→T9
[0082] Referring to Figure 1, in Path 1, T3 depends on T1, T6 depends on T3, T11 depends on T6, and T13 depends on T11. In Path 2, T4 depends on T1, T7 depends on T4, T8 depends on T7, T10 depends on T8, T12 depends on T10, and T13 depends on T12. In Path 3, T5 depends on T1, and T9 depends on T5.
[0083] It should be noted that T0, T2 and T4 in FIG1 are independent tasks and do not depend on other tasks.
[0084] With the development of big data, the architecture of computing systems such as embedded computing systems, edge computing systems, cluster computing systems in data centers, high-performance computing (HPC) supercomputing systems, and cloud computing is becoming increasingly complex. High-density interconnection, multi-core, many-core, and heterogeneous cores have become the mainstream architecture of computing systems. Task scheduling for these computing systems faces numerous challenges.
[0085] One task scheduling method involves a single scheduler (the master scheduler) performing global task scheduling for the computing system. The scheduler collects metrics such as the operating status and bandwidth of the processors (workers) through interfaces, analyzes these metrics, and schedules tasks based on the results. This task scheduling method incurs significant overhead due to the need to collect metrics, and in large-scale computing systems, the scheduler overhead is even greater. Furthermore, the scheduler can become overloaded, leading to low scheduling efficiency, unbalanced task scheduling, and low resource utilization.
[0086] Another task scheduling method is to use the instruction-level scheduler (such as the warp scheduler) originally contained in the processor of the computing system to perform parallel scheduling. This scheduling method is a static scheduling method and requires relying on prior knowledge of the task behavior (such as the consumption of the processing core, the running time of the task, and the memory access behavior). In addition, the static scheduling method is not suitable for irregular tasks.
[0087] Another task scheduling method is distributed task scheduling, such as distributed scheduling based on the message passing interface (MPI). This distributed scheduling method is a coarse-grained scheduling method. The scheduling granularity of this task scheduling method is application (i.e., application-level task scheduling). There is a problem of unbalanced task scheduling (i.e., unbalanced load), which makes the resource utilization of the computing system low.
[0088] In response to the above problems, an embodiment of the present application provides a task scheduling method, which divides multiple computing units of a computing system into multiple partitions, and configures a partition scheduler for each partition. The partition scheduler is used to schedule tasks within the partition. Each partition includes multiple task queues with different priorities set based on task criticality. The task criticality is determined based on the delay tolerance of the computing unit running a task to the task. In this method, when it is determined that cross-partition task scheduling needs to be initiated based on the partition criticality of each partition, the second partition scheduler (the partition scheduler of the second partition) among the multiple partition schedulers can schedule the first task in the first task queue of the first partition to the computing unit in the second partition, and the computing unit in the second partition executes the first task of the first partition; and the first partition scheduler (the partition scheduler of the first partition) can schedule the task in the second task queue of the second partition to the computing unit in the first partition, and the computing unit in the first partition executes the second task of the second partition. In this way, the tasks of the entire computing system can be evenly scheduled (i.e., load balancing is achieved), thereby improving the resource utilization of the computing system.
[0089] Furthermore, in the embodiments of the present application, multiple computing units of the computing system are partitioned, tasks are scheduled in parallel across the multiple partitions, and a partition scheduler is provided for each partition to perform autonomous task scheduling within the partition, thereby achieving fine-grained task scheduling. For different partitions, cross-partition task scheduling can be performed when necessary, thereby achieving coarse-grained task scheduling. It can be seen that the task scheduling method provided in the embodiments of the present application can take into account both fine-grained and coarse-grained task scheduling, and has a good scheduling effect.
[0090] The task scheduling method provided in the embodiments of the present application is applied to computing systems, which can be computing systems of different sizes, such as a small embedded system with a single processor, or a very large system with more than 10 million cores. For example, the computing system in the embodiments of the present application can be any of the following: a processor including multiple cores, a server including multiple processors (the multiple processors can be homogeneous processors and / or heterogeneous processors, and the processors can be single-core or multi-core processors), a cabinet including multiple servers, or a cluster including multiple cabinets, etc.
[0091] Exemplarily, FIG2 shows an architectural diagram of a computing system. As shown in FIG2 , the computing system includes a plurality of computing devices 200 (ie, computing nodes), and the plurality of computing devices 200 can communicate with each other.
[0092] The computing device 200 may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device 200 may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0093] For example, with continued reference to FIG2 , computing device 200 may include: one or more processors (e.g., general-purpose processor 201 and heterogeneous processor 205 shown in FIG2 ), memory 202, and communication interface 203. General-purpose processor 201, memory 202, communication interface 203, and heterogeneous processor 205 may be connected to each other via bus 204, or in other ways. Alternatively, the various components included in computing device 200 may be implemented in hardware, including one or more signal processing and / or application-specific integrated circuits, software, or a combination of hardware and software.
[0094] The computing device may include one or more general-purpose processors 201. The general-purpose processor 201 is the control center of the computing device 200. The general-purpose processor 201 may be a CPU or other general-purpose processor, such as a microprocessor or any conventional processor. Optionally, the general-purpose processor 201 may include one or more processing cores.
[0095] The controller in general-purpose processor 201 is the nerve center and command center of computing device 200. Based on instruction opcodes and timing signals, the controller generates operational control signals to control instruction fetching and execution. Optionally, general-purpose processor 201 may also include memory for storing instructions and data.
[0096] Heterogeneous processor 205 is a processor that is heterogeneous from general-purpose processor 201. Heterogeneous processor 205 may include, for example, a graphics processing unit (GPU), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Heterogeneous processor 205 may include one or more processing cores, and typically includes multiple processing cores.
[0097] The memory 202 includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical storage, magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer. In the embodiment of the present application, the memory 202 can store information such as computer instructions.
[0098] In one possible implementation, the memory 202 may exist independently of the processor (e.g., the general-purpose processor 201 or the heterogeneous processor 205). The memory 202 may be connected to the processor via a bus 204 and used to store data, instructions, or program codes. When the processor calls and executes the instructions or program codes stored in the memory, the relevant steps of the method provided in the embodiments of the present application can be implemented.
[0099] In another possible implementation, the memory 202 may also be integrated with the processor.
[0100] The communication interface 203 may be a transceiver module for communicating with other devices or communication networks, such as Ethernet, RAN, or wireless local area networks (WLAN). The communication interface 203 may receive instructions, messages, or data. The transceiver module may be a device such as a transceiver or a transceiver. Alternatively, the communication interface 203 may be a transceiver circuit located within the processor, configured to implement signal input and output for the processor. The communication interface 203 may be a wired interface (port), such as a fiber distributed data interface (FDDI) or a gigabit Ethernet (GE) interface, or a wireless interface.
[0101] Bus 204 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. This bus can be classified as an address bus, a data bus, a control bus, etc. Buses can also be classified as serial buses and parallel buses. For ease of illustration, FIG2 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.
[0102] It should be noted that the computing device in Figure 2 is only an example of a computing device. The computing device may have more or fewer components than those shown in Figure 2, may combine two or more components, or may have different component configurations. For example, the computing device may also include a smart network card, such as a data processing unit (DPU).
[0103] The task scheduling method provided in the embodiment of the present application is executed by a scheduler. Optionally, the scheduler can be a hardware-based scheduler or a software-based scheduler, which is not limited in the embodiment of the present application.
[0104] In combination with the above content, it can be seen that in the task scheduling method provided in the embodiment of the present application, multiple computing units in the computing system are partitioned, and the multiple computing units of the computing system are divided into multiple partitions. A partition scheduler (which can be called an L1 scheduler) is set for each partition. The partition scheduler has completely autonomous scheduling rights for the computing resources within the partition and is used to autonomously schedule tasks within the partition. The computing unit includes one or more processing cores of a processor. In the computing system, the multiple processing cores can be homogeneous processing cores and / or heterogeneous processing cores. Optionally, in some cases, for example, when the number of processing cores in the computing system is large (such as a multi-core computing system) and there are many partitions, a higher-level federal scheduler (which can be called an L2 scheduler) can be set on the basis of the partition scheduler. The federal scheduler is used to manage multiple partition schedulers. The L2 scheduler performs task scheduling between two partitions according to the balance state of multiple partitions (i.e., bilateral peer scheduling). When the number of processing cores in the computing system is particularly large (such as a multi-core computing system), a federal scheduler (which can be called an L3 scheduler) can also be set at the upper level of the L2 scheduler.
[0105] For example, referring to (a) in FIG3 , in a computing system, six computing units can be divided into a partition, and each partition corresponds to an L1 scheduler. Referring to (b) in FIG3 , six computing units can be divided into a partition, and each partition corresponds to an L1 scheduler. Furthermore, when the number of partitions in the computing system is large, the number of L1 schedulers is also large. In this case, multiple L1 schedulers can also be cascaded to an L2 scheduler, and the L2 scheduler manages multiple L1 schedulers. Referring to (c) in FIG3 , when the number of partitions in the computing system is particularly large, multiple L1 schedulers can be divided into multiple groups, and each group of L1 schedulers is cascaded to an L2 scheduler. In this case, there are multiple L2 schedulers in the computing system. When the number of L2 schedulers is relatively large, multiple L2 schedulers are cascaded to an L3 scheduler, and the L3 scheduler manages multiple L2 schedulers.
[0106] Optionally, a reasonable number of L1 schedulers are set according to the scale and actual needs of the computing system, and it is selected whether to cascade the L2 scheduler and L3 scheduler. Of course, as the scale of the computing system increases, one or more levels of schedulers can be cascaded on top of the L3 scheduler. The scheduler is set according to actual needs, and the embodiments of this application do not limit it.
[0107] As can be seen from Figure 3, the L1 scheduler schedules tasks within the partition according to the task criticality of the tasks in each partition, realizing autonomous scheduling within the partition (i.e., partition autonomy). Autonomous scheduling within the partition is a fine-grained task scheduling strategy. The L2 manages multiple L1 schedulers, and initiates cross-partition task scheduling between two partitions (i.e., bilateral peer scheduling) when the difference in partition criticality reported by the L1 scheduler is large. When an L3 scheduler is provided, multiple L2 schedulers continue to report multiple domain criticalities to the L3 scheduler (domain criticality is the criticality obtained by each L2 by performing relevant operations on multiple partition criticalities collected, called domain criticality, which can also be understood as classifying multiple partitions into one domain). When the L3 scheduler determines that the difference in the criticality of each domain is large, it determines to initiate cross-domain task scheduling between the two domains (the essence of cross-domain task scheduling here is also cross-partition task scheduling). The detailed process of the scheduling method will be introduced in the following embodiments.
[0108] Referring to the computing system architecture diagram shown in FIG4 (including the computing system's hardware and software systems), as one possible implementation, when the aforementioned scheduler is in software form, the scheduler can be deployed in the runtime software system in FIG4 . It is understood that the runtime is the environment required by the program during execution and can also be referred to as the runtime environment or execution environment.
[0109] The runtime software system may include a runtime host system and an application runtime system (the application runtime system is the environment in which the application runs). Optionally, a scheduler may be deployed in the host system in the runtime system, and the scheduler may be a scheduling application. In this way, the scheduling application may be loaded as a runtime service.
[0110] As a possible implementation, if the scheduler is implemented in software, it can be deployed in the runtime library software module shown in Figure 4. As can be understood, a runtime library is a library file required by a program at runtime. It generally includes commonly used programming functions, such as string operations, file operations, and user interfaces. The runtime library can provide an API that calls and integrates the scheduler. Application developers can implement the scheduler's functionality by explicitly calling the API, or the compiler can automatically integrate the scheduler during the linking phase.
[0111] As a possible implementation method, when the above-mentioned scheduler is in hardware form, the scheduler can be deployed on the processing core in the computing system. In the embodiment of the present application, the scheduler is deployed as much as possible in a local deployment manner (deployed on the same server). This can reduce the bandwidth consumption of the task scheduling process and improve the efficiency of task scheduling.
[0112] The following examples introduce possible deployment methods of various schedulers for computing systems of several different sizes.
[0113] 1. Single-processor computing system including multiple processing cores
[0114] In a single-processor computing system with multiple processing cores, when partitioning, the processing cores are partitioned based on the granularity of the computing unit. A partition includes one or more processing cores.
[0115] In one implementation, the L1 scheduler for each partition can be deployed on any processing core within that partition. When there are a large number of partitions, at least one L2 scheduler can also be deployed, and at least one L2 scheduler can be deployed on any one or more processing cores of the processor. This deployment method allows for tighter integration of the scheduler with multiple cores, resulting in more efficient task scheduling.
[0116] For example, assuming that each partition includes two processing cores, for each partition, the L1 scheduler can be integrated into any processing core within the partition. (a) in Figure 5 shows an example of the deployment of the L1 scheduler corresponding to the partition and an example of the deployment of the L2 scheduler (assuming that the computing system deploys only one L2 scheduler).
[0117] In another implementation, the L1 scheduler of each partition can be deployed on a dedicated processing core, that is, a dedicated processing core outside the partition is defined within a processor for deploying the scheduler, and the partition scheduler of each partition is deployed on the dedicated processing core. When the number of partitions in the computing system is large, the L2 scheduler can also be deployed on the dedicated processing core. In this deployment method, the process of task scheduling by the scheduler does not occupy the computing resources of the processing core within the partition, which is conducive to the smooth execution of tasks. In addition, the L1 scheduler and the L2 scheduler are located on the same processor, and the processing cores on the same processor are packaged together. The task scheduling process is an on-chip interaction, and the task scheduling efficiency is high. (b) in Figure 5 shows an example of the deployment of the L1 scheduler corresponding to the partition and an example of the deployment of the L2 scheduler (assuming that the computing system only deploys one L2 scheduler).
[0118] It is understandable that small computing systems with single-core multi-processors generally do not need to deploy L3 and higher-level schedulers.
[0119] 2. Single-server computing system with multiple processors
[0120] In one implementation, in a single-server computing system with multiple processors, the deployment method of the L1 scheduler can be similar to the deployment method in the above-mentioned single-processor computing system with multiple processing cores, that is, the L1 scheduler is deployed on any processing core in each partition. The deployment method of the L2 scheduler can be similar to the deployment method in the above-mentioned single-processor computing system with multiple processing cores, and at least the L2 scheduler is deployed on any processing core of any one or more processors among the multiple processors. In this deployment method, the L1 scheduler and the L2 scheduler are deployed in the same server, and the task scheduling efficiency is high. (a) in Figure 6 shows an example of the deployment of the L1 scheduler corresponding to the partition and an example of the deployment of the L2 scheduler (assuming that the computing system only deploys one L2 scheduler).
[0121] In another implementation, the L1 scheduler for each partition can be deployed on a dedicated processor or dedicated processing core, and at least one L2 scheduler can also be deployed, and at least one L2 scheduler can also be deployed on a dedicated processor or dedicated processing core. Figure 6(b) shows an example deployment of the L1 scheduler and the L2 scheduler corresponding to the partition (assuming that the computing system deploys only one L2 scheduler).
[0122] 3. Single rack computing system containing multiple servers
[0123] Similarly, in one implementation, in a single-cabinet computing system with multiple servers, the L1 scheduler can be deployed in a manner similar to that in the aforementioned single-processor computing system with multiple processing cores, namely, deploying the L1 scheduler on any processing core within each partition. Multiple L1 schedulers and their corresponding L2 schedulers are preferably deployed on the same server, for example, on any one or more processing cores of any one or more processors on the same server. This avoids cross-partition scheduling and avoids cross-server scheduling, thereby improving task scheduling efficiency.
[0124] Optionally, in a single-cabinet computing system with multiple servers, when the number of L2 schedulers is relatively large, an L3 scheduler can also be deployed. When multiple L2 schedulers are deployed on the same server, multiple L2 schedulers and their corresponding L3 schedulers can be deployed on the same server as much as possible to avoid cross-server scheduling. If the computing system deploys only one L3 scheduler, the L3 scheduler can be deployed on any processing core of any processor in any server. Figure 7 shows an example of the deployment of the L1 scheduler corresponding to the partition, an example of the deployment of the L2 scheduler, and an example of the deployment of the L3 scheduler (assuming that the computing system deploys only one L3 scheduler and each server deploys one L2 scheduler).
[0125] In another implementation, in some cases, for example, when an L1 scheduler and its corresponding L2 scheduler and L3 scheduler cannot be deployed on the same server, and the computing system includes a management server, the L2 scheduler and L3 scheduler can be deployed on the management server. Figure 8 shows an example deployment of an L2 scheduler and an example deployment of an L3 scheduler (assuming that the computing system deploys only one L3 scheduler).
[0126] In another implementation, when each server has a DPU network card, the L2 scheduler or L3 scheduler can also be deployed on the DPU network card's processing core. Leveraging the DPU's computing power can fully accelerate task scheduling, thus conserving server processing core resources. Figure 9 shows an example deployment of an L2 scheduler and an example deployment of an L3 scheduler (assuming the computing system deploys only one L3 scheduler).
[0127] It is understandable that a single-server computing system with multiple processors generally does not need to deploy L3 and higher-level schedulers.
[0128] Based on the above description of technical terms, system architecture, etc., the following is a detailed introduction to the task scheduling method provided by the embodiment of the present application. It can be understood that the embodiment of the present application proposes a strategy for partitioning computing units in a computing system. Based on this, task scheduling includes partition autonomous task scheduling and cross-partition task scheduling.
[0129] The multiple computing units in the computing system include processing cores of one or more processors, wherein the multiple processors include a heterogeneous first processor and / or a second processor, where the first processor may be, for example, the general-purpose processor 201 in FIG. 2 , and the second processor may be, for example, the heterogeneous processor 205 in FIG. 2 .
[0130] For example, in a homogeneous multi-core computing system (homogeneous means that the computing system contains one type of processor, such as only CPU), suppose that for a server containing four 48-core CPUs, the 48 cores in the server are divided into 8 partitions, namely {P0, P1,......., P7}, and each partition contains 24 processing cores.
[0131] For another example, in a heterogeneous multi-core computing system (heterogeneous means that the computing system contains at least two different types of processors, such as CPU and GPU), suppose that for a server with a heterogeneous processor containing one 48-core CPU and four 128-core GPUs, the processor cores in the server are divided into 10 partitions, namely {P0, P1,......., P9}, where the 48-core CPU is divided into two partitions, each partition contains 24 processing cores, and the four 128-core GPUs are divided into 8 partitions, each partition contains 64 processing cores.
[0132] Optionally, in an embodiment of the present application, a partition vector can be used to identify a partition. The partition vector of a partition can be represented as a vector P, P = (n, p, t, e0, e1, ...). The embodiment of the present application does not limit the dimension of the partition vector. The partition vectors of multiple partitions of the computing system can form a matrix. The meaning of each element in the partition vector is shown in Table 1 below.
[0133] Table 1
[0134] Regarding processor indexing, in some cases, some computing systems, such as those using the big.LITTLR architecture, include both large and small cores within a single processor. These cores can be classified into two different types based on their characteristics. Alternatively, a single processor can be used to virtually represent two different types of processors. For example, CPUs can be classified into different types based on whether they have vector or matrix scalability.
[0135] The extended index can be added according to actual needs. For example, e0 can usually be used to indicate the group index of the processing cores in the multi-core processor. The extended index can also include an index indicating a task timeout coefficient, a queue idle coefficient, etc.
[0136] Multi-dimensional indexing based on partition vectors can be used to obtain the topology of the computing system, facilitating task routing across the entire computing system. This information is also provided during cross-partition task scheduling, enabling near-term bilateral scheduling. For example, cross-partition scheduling between two partitions within the same server or cabinet is prioritized, minimizing peer-to-peer scheduling across servers and cabinets, thus reducing the overhead of cross-node scheduling. In some practical applications, partition vectors can be provided to developers for explicit resource invocation. Because partition vectors contain compute node indices and processor type indices, they can be directly used for heterogeneous and distributed programming.
[0137] In conjunction with (a) in FIG3 , the method for scheduling tasks within each of the multiple partitions of the computing system is the same. The following embodiment uses one partition as an example to describe the process of scheduling tasks within a partition. The main idea of scheduling tasks within a partition is to schedule tasks using a multi-level queue scheduling strategy based on the criticality of the tasks within the partition. As shown in FIG10 , the task scheduling method provided in the embodiment of the present application includes S901-S903.
[0138] S901: The partition scheduler determines the initial task criticality of the acquired task.
[0139] It can be understood that the tasks to be processed by the processing cores in the partition include multiple tasks. After the scheduler starts to obtain each task, it determines the task criticality of the initial state (i.e., the initial task criticality) for each task. During the task processing process, the task criticality of each task may change dynamically.
[0140] For a detailed introduction to task criticality and how to determine task criticality, please refer to the technical terminology section above and will not be repeated here.
[0141] In combination with FIG10 , as shown in FIG11 , in one implementation, a method for a partition scheduler to determine an initial task criticality of a task (one task or multiple tasks) includes S9011 - S9012 .
[0142] S9011. The partition scheduler obtains task dependencies between multiple tasks in the partition.
[0143] In an embodiment of the present application, the task dependency relationship between tasks can be obtained through the task dependency graph of the computing system, and the task dependency path can be obtained through the task dependency graph. According to the task dependency path, it can be known which tasks have dependencies, how the tasks are dependent on each other, and the task dependency depth. For an introduction to the task dependency graph, please refer to Figure 1 and the description of related technical terms.
[0144] S9012: The partition scheduler determines initial task criticalities of the multiple tasks respectively according to the task dependency depths indicated by the task dependency relationships.
[0145] The task dependency depth can be understood as the position of a task in the task dependency path to which it belongs (a task may have multiple task dependency paths).
[0146] For example, referring to the task dependency graph shown in FIG1 above, among the three task dependency paths, T1 is simultaneously located in path 1 (T1→T3→T6→T11→T13), path 2 (T1→T4→T7→T8→T10→T12→T13) and path 3 (T1→T5→T9). The task dependency depth of T1 in path 1 is 4, the task dependency depth of T1 in path 2 is 6, and the task dependency depth of T1 in path 3 is 2.
[0147] Specifically, the process of determining the initial task criticality of multiple tasks according to the task dependency depth is as follows:
[0148] First, the task dependency graph is analyzed to obtain the attribute parameter of the task, which is the task dependency depth. A bottom-level (BL) algorithm is then used to determine the maximum path depth in the dependency path where the task resides. For example, if T1 is located on different task dependency paths, with task dependency paths of 4, 6, and 2, respectively, then T1's task dependency depth of 6 should be selected to determine T1's initial task criticality. It can be understood that the value in the circle corresponding to each task in Figure 1 is the determined task dependency depth.
[0149] Secondly, the initial task criticality of the task is determined according to the maximum value of the task dependency depth.
[0150] In one implementation, assuming that the initial task criticality of the task is also determined by the dichotomy formula, the initial task criticality (C p ) can be:
[0151] Wherein, Dpath is the maximum value of the task dependency depth of the task. When the maximum value of the task dependency depth is greater than or equal to the second threshold, the initial task criticality is 1 (high task criticality). When the maximum value of the task dependency depth is less than the second threshold, the initial task criticality is 0 (low task criticality).
[0152] In summary, the embodiment of the present application can set an initial task criticality for a task based on the task dependency graph. The initial task criticality is a static value determined when the task is generated.
[0153] In one implementation, when the task dependencies of some tasks cannot be obtained (the computing system has no task dependency graph or it is known from the dependency graph that some tasks are independent tasks and do not depend on other tasks), the initial task criticality of the task is set to the lowest task criticality, that is, the task is assumed to be a low-latency task (such as a computing task). For example, referring to the introduction of the concept of task criticality in the above embodiment, when the value of the task criticality is determined by a binary formula, that is, the task criticality includes 0 and 1, for the new task, the task criticality can be set to a low task criticality of 0. Referring to Figure 1, T0, T2 and T4 in Figure 1 are independent tasks, so the initial task criticalities of T0, T2 and T4 can all be set to 0.
[0154] S902: The partition scheduler determines an initial queue for the task from queues of different priorities of the partition according to the criticality of the initial task.
[0155] In an embodiment of the present application, multiple task queues with different priorities are set for each partition. When the processing core runs tasks, the tasks in each queue are executed in order from high to low priority. For tasks in the same queue, the processing core adopts a round robin strategy to run tasks.
[0156] For example, referring to the task queue shown in FIG12 , each partition may include four task queues of different priorities, namely Q L0 , Q L1 , Q L2 , Q L3 , the priority of the queue increases in turn, queue Q L3 The tasks in T a and T b , queue Q L2 The task in T c , queue Q L0 The task in T d According to the priority of the task queue, when executing a task, the queue with the highest priority Q is executed first. L3 For tasks in queue Q L3 T in a and T b , adopting the round-robin scheduling strategy, that is, executing T a , consumed as T a After the allocated time slice, execute T b , consumed as T b After the allocated time slice, execute T a , or other tasks in the queue; then, execute queue Q L2 T in c ;Finally, execute queue Q L0 T ind .
[0157] In an embodiment of the present application, time slices of a processing core are dynamically allocated to tasks based on queue priority from high to low. Different task queues have different numbers of time slices, and the number of time slices corresponding to a task queue decreases as the priority of the task queue increases. Optionally, the number of time slices corresponding to the multiple task queues of different priorities maintains a multiple relationship.
[0158] For example, referring to the four priority task queues shown in FIG12 above, the number of time slices corresponding to the four task queues is recorded as Quantum_Q L0 、Quantum_Q L1 、Quantum_Q L2 、Quantum_Q L3 , then the quantity relationship of the time slices corresponding to the four task queues satisfies: Quantum_Q L3 <Quantum_Q L2 <Quantum_Q L1 <Quantum_Q L0 .
[0159] In the embodiment of the present application, tasks are placed in task queues of different priorities according to the criticality of the tasks, that is, the task criticality of the above-mentioned tasks is used to determine the task queue to which the tasks belong (the queue or task queue in the following embodiments all refers to the queue of tasks, which is the same concept). It can also be understood that tasks in task queues of different priorities have different processing priorities, the priority of a task is related to the level of the task criticality of the task, and tasks in the same task queue have the same processing task priority.
[0160] It should be understood that as the criticality of a task increases, the priority of the task queue to which the task belongs increases, that is, when the task criticality of a task increases, the task is moved to a queue with a higher priority, and when the task criticality of a task decreases, the task is moved to a queue with a lower priority.
[0161] In one implementation, there may be a correspondence between the task criticality of a task and the priority of a task queue. For example, the task criticality includes four levels, and the task queue also includes four task queues of different priorities. Then, based on the correspondence between the task criticality and the priority, the task is divided into the task queue corresponding to the task criticality.
[0162] In another implementation, if the task criticality is in binary form (i.e., the high task criticality of 1 and the low task criticality of 0 mentioned above), when a new task is obtained, the task is divided into the corresponding queue according to the initial criticality of the task. When the initial task criticality is high task criticality of 1, the task is divided into the task queue with the highest priority (such as the Quantum_Q L3 ); When the initial task criticality is low task criticality 0, the task is assigned to the task queue with the lowest priority (such as the above Quantum_Q L0 ).
[0163] Optionally, in an embodiment of the present application, the tasks in the task queue may be tasks of the same application or tasks of multiple applications, without specific limitation.
[0164] Combined with the above description of task criticality, task queues, and time slices, it can be seen that the higher the priority of the task queue, the fewer time slices the queue corresponds to, the higher the task criticality of the tasks in the queue, and the lower the delay tolerance for processing and checking tasks.
[0165] S903 . During the process of executing the task by the computing unit, the partition scheduler updates the queue to which the task belongs, so that the processing core executes the task according to the priority of the queue to which the task belongs.
[0166] In the embodiment of the present application, during the task scheduling process, the task criticality and execution status of the task may change. For example, if a task consumed more time slices (in the computing state) at the previous moment, but consumes fewer time slices (in the memory access state) at the current moment, the task criticality of the task will increase. For another example, if a task consumed fewer time slices (in the memory access state) at the previous moment, but consumes more time slices (in the computing state) at the current moment, the task criticality of the task will decrease.
[0167] In one implementation, in combination with FIG10 , as shown in FIG13 , the process of the partition scheduler updating the queue to which the task belongs includes S9031 - S9032 .
[0168] S9031. Update the task criticality of the task.
[0169] As you can understand, the criticality of a task may change during execution, for example, when it switches between memory access and computation. The criticality of a task is determined based on the time slice consumed by the task and the core latency, that is, the core's tolerance for task latency.
[0170] S9032. Update the queue to which the task belongs based on the task criticality.
[0171] In an embodiment of the present application, task criticality is the basis for determining whether to switch queues for a task. Taking a task as an example, when the task criticality of the task increases, the partition scheduler moves the task from the current queue to a queue with a higher priority than the current queue; when the task criticality of the task decreases, the partition scheduler moves the task from the current queue to a queue with a lower priority than the current queue.
[0172] In some embodiments, a task is associated with the ID of the processing core to which it was initially assigned, and time slices of the same processing core are preferentially allocated. When a task encounters a cache miss and initiates memory access, the latency of the task increases, causing the criticality of the task to increase. The task is then moved into a high-priority queue and scheduled for execution with priority.
[0173] Optionally, when the task criticality increases, the task is moved to a high-priority task queue step by step, and when the task criticality decreases, the task is moved to a low-priority task queue step by step. For example, a task is currently in queue Q L1 If the criticality of the task increases, the task will be removed from the queue Q L1 Move into queue Q L2 In the process of task execution, if the criticality of subsequent tasks continues to increase, the task will be moved from Q L2 Move into queue Q L3 middle.
[0174] In another implementation, the queue to which the above-mentioned update task belongs may also include the following situations.
[0175] Case 1: When the time slice corresponding to the task is consumed and the task has not yet finished running, the partition scheduler moves the task from the current queue to a queue with a lower priority than the current queue.
[0176] In the embodiment of the present application, if the time slice corresponding to a task has been consumed but the task has not yet completed, it means that the computing resources required to run the task are insufficient. Since the lower-priority queues have more time slices, the task is moved from the current queue to a queue with a lower priority than the current queue. This allows the task to obtain more computing resources, thus ensuring that the task runs quickly and improving the efficiency of task scheduling. Furthermore, switching tasks between queues within a partition based on their running status can improve the overall resource utilization of the processing cores within the partition.
[0177] Case 2: When the task waiting timeout is serious, the partition scheduler moves the task from the current queue to a queue with a higher priority than the current queue.
[0178] In the embodiment of the present application, when a task in a low-priority queue times out, it indicates that the task in the low-priority queue cannot be allocated computing resources (i.e., time slices) in a timely manner, resulting in starvation. The task is moved from the current queue to a queue with a higher priority than the current queue. This can alleviate task timeouts within the partition and enable the task to run smoothly. Furthermore, by switching tasks between queues within the partition based on their running status, the resource utilization of the processing cores within the partition can be improved.
[0179] Case 3: When the time slice corresponding to the task queue is not consumed and the processing core executing the task is released, the queue to which the task belongs remains unchanged.
[0180] The process of task scheduling within a partition described in S901-S903 above is a dynamic scheduling process based on task criticality. The partition scheduler of each partition performs autonomous task scheduling for the partition, and schedules based on the task criticality of the tasks within the partition. The scheduling granularity is finer, so that the task scheduling within the partition is balanced, which can improve the resource utilization of the processing core within the partition; and for each task, the overhead required to obtain the task criticality is relatively small, which can significantly reduce the scheduling overhead required for task scheduling.
[0181] Furthermore, compared with the static scheduling method, the embodiment of the present application updates the queue based on the task criticality, does not rely on prior knowledge, and does not require any profiling performance collection data, thereby greatly simplifying development and engineering optimization.
[0182] The following describes the process of cross-partition task scheduling in detail. The main idea of cross-partition task scheduling is to determine whether cross-partition task scheduling needs to be initiated based on the partition criticality of each partition. If cross-partition task scheduling is required, the partition schedulers of the two partitions to be cross-partitioned perform cross-partition task scheduling.
[0183] Optionally, when the number of partitions in a computing system is small (e.g., the number of partitions is less than or equal to 4), cross-partition task scheduling between partitions is performed through point-to-point negotiation between the partition schedulers, without the need to cascade the upper-level federal scheduler (i.e., the L2 scheduler). When the number of partitions in a computing system is large (e.g., the number of partitions is greater than 4), a higher-level federal scheduler (i.e., the L2 scheduler) can be cascaded, and the L2 scheduler aggregates the criticality of each partition and determines whether to initiate cross-partition task scheduling, as well as which two partitions to perform cross-partition task scheduling between.
[0184] For the scenario of cross-partition task scheduling, as shown in FIG14 , the task scheduling method provided in the embodiment of the present application includes S1201 - S1206 .
[0185] S1201: Determine the partition criticality of each partition in a plurality of partitions of a computing system.
[0186] Partition criticality is used to indicate the latency tolerance of computing units within a partition for scheduled tasks, and partition criticality is the basis for determining whether to switch tasks across partition queues (i.e., the basis for whether to schedule tasks across partitions).
[0187] In an embodiment of the present application, the partition criticality of each partition is determined by the partition scheduler corresponding to the partition. The partition criticality is the proportion of the core criticality values of multiple processing cores within a partition whose core criticality values are greater than a first threshold. The core criticality is determined based on the latency tolerance of the processing cores to the tasks running on the processing cores. For details on the process of determining the partition criticality, please refer to the description of the partition criticality in the above technical terminology introduction and will not be repeated here.
[0188] S1202: Calculate criticality differences of the multiple partitions based on the partition criticalities of the multiple partitions.
[0189] The criticality difference is used to indicate the difference between the partition criticalities of multiple partitions. The criticality difference of multiple partitions can measure whether the task scheduling in the entire computing system is balanced.
[0190] In one implementation, when an L2 scheduler does not exist in a computing system, after the partition scheduler for each partition determines the partition criticality, multiple partition schedulers communicate point-to-point, transmitting the partition criticality of each partition they manage. This allows each partition scheduler to aggregate the partition criticality of all partitions, and then one or more schedulers calculate the criticality difference between multiple partitions based on the multiple partition criticalities. For example, the partition schedulers corresponding to the partition with the highest and lowest partition criticality can calculate the criticality difference between multiple partitions and determine whether to initiate cross-partition task scheduling based on the criticality difference.
[0191] In another implementation, when an L2 scheduler exists in the computing system, the L2 scheduler can calculate the criticality differences of multiple partitions. The L2 scheduler obtains the partition criticality of each of the multiple partitions from the partition scheduler (i.e., the L2 scheduler aggregates the partition criticalities of all partitions in the computing system), calculates the criticality differences of the multiple partitions based on the multiple partition criticalities, and then determines whether to initiate cross-partition task scheduling based on the criticality differences.
[0192] Optionally, a method for the L2 scheduler to summarize the partition criticalities of all partitions may be: the partition scheduler of each partition reports the partition criticality of the partition it manages to the L2 scheduler respectively.
[0193] Alternatively, another method for the L2 scheduler to aggregate the partition criticalities of all partitions may be to use a Ring-All Reduce algorithm to aggregate the partition criticalities of multiple partitions. For example, referring to FIG15 , assume that the computing system includes n partitions, numbered from 0 to n-1, and the n partitions are denoted as Partition 0 to Partition n-1. The n partition schedulers corresponding to the n partitions are denoted as Scheduler L1_0 to Scheduler L1_n-1, respectively. The partition criticalities of the multiple partitions are transferred between the partition schedulers using the Ring-All Reduce algorithm.
[0194] For example, scheduler L1_0 in Figure 15 sets the partition criticality C of partition 0 partition {0} is reported to scheduler L1_1, which aggregates the partition criticality of partition 0 and partition 1 to obtain C partition {0,1}, and report it to the scheduler L1_2; the scheduler L1_2 summarizes the partition criticality of partition 0, the partition criticality of partition 1, and the partition criticality of partition 2 to obtain C partition {0,1,2}, and report it to the next L1 scheduler; and so on, scheduler L1_n-2 receives C sent by scheduler L1_n-3 partition After {0,1,2,n-3}, the previous partition criticality and the partition criticality of partition n-2 are summed up to get C partition {0,1,2,n-2}, and report it to the scheduler L1_n-1; finally, the scheduler L1_n-1 summarizes the received partition criticality with the partition criticality of partition n-1 to obtain C partition {0,1,2,n-1}, and report it to the upper-level L2 scheduler.
[0195] The Ring-All Reduce algorithm is an efficient method for information aggregation. Using this method to aggregate partition criticality to the L2 scheduler can save the amount of data for summarizing partition criticality and reduce overhead.
[0196] Optionally, the difference in criticality of the plurality of partitions is the covariance of the criticality of the plurality of partitions. The covariance of the criticality of the plurality of partitions is calculated as follows:
[0197] Among them, COV represents the criticality difference, σ represents the covariance of the partition criticality, C represents the mean of the partition criticality, and C i represents the partition criticality of partition i, and n represents the number of partitions in the computing system.
[0198] S1203: When the criticality difference between the multiple partitions is greater than a second threshold, determine that cross-partition task scheduling needs to be initiated.
[0199] It should be understood that when the criticality of multiple partitions differs significantly, this indicates an imbalance in task scheduling, or load, across the entire computing system. In other words, when the criticality of multiple partitions differs significantly, some partitions in the computing system may have sufficient computing resources while others may have insufficient resources. Therefore, cross-partition task scheduling can be used to ensure balanced scheduling of tasks across partitions.
[0200] When it is determined that cross-partition task scheduling needs to be initiated, the following S1204 - S1205 are executed to complete the cross-partition task scheduling.
[0201] S1204 . The second partition scheduler corresponding to the second partition moves the task in the first task queue of the first partition among the multiple partitions to the task queue of the second partition among the multiple partitions, so that the computing unit in the second partition processes the task.
[0202] S1205 . The first partition scheduler corresponding to the first partition moves the task in the task queue of the second partition among the multiple partitions into the task queue of the first partition, so that the computing unit in the first partition processes the task.
[0203] The first partition and the second partition are two partitions from among the multiple partitions of the computing system that require cross-partition task scheduling, i.e., two partitions that require bilateral peer scheduling. Among the multiple partitions, the difference in partition criticality between the first partition and the second partition is the largest. For example, the first partition corresponds to the partition criticality with the largest median value among the multiple partition criticalities, and the second partition corresponds to the partition criticality with the smallest median value among the multiple partition criticalities. The first task queue is the highest priority queue among the multiple task queues of the first partition, and the second task queue is the lowest priority queue among the multiple task queues of the second partition.
[0204] Optionally, the method for determining the first partition and the second partition from multiple partitions includes the following two implementation manners.
[0205] In one implementation, when there is no L2 scheduler in the computing system, when cross-partition task scheduling needs to be initiated, each partition scheduler can obtain the partition criticality of all partitions. Therefore, each partition scheduler can determine whether the partition in which it is located is the partition with the highest or lowest partition criticality, and then the two partition schedulers corresponding to the partition with the highest partition criticality and the partition with the lowest partition criticality interact to obtain tasks from each other's task queue to the local task queue.
[0206] For example, a computing system includes four partitions (partition 1, partition 2, partition 3, and partition 4). The partition with the highest partition criticality is partition 1, and the partition with the lowest partition criticality is partition 4. After the partition scheduler of partition 1 obtains the partition criticalities of each of the four partitions, the partition scheduler of partition 1 can determine, based on the four partition criticalities, that partition 1, with the highest partition criticality, is the first partition. Similarly, after the partition scheduler of partition 4 obtains the partition criticalities of each of the four partitions, the partition scheduler of partition 4 can determine, based on the four partition criticalities, that partition 4, with the lowest partition criticality, is the second partition.
[0207] In another implementation, when an L2 scheduler exists in a computing system, when cross-partition task scheduling needs to be initiated, the L2 scheduler determines the partition with the highest partition criticality from multiple partitions as the first partition, and determines the partition with the lowest partition criticality as the second partition, and the L2 scheduler sends bilateral scheduling instructions to the first partition scheduler and the second partition scheduler respectively, and the bilateral scheduling instructions are used to instruct cross-partition task scheduling of the first partition and the second partition.
[0208] In an embodiment of the present application, after the L2 scheduler summarizes the partition criticalities of all partitions, it determines the partition with the highest partition criticality as the first partition, and determines the partition with the lowest partition criticality as the second partition; and the L2 scheduler sends a first bilateral scheduling instruction to the L1 scheduler of the first partition, and the first bilateral scheduling instruction instructs the first partition and the second partition to perform cross-partition scheduling, and the second bilateral scheduling instruction may include first indication information, and the first indication information is used to indicate that the first partition is the partition with the highest partition criticality; the L2 scheduler sends a second bilateral scheduling instruction to the second partition criticality of the second partition, and the second bilateral scheduling instruction instructs the first partition and the second partition to perform cross-partition scheduling, and the second bilateral scheduling instruction may include second indication information, and the second indication information is used to indicate that the second partition is the partition with the lowest partition criticality; then, the first partition scheduler and the second partition scheduler interact to obtain tasks from each other's task queue to the local task queue.
[0209] The above-mentioned second partition scheduler moves the tasks in the first task queue of the first partition (the task queue with the highest priority) into the task queue of the second partition, and the task is processed by the processing core of the second partition (in order to distinguish it from other tasks, the task is referred to as the first task below); and the first scheduler moves the tasks in the second task queue of the second partition (the task queue with the lowest priority) into the task queue of the first partition, and the task is processed by the processing core of the first partition (the task is referred to as the second task below).
[0210] Optionally, the first task is the task at the end of the first task queue of the first partition. After the second partition scheduler receives the bilateral scheduling instruction, the second partition scheduler does not first accept new tasks, but instead preferentially obtains the first task from the first task queue of the first partition and moves the first task to the task queue with the highest priority in the task queue of the second partition.
[0211] Optionally, the second task is the task at the end of the second task queue of the second partition. The first partition scheduler moves the second task to the task queue with the lowest priority among the task queues of the first partition.
[0212] In some embodiments, if the first partition and the second partition to be cross-partitioned task scheduling share memory or cache, then during the cross-partition task scheduling process, the second partition can move the first task in the first partition to any queue in the task queue of the second partition, so as not to cause additional memory access and reduce overhead.
[0213] In some embodiments, if the first partition and the second partition to be scheduled for cross-partition tasks do not have shared memory, for the second partition, the first task is a new task, and data needs to be prepared before the new task is executed, so there is a memory access operation. In this case, the second partition will preferentially move the first task in the first partition to the high-priority queue of the second partition.
[0214] For example, referring to FIG16 , it is assumed that the task queue of the first partition includes four queues with different priorities, and the four queues from low to high priority are Q L0 , Q L1 , Q L2 , Q L3 The task queue of the second partition also includes 4 queues with different priorities. The four queues from low to high priority are Q' L0 , Q' L1 , Q' L2 , Q' L3 When cross-partition task scheduling is required and the criticality of the task in the first partition is higher than that of the task in the second partition, the second partition scheduler selects the task with the highest priority Q from the task queue of the first partition. L3 Get task T a , and the task T a Move to the task queue with the highest priority Q' in the second partition L3 and the first partition scheduler obtains task T from the task queue Q0 with the lowest priority in the second partition i , and the task T i Move to the lowest priority Q' in the task queue of the first partition L0 middle.
[0215] Based on the above description, in the process of cross-partition task scheduling, the partition with high partition criticality obtains tasks from the low-priority queue of the partition with low partition criticality, which can reduce the criticality of the high partition; the partition with low partition criticality obtains tasks from the high-priority queue of the partition with high partition criticality, which can increase the criticality of the low partition, thereby reducing the difference between the partition criticalities of the two partitions and achieving balanced task scheduling between the two partitions.
[0216] It should be noted that in some cases, the differences between the partition criticalities of various partitions of the computing system are small (i.e., the differences in the criticalities of multiple partitions are less than the second threshold), but there are some partitions where task timeouts are very serious (such as a very high task timeout coefficient). At this time, the task schedulers of other partitions can obtain timed-out tasks (task delays exceed the threshold) from the task queue of the partition where the task has timed out, and move them into a high-priority queue. In this way, the problem of task timeouts can be solved.
[0217] Optionally, in an embodiment of the present application, for a heterogeneous multi-core computing system, when cross-partition task scheduling is performed, cross-partition task scheduling can be performed between partitions with the same processor type, and cross-partition task scheduling can also be performed between partitions with different processor types.
[0218] For example, for a computing system with heterogeneous processors including a 48-core CPU (the CPU also has vector processing or matrix processing capabilities) and four 128-core GPUs, the processor cores are divided into 10 partitions, namely {P0, P1, ..., P9}. The 48-core CPU is divided into two partitions, P0 and P1, each containing 24 processing cores. The four 128-core GPUs are divided into eight partitions, P2 to P9, each containing 64 processing cores. When scheduling tasks across partitions, cross-partition scheduling can be performed between P0 and P1, cross-partition scheduling can be performed between P2 to P9, and cross-partition task scheduling can also be performed between partitions of the heterogeneous processors, for example, cross-partition task scheduling can be performed between P1 and P2.
[0219] It should be noted that when cross-partition task scheduling is performed between partitions divided by heterogeneous processors, it is necessary to generate machine code modules suitable for processor class 1 (such as CPU) and processor class 2 (GPU) to run the same task through a heterogeneous compiler. In other words, two executable libraries suitable for processor class 1 (CPU) and processor class 2 (GPU) are pre-compiled. For example, for task 1, it is necessary to pre-compile executable library 1 for the CPU to execute task 1 and executable library 2 for the GPU to execute task 1. When task 1 is originally executed by the CPU of the partition P1, task 1 is executed based on executable library 1. When P1 and P2 perform cross-partition task scheduling, task 1 is subsequently executed by the GPU of partition 2. When the GPU executes task 1, task 1 is executed based on executable library 2. This implements task scheduling across heterogeneous cores and can improve the processing core utilization of the computing system.
[0220] S1206: When the criticality difference is less than or equal to the second threshold, each partition scheduler schedules tasks within each partition.
[0221] It should be understood that when the criticality differences among multiple partitions are small, it indicates that for the entire computing system, the global task scheduling is relatively balanced, and there is no need for cross-partition task scheduling. Autonomous task scheduling is performed within each partition.
[0222] The detailed description of task scheduling within a partition can refer to the detailed description of S901-S903 and related contents in the above embodiment, which will not be repeated here.
[0223] Based on the above content, the task scheduling method provided in the embodiment of the present application is a dynamic task scheduling method, which can partition the computing system and perform autonomous task scheduling within the partition, which can improve the resource utilization of the processing cores within the partition; when the criticality of each partition is greatly different, cross-partition task scheduling is performed, which can improve the resource utilization of the computing system.
[0224] Furthermore, this task scheduling method is applicable to a variety of regular or irregular tasks and has good generalization. For example, it can achieve balanced parallel execution across multi-core processors for various CPU parallel tasks, regardless of whether they are computationally intensive, memory-intensive, have complex task dependencies, or have complex control flows (such as QR decomposition, Cholesky decomposition, and Heat Diffusion), thereby improving the overall performance of the computing system.
[0225] Optionally, in an embodiment of the present application, in a many-core computing system, the computing system has many partitions, the L1 scheduler cascades multiple L2 schedulers, and multiple L2 schedulers can also cascade L3 schedulers. The L2 scheduler and the L3 scheduler are responsible for coarse-grained balanced task scheduling, and coarser-grained scheduling is achieved through L3 scheduling.
[0226] Exemplarily, a computing system includes multiple computing nodes, each computing node includes multiple processors, and each processor includes multiple processing cores. In this case, an L1 scheduler manages partitions of multiple cores, an L2 scheduler manages multiple L1 schedulers (equivalent to an L2 scheduler managing multiple partitions, and these multiple partitions can be defined as a domain), and an L3 scheduler manages multiple L2 schedulers (equivalent to an L3 scheduler managing multiple domains of the computing system). Optionally, a domain can be a cabinet. For example, multiple servers (for example, 8) can be defined as a cabinet, and one L2 is responsible for managing multiple servers. When the number of cabinets is large, an L3 scheduler is cascaded on top of multiple L2 schedulers, and the L3 scheduler is responsible for the global scheduling of the computing system.
[0227] Each L1 caller is responsible for scheduling tasks within the partition it manages based on the task criticality; after the L1 scheduler determines the partition criticality, it reports the partition criticality to the L2 scheduler. The L2 scheduler determines whether cross-partition task scheduling is required between the multiple partitions managed by the L2 scheduler based on the aggregated multiple partition criticalities; further, each L2 scheduler can determine the domain criticality of the domain managed by the L2 scheduler based on the multiple partition criticalities it collects, and report the domain criticality to the L3 scheduler. The L3 scheduler determines the difference between the domain criticality of the multiple domains managed by the L3 caller based on the aggregated multiple domain criticalities, and determines whether cross-domain task scheduling is required.
[0228] The above-mentioned method for determining and reporting domain criticality may be similar to the above-mentioned method for determining and reporting partition criticality. The method for the L3 scheduler to determine the difference between domain criticalities may also be similar to the process for the L2 scheduler to determine the difference between partition criticalities. For details, please refer to the relevant description in the above embodiments.
[0229] In an embodiment of the present application, when the L3 scheduler determines, based on multiple domain criticalities, that the difference between the multiple domain criticalities exceeds a threshold, that cross-domain task scheduling is required, the L3 scheduler determines, from the multiple domains, two domains with the largest difference in domain criticality, then selects a partition from each of the two domains, and then performs cross-partition task scheduling between the two partitions. For example, a partition with the largest partition criticality (referred to as the first target partition) is selected from the domain with the largest domain criticality, and a partition with the smallest partition criticality (referred to as the second target partition) is selected from the domain with the smallest domain criticality, and then instructs the first target partition and the second target partition to perform cross-partition task scheduling.
[0230] The L1 scheduler, L2 scheduler, and L3 scheduler are responsible for scheduling at different granularities in the computing system, which can make the task scheduling of the computing system more balanced and improve the resource utilization of the computing system.
[0231] An embodiment of the present application also provides a computing system, as shown in Figure 17, which includes multiple partition schedulers, each partition scheduler includes a determination module 1701 and a scheduling module 1702, the determination module 1702 is used to execute S901, S9012, S902, and S1203 in the above method embodiment, and the scheduling module 1702 is used to execute S903 (including S9031-S9032), S1201, S1204, and S1205 in the above method embodiment.
[0232] Optionally, each of the multiple partition schedulers includes a calculation module 1703 and an acquisition module 1704 , the calculation module 1703 is used to execute S1202 in the above method embodiment, and the acquisition module 1704 is used to execute S9012 in the above method embodiment.
[0233] Optionally, the computing system further includes a federated scheduler, which is used to manage multiple partition schedulers. The federated scheduler includes an acquisition module 1705, a calculation module 1706, a determination module 1707, and a sending module 1708. Acquisition module 1705 is used to acquire partition criticalities of multiple partitions, calculation module 1706 is used to calculate the criticality differences between partitions, determination module 1707 is used to determine whether to initiate cross-partition task scheduling, and sending module 1708 is used to send bilateral scheduling instructions to the partition schedulers.
[0234] For more details on how the above modules implement the above functions, please refer to the descriptions in the previous method embodiments, which will not be repeated here. The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments.
[0235] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions in accordance with the embodiments of the present application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (eg, a floppy disk, a magnetic disk, a magnetic tape), an optical medium (eg, a digital video disc (DVD)), or a semiconductor medium (eg, a solid state drive (SSD)).
[0236] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0237] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0238] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0239] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0240] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0241] The above is only a specific embodiment of the present application, but the scope of protection of this application is not limited to this. Any changes or substitutions within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A task scheduling method, characterized in that: The method is used for scheduling tasks for a computing system, the computing system comprising a plurality of computing units, the plurality of computing units being divided into a plurality of partitions, each of the plurality of partitions corresponding to a partition scheduler, each partition comprising a plurality of task queues with different priorities set based on task criticality, the task criticality being determined based on the latency tolerance of a computing unit running a task to the task; the method comprising: Determine a partition criticality of each of the plurality of partitions, the partition criticality being used to indicate a latency tolerance of a computing unit within the partition to a scheduled task; When it is determined that cross-partition task scheduling needs to be initiated according to the partition criticality of each partition in the multiple partitions, the second partition scheduler moves the task in the first task queue of the first partition in the multiple partitions to the task queue of the second partition in the multiple partitions, so that the computing unit in the second partition processes the task; the second partition scheduler is the partition scheduler corresponding to the second partition; The first partition scheduler moves tasks in the second task queue of the second partition among the multiple partitions into the task queue of the first partition so that the computing unit in the first partition processes the tasks; the first partition scheduler is the partition scheduler corresponding to the first partition; the priority of the first task queue is different from the priority of the second task queue.
2. The method according to claim 1, characterized in that The plurality of computing units include processing cores of one or more processors; wherein the one or more processors include a heterogeneous first processor and / or a second processor.
3. The method according to claim 1 or 2, characterized in that: The partition criticality is the proportion of the number of processing cores whose core criticality values are greater than a first threshold among the core criticalities of multiple processing cores in a partition; the core criticality is determined based on the latency tolerance of the processing core to the task running on the processing core.
4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Calculating a criticality difference among the plurality of partitions according to the partition criticalities of the plurality of partitions; the criticality difference is used to indicate a difference between the plurality of partition criticalities of the plurality of partitions; When the criticality difference is greater than a second threshold, determining that cross-partition task scheduling needs to be initiated; When the criticality difference is less than or equal to the second threshold, each of the partition schedulers schedules the tasks in each partition.
5. The method according to any one of claims 1 to 4, characterized in that: The multiple partitions correspond to a federated scheduler, and the federated scheduler is used to manage the multiple partition schedulers; Determining the partition criticality of each of the plurality of partitions includes: The federated scheduler obtains the partition criticality of each of the plurality of partitions from the partition scheduler; The method further comprises: When the federated scheduler determines that cross-partition task scheduling needs to be initiated according to the partition criticality of each partition in the multiple partitions, the federated scheduler determines the first partition and the second partition from the multiple partitions; wherein the first partition is the partition corresponding to the partition criticality with the largest value among the multiple partition criticalities, and the second partition is the partition corresponding to the partition criticality with the smallest value among the multiple partition criticalities; The federated scheduler sends a bilateral scheduling instruction to the first partition scheduler and the second partition scheduler respectively; the bilateral scheduling instruction is used to instruct to perform cross-partition task scheduling on the first partition and the second partition.
6. The method according to claim 5, characterized in that The priority of the task queue increases with the increase of the criticality of the task, and the number of time slices allocated to queues of different priorities is different; the criticality of the task increases with the decrease of the delay tolerance of the processing core to the task; The first task queue is a queue with the highest priority among the task queues of the first partition, and the second task queue is a queue with the lowest priority among the task queues of the second partition.
7. The method according to any one of claims 1 to 6, characterized in that: Each of the partition schedulers schedules tasks in each partition, including: For any one of the multiple tasks in each partition, during the process of the computing unit executing the task, the partition scheduler updates the queue to which the task belongs, so that the processing core executes the task according to the priority of the queue to which the task belongs.
8. The method according to claim 7, characterized in that The updating of the queue to which the task belongs includes: updating the task criticality of the task; According to the task criticality of the task, the queue to which the task belongs is updated; wherein, when the task criticality of the task increases, the partition scheduler moves the task from the current queue to a queue with a priority higher than the priority of the current queue; when the task criticality of the task decreases, the partition scheduler moves the task from the current queue to a queue with a priority lower than the priority of the current queue.
9. The method according to any one of claims 1 to 8, characterized in that: The method further comprises: Each of the partition schedulers obtains task dependencies between multiple tasks in a partition; The partition scheduler determines the initial task criticalities of the plurality of tasks respectively according to the task dependency depths indicated by the task dependency relationships; The partition scheduler determines initial queues for the multiple tasks respectively from queues of different priorities of the partition according to the initial task criticalities of the multiple tasks.
10. A computing system, characterized in that: The method comprises a plurality of partition schedulers, wherein the plurality of computing units of the computing system are divided into a plurality of partitions, each of the plurality of partitions corresponds to a partition scheduler, and each partition comprises a plurality of task queues with different priorities set based on task criticality, wherein the task criticality is determined based on the latency tolerance of a computing unit running a task to the task; Each partition scheduler includes a determination module and a scheduling module; A determination module of each of the plurality of partition schedulers, configured to determine a partition criticality of each of the plurality of partitions, wherein the partition criticality is used to indicate a latency tolerance of a computing unit within the partition to a scheduled task; a scheduling module of the second partition scheduler, configured to, when it is determined that cross-partition task scheduling needs to be initiated according to the partition criticality of each partition among the multiple partitions, move tasks in a first task queue of a first partition among the multiple partitions to a task queue of a second partition among the multiple partitions, so that a computing unit in the second partition processes the tasks; The second partition scheduler is a partition scheduler corresponding to the second partition; The scheduling module of the first partition scheduler is used for, when it is determined that cross-partition task scheduling needs to be initiated according to the partition criticality of each partition of the multiple partitions, moving tasks in the second task queue of a second partition of the multiple partitions into the task queue of the first partition, so that the computing unit in the first partition processes the tasks; The first partition scheduler is a partition scheduler corresponding to the first partition; the priority of the first task queue is different from the priority of the second task queue.
11. The computing system according to claim 10, characterized in that: The plurality of computing units include processing cores of one or more processors; wherein the one or more processors include a heterogeneous first processor and / or a second processor.
12. The computing system according to claim 10 or 11, characterized in that: The partition criticality is the proportion of the number of processing cores whose core criticality values are greater than a first threshold among the core criticalities of multiple processing cores in a partition; the core criticality is determined based on the latency tolerance of the processing core to the task running on the processing core.
13. The computing system according to any one of claims 10 to 12, characterized in that: Each of the plurality of partition schedulers comprises a computing module; The calculation module of the partition scheduler is used to calculate the criticality difference of the multiple partitions according to the partition criticality of the multiple partitions; The criticality difference is used to indicate the difference between the plurality of partition criticalities of the plurality of partitions; The determination module of the partition scheduler is further used to determine that cross-partition task scheduling needs to be initiated when the criticality difference is greater than a second threshold; The scheduling module of the partition scheduler is used to schedule the tasks in each partition when the criticality difference is less than or equal to the second threshold.
14. The computing system according to any one of claims 10 to 12, characterized in that: The computing system further includes a federated scheduler, the multiple partitions correspond to one federated scheduler, and the federated scheduler is used to manage the multiple partition schedulers; the federated scheduler includes an acquisition module, a calculation module, a determination module, and a sending module; The acquisition module of the federated scheduler is used to acquire the partition criticality of each of the multiple partitions from the partition scheduler; The calculation module of the federated scheduler is used to calculate the criticality difference of the plurality of partitions according to the partition criticalities of the plurality of partitions; The criticality difference is used to indicate the difference between the plurality of partition criticalities of the plurality of partitions; The determination module of the federated scheduler is used to determine the first partition and the second partition from the multiple partitions when it is determined that cross-partition task scheduling needs to be initiated when the criticality difference is greater than a second threshold; wherein the first partition is the partition corresponding to the partition criticality with the largest median value of the criticalities of the multiple partitions, and the second partition is the partition with the smallest median value of the criticalities of the multiple partitions The partition corresponding to the critical degree; The sending module of the federated scheduler is used to send bilateral scheduling instructions to the first partition scheduler and the second partition scheduler respectively; the bilateral scheduling instructions are used to instruct to perform cross-partition task scheduling on the first partition and the second partition.
15. The computing system according to claim 14, characterized in that: The priority of the task queue increases with the increase of the criticality of the task, and the number of time slices allocated to queues of different priorities is different; the criticality of the task increases with the decrease of the delay tolerance of the processing core to the task; The first task queue is a queue with the highest priority among the task queues of the first partition, and the second task queue is a queue with the lowest priority among the task queues of the second partition.
16. The computing system according to any one of claims 10 to 15, characterized in that: The scheduling module of each partition scheduler is specifically used to update the queue to which any one of the multiple tasks in each partition belongs during the process of the computing unit executing the task, so that the processing core executes the task according to the priority of the queue to which the task belongs.
17. The computing system according to claim 16, characterized in that: A scheduling module is specifically used to update the task criticality of the task; and update the queue to which the task belongs according to the task criticality of the task; wherein, when the task criticality of the task increases, the partition scheduler moves the task from the current queue to a queue with a priority higher than the priority of the current queue; when the task criticality of the task decreases, the partition scheduler moves the task from the current queue to a queue with a priority lower than the priority of the current queue.
18. The computing system according to any one of claims 10 to 17, characterized in that: Each partition scheduler also includes an acquisition module; An acquisition module of each partition scheduler is used to acquire task dependencies between multiple tasks in a partition; The determination module of each partition scheduler is also used to determine the initial task criticality of the multiple tasks respectively according to the task dependency depth indicated by the task dependency relationship; and according to the initial task criticality of the multiple tasks, determine the initial queues for the multiple tasks respectively from the multiple queues of different priorities of the partition.
19. A computing device, characterized in that The method comprises a memory and at least one processor connected to the memory, wherein the memory is used to store computer program code, and the computer program code comprises computer instructions. When the computer instructions are executed by at least one processor, the computing device executes the method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that: Computer instructions are stored, and when the computer instructions are executed on a computer, the method according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
Task scheduling method and device and computing system
CN119917229A
Method and system for balancing physical system resource access between logic partitions
CN101276293A
Task scheduling apparatus and method for embedded operating system
CN101452404A
Zoning scheduling management method of cluster computing resources
CN102902592A
Cluster resource adjustment method and apparatus, and cloud platform
CN108427604A
Cited By
Parallel task graph incremental scheduling method based on OCaml functional language
CN121070543A