Task scheduling methods, task schedulers, processing devices, electronic devices, media
Patent Information
- Application Number
- CN202510686521.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-05-26
AI Technical Summary
然而,相关技术中,抢占式调度需要较大的硬件代价,且任务调度的性能和吞吐率较低;多队列调度的切换粒度过大,任务切换缓慢;线程束调度需要硬件实现复杂的资源控制逻辑,带来的额外开销较大
[0025] The embodiments provided in this disclosure introduce priority comparison, timeout management, and task boundary judgment mechanisms into the task scheduler of the processing device. Based on the priority comparison results, load timeout results, and task boundary judgment results, the scheduling and switching of high and low priority tasks can be realized, achieving thread bundle-level, fine-grained priority scheduling, thereby improving the efficiency of task scheduling and enhancing the overall performance and processing efficiency of the processing device.
Smart Images

Figure CN120631530B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a task scheduling method, a task scheduler, a processing device, an electronic device, and a computer-readable storage medium. Background Technology
[0002] In current hardware systems (such as central processing unit (CPU), graphics processing unit (GPU), etc.), it is usually necessary to break down the hardware system's work objective (such as rendering a frame) into tasks of different processing categories (such as computational category and graphics category) in order to process them.
[0003] During the processing, tasks of different priorities need to be scheduled, with high-priority tasks (such as those used for real-time graphics rendering) being executed first, followed by low-priority tasks (such as those used for background computation) to improve the system's execution efficiency.
[0004] Related task scheduling schemes include preemptive scheduling, multi-queue scheduling, and thread-binding scheduling. However, among these technologies, preemptive scheduling requires significant hardware investment and has low task scheduling performance and throughput; multi-queue scheduling has too large a granularity for switching, resulting in slow task switching; and thread-binding scheduling requires complex resource control logic to be implemented in hardware, leading to substantial additional overhead. Summary of the Invention
[0005] This disclosure provides a task scheduling method, a task scheduler, a processing device, an electronic device, and a computer-readable storage medium.
[0006] In a first aspect, this disclosure provides a task scheduling method, the method comprising a task scheduler applied in a processing device, the task scheduler corresponding to a first task queue, the tasks in the first task queue having the same processing category, the processing category including a computational category and a graphics category, and each task including multiple thread bundles; the method comprising:
[0007] The priority of the first task in the first task queue is compared with the priority of the second task in the second task queue to determine the priority comparison result. The processing categories of the tasks in the first task queue and the second task queue are different.
[0008] Based on the priority comparison result, the load timeout result of the target task queue is determined, wherein the target task queue is a task queue with high priority or equal priority between the first task queue and the second task queue;
[0009] Based on the shared resource usage information of the first task queue, determine the task boundary judgment result of the first task queue;
[0010] Based on the priority comparison result, the load timeout result, and the task boundary judgment result, the scheduling result of the first task currently to be executed in the first task queue is determined;
[0011] If the scheduling result allows the task to be sent, the first task is sent to the backend module so that the backend module allocates hardware resources for the thread bundle of the first task and executes it.
[0012] Secondly, this disclosure provides a task scheduler corresponding to a first task queue, wherein the tasks in the first task queue have the same processing category, the processing category including a computation category and a graphics category, and each task includes multiple thread bundles;
[0013] The task scheduler includes a priority judgment module, a timeout control module, a boundary judgment module, and an output control module.
[0014] The priority determination module is connected to the timeout control module and the output control module, and is configured to: read the first task priority of the first task queue and the second task priority of the second task queue respectively; compare the first task priority and the second task priority, and output the priority comparison result;
[0015] The timeout control module is connected to the output control module and is configured to: determine the load timeout result of the tasks in the target task queue based on the priority comparison result, wherein the target task queue is a task queue with higher priority or equal priority in the first task queue and the second task queue;
[0016] The boundary judgment module is connected to the output control module and is configured to: determine the task boundary judgment result of the first task queue based on the shared resource usage information of the first task queue;
[0017] The output control module is configured to: determine the scheduling result of the first task currently to be executed in the first task queue based on the priority comparison result, the load timeout result, and the task boundary judgment result; and, if the scheduling result allows the first task to be executed, issue the first task so that the backend module allocates hardware resources to the thread bundle of the first task and executes it.
[0018] Thirdly, this disclosure provides a processing apparatus, which includes a front-end module, multiple task schedulers, and a back-end module, wherein the task schedulers include those described above, wherein...
[0019] The front-end module is connected to multiple task schedulers and is configured to: package threads of the same processing category in the received work targets into thread bundles to obtain tasks including multiple thread bundles; and send the tasks to the corresponding task queues according to the processing category of the tasks.
[0020] Multiple task schedulers are connected to the backend module, and each task scheduler corresponds to a task queue;
[0021] The task scheduler is configured to: receive tasks from the corresponding first task queue and obtain the task priorities of multiple task queues; determine the scheduling result of the first task currently to be executed in the first task queue; and, if the scheduling result allows the task to be sent, send the first task to the backend module.
[0022] The backend module is configured to allocate hardware resources for the thread bundle of the first task and execute the thread bundle of the first task.
[0023] Fourthly, this disclosure provides an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores one or more computer programs executable by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the task scheduling method described above.
[0024] Fifthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the above-described task scheduling method.
[0025] The embodiments provided in this disclosure introduce priority comparison, timeout management, and task boundary judgment mechanisms into the task scheduler of the processing device. Based on the priority comparison results, load timeout results, and task boundary judgment results, the scheduling and switching of high and low priority tasks can be realized, achieving thread bundle-level, fine-grained priority scheduling, thereby improving the efficiency of task scheduling and enhancing the overall performance and processing efficiency of the processing device.
[0026] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0027] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which:
[0028] Figure 1 A flowchart of a task scheduling method provided in this embodiment of the disclosure;
[0029] Figure 2 A schematic diagram of a processing apparatus provided in an embodiment of this disclosure;
[0030] Figure 3 A schematic diagram of a task scheduler provided in an embodiment of this disclosure;
[0031] Figure 4 A schematic diagram of a timeout control module provided in an embodiment of this disclosure;
[0032] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation
[0033] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0034] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.
[0035] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.
[0036] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.
[0037] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.
[0038] As mentioned earlier, the relevant technologies in hardware systems include preemptive scheduling, multi-queue scheduling, and thread bundle scheduling.
[0039] In a preemptive scheduling scheme, a high-priority task can preempt a low-priority task at an interrupt or synchronization point; by using context switching, when the system receives a high-priority task, it can save the current task state and context information, switch to high-priority task scheduling and execution; after the high-priority task completes, it can restore the context of the preempted task.
[0040] However, this scheduling scheme requires the system hardware to support context switching mechanisms for different task types, including the ability to save and restore task contexts for graphics rendering and general computing, which requires significant hardware implementation costs. At the same time, context switching requires the entire system hardware to save information and clear the state information of the current task in the system pipeline, which will generate a certain delay, affecting the performance and throughput of task scheduling, resulting in low performance and throughput of task scheduling.
[0041] In a multi-queue scheduling scheme, tasks are typically assigned to different queues within the hardware system (e.g., GPU) based on their processing category, such as graphics task queues and computation task queues. The hardware scheduler performs priority scheduling among multiple task queues. When a task in a queue reaches its defined end boundary, the scheduler can switch to a higher-priority queue to distribute the task.
[0042] However, this scheduling scheme requires other hardware systems to allocate tasks to different queues within this hardware system. This requires the hardware of this system to have the ability to independently parse queue commands. Typically, one or more front-end modules are needed to read the load of different task queues, increasing additional hardware overhead. At the same time, switching by queue requires the system to execute the tasks in the current queue to the end boundary before switching to a higher priority queue. The switching granularity is too large, the task switching is slow, and the switching efficiency is low.
[0043] In thread bundle scheduling schemes, the workload of a hardware system (e.g., a GPU) (e.g., rendering a frame) is typically broken down into threads of different processing categories (e.g., computational and graphics). Threads of the same processing category within the same or different workloads are bundled together into a thread bundle. Multiple thread bundles form a task to fully utilize the instruction and data parallelism capabilities of the hardware system. The hardware scheduler can dynamically adjust task distribution based on priority and the resource requirements of the task, such as register requirements and cache utilization. When a high-priority task is received, hardware resources need to be reserved, and the high-priority task is immediately distributed once available resources can meet its needs.
[0044] However, this scheduling scheme requires a hardware scheduler to arbitrate and schedule tasks across multiple thread bundles of different tasks. This can easily lead to dependency handling anomalies or resource allocation deadlocks, causing system crashes. For example, in general computing tasks, multiple computing groups often need to use shared memory for data communication. This resource may already be occupied by graphics tasks. Scheduling a computing task in this situation can cause the shared memory to be insufficient, resulting in a deadlock. Therefore, this scheduling scheme requires complex resource control logic or thread bundle-level context switching in hardware, incurring significant overhead.
[0045] According to the task scheduling method of this disclosure, a priority comparison, timeout management, and task boundary judgment mechanism are introduced into the task scheduler of the processing device. Based on the priority comparison results, load timeout results, and task boundary judgment results, the method can automatically manage and switch between high and low priority tasks, thereby fully utilizing the hardware computing power of the processing device. According to the embodiments of this disclosure, fine-grained priority scheduling at the thread bundle level can be achieved at a relatively low cost, improving the efficiency of priority scheduling and avoiding deadlock problems that may occur in thread bundle scheduling schemes of related technologies, thereby improving the overall performance and processing efficiency of the processing device.
[0046] The processing device according to the embodiments of this disclosure can be any hardware system itself, or a device or component in a hardware system. The hardware system can be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), or a combination of multiple hardware. This disclosure does not limit the specific type of hardware system corresponding to the processing device.
[0047] The processing apparatus according to embodiments of the present disclosure can process various work objectives (also referred to as work units), such as rendering a frame, background calculation, etc.
[0048] In some possible implementations, the processing device includes a front-end module, multiple task schedulers, and a back-end module. The front-end module is connected to the multiple task schedulers, and the multiple task schedulers are connected to the back-end module. The front-end module is used to parse and generate tasks; the multiple task schedulers are used to perform task scheduling; and the back-end module is used to allocate hardware resources for the thread bundles of tasks issued by the multiple task schedulers and execute them. The front-end module, multiple task schedulers, and back-end module can all be implemented in hardware, and this disclosure does not limit the specific structure of each module of the processing device.
[0049] In some possible implementations, after receiving a task target, the processing device can decompose the task target into threads of different processing categories (e.g., computation and graphics) through the front-end module. Threads of the same processing category in the same or different task targets are packaged into a thread bundle, and multiple thread bundles form a task. The task is then sent to the task queue of the corresponding category.
[0050] In some possible implementations, the task objective can also be broken down into threads of different processing categories by other devices (such as a CPU). After receiving the threads of the task objective, the processing device packages threads of the same processing category from the same or different task objectives into a thread bundle, and multiple thread bundles form a task; then the task is sent to the task queue of the corresponding category. This disclosure does not limit this.
[0051] In this system, each of the multiple task schedulers corresponds to a task queue, and each task queue corresponds to a task priority, such as a first priority (highest priority), a second priority (normal priority), and a third priority (lowest priority). The first priority is applied to latency-sensitive tasks, such as real-time graphics rendering tasks; the second priority is applied to general graphics rendering and ordinary computational tasks; and the third priority is applied to background computational tasks or tasks with non-real-time feedback. It should be understood that those skilled in the art can set the number and type of priorities according to actual circumstances, and this disclosure does not impose any restrictions on this.
[0052] In some possible implementations, the task queue may include tasks and priority configuration commands. When a priority configuration command in the task queue is executed, the contents of the register corresponding to that task queue can be updated to the priority value corresponding to that command. For example, 11 represents high priority, 10 represents normal priority, and 01 represents low priority. Thus, the priorities of tasks following this priority configuration command are all the values configured in the register, until the next priority configuration command is executed, at which point the register is updated again to change the task priority in the task queue.
[0053] In some possible implementations, tasks in the same task queue belong to the same processing category, such as all computational tasks or all graphics tasks. Computational tasks may include computation-related tasks such as deep learning (e.g., network training, network inference), scientific computing (e.g., numerical simulation), and data processing (e.g., image and video processing); graphics tasks may include various tasks in the graphics rendering process.
[0054] In some possible implementations, the graphics category may also include a vertex category and a pixel category. That is, tasks in the same task queue can all be of the vertex category or all of the pixel category. Vertices are the basic elements constituting 3D graphics; they define the shape and position of the graphics. Vertex category tasks are used to manipulate vertex data in 3D space during graphics rendering. Pixel category tasks operate on pixels on the screen during graphics rendering to determine the final color and other attributes of each pixel. It should be understood that those skilled in the art can set the task processing category according to actual circumstances, and this disclosure does not impose any limitations on this.
[0055] For any one of the multiple task schedulers, there is a corresponding task queue (hereinafter referred to as the first task queue) used to schedule the tasks in the first task queue. The tasks in the first task queue have the same processing category.
[0056] Figure 1 A flowchart illustrating a task scheduling method provided in an embodiment of this disclosure. (Refer to...) Figure 1 The method includes:
[0057] In step S11, the priority of the first task in the first task queue and the priority of the second task in the second task queue are compared to determine the priority comparison result. The processing categories of the tasks in the first task queue and the second task queue are different.
[0058] In step S12, the load timeout result of the target task queue is determined according to the priority comparison result. The target task queue is a task queue with high priority or equal priority between the first task queue and the second task queue.
[0059] In step S13, the task boundary judgment result of the first task queue is determined based on the shared resource usage information of the first task queue.
[0060] In step S14, the scheduling result of the first task currently to be executed in the first task queue is determined based on the priority comparison result, the load timeout result, and the task boundary judgment result.
[0061] In step S15, if the scheduling result allows the task to be sent, the first task is sent to the backend module so that the backend module allocates hardware resources for the thread bundle of the first task and executes it.
[0062] For example, the task scheduler of the first task queue can perform the current round of scheduling after the previous round of scheduling is completed, or it can perform scheduling at a certain period of time. This disclosure does not limit this.
[0063] In some possible implementations, at the start of the current round of scheduling, the task scheduler can obtain the first task priority of the first task queue and the second task priority of the second task queue. The tasks in the first and second task queues have different processing categories. If the tasks in the first task queue are graphics-related (vertex or pixel), then the tasks in the second task queue are computation-related; conversely, if the tasks in the first task queue are computation-related, then the tasks in the second task queue are graphics-related (vertex or pixel).
[0064] In some possible implementations, the task priority of the vertex category task queue is the same as that of the pixel category task queue; that is, the vertex category task queue and the pixel category task queue can share task priorities, and the task priorities of the vertex category task queue and the pixel category task queue are stored in the same register. This disclosure does not limit the specific configuration method.
[0065] In some possible implementations, in step S11, the priorities of the first task and the second task can be compared to obtain a priority comparison result. If the priority of the first task is higher than the priority of the second task, the priority comparison result includes the first task queue being the task queue with higher priority and the second task queue being the task queue with lower priority. Let 1 represent high priority or equal priority, and 0 represent low priority. In this case, the priority comparison result is represented as 10. In practice, it is possible that the first task queue has the first priority, and the second task queue has the second or third priority; or the first task queue has the second priority, and the second task queue has the third priority.
[0066] In some possible implementations, if the priority of the first task is lower than that of the second task, the priority comparison result includes the first task queue being a lower-priority task queue and the second task queue being a higher-priority task queue, which can be represented as 01. If the priority of the first task is equal to that of the second task, the priority comparison result includes the first task queue and the second task queue being task queues of equal priority, which can be represented as 11.
[0067] In some possible implementations, when the first task queue has high priority or equal priority, the corresponding task scheduler can prioritize assigning tasks from the first task queue to improve system processing efficiency. When the first task queue has low priority, the corresponding task scheduler can further determine whether to assign tasks from the first task queue based on other conditions. For example, if the high-priority second task queue is empty and no load is assigned within a specified time (i.e., no new tasks arrive), then tasks from the low-priority first task queue can be assigned to more effectively utilize the hardware resources and computing power of the processing device.
[0068] In some possible implementations, in step S12, based on the priority comparison result, a high-priority or equal-priority task queue can be determined, hereinafter referred to as the target task queue; then, based on the status of the tasks in the target task queue, the load timeout result of the target task queue is determined. The task scheduler is equipped with a counter to calculate the duration; the counter's initial value is 0 and it is in a closed state.
[0069] For the task scheduler of the first task queue, if the target task queue includes the first task queue (high priority or equal priority), it can determine whether the first task queue is empty; if the first task queue is empty, a counter is started, the counter counts at a preset counting period, and the count value is output.
[0070] For the current round of scheduling, if the count value does not reach the preset counting threshold, the load timeout result of the first task queue is determined to be that the task has not timed out. For example, a 1-bit signal is output, where 0 indicates no timeout and 1 indicates timeout. In this case, the output load timeout result is 0, and it is necessary to continue waiting for new tasks in the first task queue. If the count value reaches the preset counting threshold, the load timeout result of the first task queue is determined to be that the task has timed out, and the output load timeout result is 1. Tasks in the lower priority second task queue can be assigned.
[0071] The counter's counting period can be a preset clock period, and the counting threshold is a configurable value, such as 1024. This disclosure does not impose any restrictions on the specific duration of the clock period or the specific value of the counting threshold.
[0072] In some possible implementations, if the task execution process requires resources shared by multiple thread bundles from different tasks, such as shared memory for communication between multiple computation groups in a computation-type task, then switching to other task queues can only proceed after all tasks corresponding to the thread bundles using these shared resources have been deployed. In this case, additional checks are required.
[0073] In some possible implementations, in step S13, the task boundary judgment result of the first task queue is determined based on the shared resource usage information of the first task queue. That is, the shared resource usage information of the first task queue can be determined based on the resource requirement information of the thread bundles in each task of the first task queue. If a first shared resource exists, and there are unissued tasks in the first task queue that require the use of the first shared resource, then the task boundary judgment result is that the switching boundary has not been reached, and the tasks in the first task queue need to continue execution; switching to tasks in other task queues is not allowed. In this case, if a forced switch is performed, dependency handling anomalies or resource allocation deadlocks may occur, leading to system crashes. The task boundary judgment result can be output as a 1-bit signal, where 0 indicates that the switching boundary has not been reached, and 1 indicates that the switching boundary has been reached. In this case, the output task boundary judgment result is 0.
[0074] In some possible implementations, if there is no first shared resource, or if there is a first shared resource but no unissued task that needs to use the first shared resource, the task boundary judgment result is that the switching boundary has been reached and the task can be switched to other task queues. In this case, the output task boundary judgment result is 1.
[0075] In step S14, the scheduling result of the first task currently to be executed in the first task queue can be determined by comprehensively judging the priority comparison result, load timeout result and task boundary judgment result.
[0076] In some possible implementations, if the first task queue has a low priority (corresponding bit 0) and the task boundary judgment result indicates that the switching boundary has not been reached (corresponding bit 0), then regardless of whether the load of the high-priority task queue has timed out, the first task currently waiting to be executed in the first task queue should continue to be allowed, and the scheduling result should be determined as allowing delivery. The scheduling result is, for example, outputting a 1-bit signal, where 0 indicates that delivery is not allowed and 1 indicates that delivery is allowed. In this case, the output scheduling result is 1.
[0077] In some possible implementations, if the first task queue has a low priority (corresponding bit 0) and the task boundary judgment result indicates that the switching boundary has been reached (corresponding bit 1), then it is necessary to determine whether the load of the high-priority task queue has timed out. If the load timeout result of the target task queue indicates that the task has not timed out, it means that a new task has arrived in the target task queue within the specified time, and the task issuance of the first task queue needs to be stopped. In this case, the scheduling result is determined to be disallowed, and the output scheduling result is 0. Conversely, if the load timeout result of the target task queue indicates that the task has timed out, it means that no new task has arrived in the target task queue within the specified time. In this case, the first task currently waiting to be executed in the first task queue is allowed to proceed, the scheduling result is determined to be allowed, and the output scheduling result is 1.
[0078] In some possible implementations, if the first task queue has a high priority and the corresponding bit is 1, then the first task currently waiting to be executed in the first task queue must always be allowed to be sent down. In this case, the scheduling result is determined to be allowed to be sent down, and the output scheduling result is 1.
[0079] In some possible implementations, in step S15, if the scheduling result is "allow to send", the output scheduling result is 1, then the first task is sent to the backend module so that the backend module allocates hardware resources for the thread bundle of the first task and executes the first task, thereby completing the scheduling process of the task scheduler corresponding to the first task queue.
[0080] In this way, the task scheduler corresponding to each task queue processes the tasks in the manner described above, thus completing the entire scheduling process of the tasks in the processing device.
[0081] According to the task scheduling method of this disclosure, a priority comparison, timeout management, and task boundary judgment mechanism are introduced into the task scheduler of the processing device. Based on the priority comparison results, load timeout results, and task boundary judgment results, the method can automatically manage and switch between high and low priority tasks, thereby fully utilizing the hardware computing power of the processing device. According to the embodiments of this disclosure, fine-grained priority scheduling at the thread bundle level can be achieved at a relatively low cost, improving the efficiency of priority scheduling and avoiding deadlock problems that may occur in thread bundle scheduling schemes of related technologies, thereby improving the overall performance and processing efficiency of the processing device.
[0082] The task scheduling method according to the embodiments of this disclosure will now be described in detail.
[0083] Figure 2 This is a schematic diagram of a processing apparatus provided according to an embodiment of the present disclosure. (Refer to...) Figure 2 The processing apparatus according to an embodiment of this disclosure includes a front-end module 21, multiple task schedulers (task scheduler 1, task scheduler 2, ..., task scheduler N), and a back-end module 22, where N is an integer greater than 1. The multiple task schedulers include the aforementioned task schedulers.
[0084] In some possible implementations, the front-end module 21 is connected to multiple task schedulers and is configured to: package threads of the same processing category in the received work targets into thread bundles to obtain tasks including multiple thread bundles; and send the tasks to the corresponding task queues according to the processing category of the tasks.
[0085] In this process, after receiving the threads of the work target, the front-end module 21 packages threads of the same or different work targets and the same processing category into a thread bundle. Multiple thread bundles form a task, for example, 32 threads are packaged into a thread bundle, and 16 thread bundles form a task. The task is then sent to the task queue of the corresponding category. This disclosure does not limit the number of threads in a thread bundle or the number of thread bundles in a task.
[0086] In some possible implementations, multiple task schedulers are connected to the backend module, with each task scheduler corresponding to a task queue, such as... Figure 2 As shown, task schedulers 1-N correspond to task queues 1-N respectively. The task scheduler is configured to: receive tasks from the corresponding first task queue and obtain the task priorities of multiple task queues; determine the scheduling result of the first task to be executed in the first task queue; and, if the scheduling result allows the first task to be sent, send the first task to the backend module.
[0087] In some possible implementations, the backend module is configured to allocate hardware resources for the thread bundle of the first task and execute the thread bundle of the first task. For example... Figure 2 As shown, the backend module may include a vertex shader thread bundle resource allocator, a pixel shader thread bundle resource allocator, a compute shader thread bundle resource allocator, and a thread bundle executor.
[0088] After receiving the first task, the backend module allocates hardware resources, such as computing resources and memory resources, to the thread bundle of the first task in the vertex shader thread bundle resource allocator, pixel shader thread bundle resource allocator, or compute shader thread bundle resource allocator, according to the processing category of the first task. Then, it sends the thread bundle of the first task to the thread bundle executor for execution, thereby completing the execution process of the first task.
[0089] In this way, the scheduling of corresponding task queues can be implemented in multiple task schedulers, thereby achieving thread-beam-level, fine-grained priority scheduling and improving the efficiency of priority scheduling.
[0090] Figure 3 This is a schematic diagram of a task scheduler provided in an embodiment of this disclosure. (Refer to...) Figure 3 The task scheduler according to embodiments of the present disclosure includes a priority determination module 31, a timeout control module 32, a boundary determination module 33, and an output control module 34.
[0091] The priority judgment module 31 is connected to the timeout control module 32 and the output control module 34, and is configured to: read the first task priority of the first task queue and the second task priority of the second task queue respectively; compare the first task priority and the second task priority; and output the priority comparison result.
[0092] The timeout control module 32 is connected to the output control module 34 and is configured to: determine the load timeout result of the task in the target task queue based on the priority comparison result, wherein the target task queue is a task queue with higher priority or equal priority in the first task queue and the second task queue;
[0093] The boundary judgment module 33 is connected to the output control module 34 and is configured to: determine the task boundary judgment result of the first task queue based on the shared resource usage information of the first task queue;
[0094] The output control module 34 is configured to: determine the scheduling result of the first task currently to be executed in the first task queue based on the priority comparison result, the load timeout result and the task boundary judgment result; and, if the scheduling result allows the first task to be executed, issue the first task so that the backend module allocates hardware resources to the thread bundle of the first task and executes it.
[0095] For example, the priority determination module 31 may include at least one comparator. At the start of the current round of scheduling, the priority determination module 31 is configured to execute step S11, which reads the value of the first task priority of the first task queue and the value of the second task priority of the second task queue from the registers corresponding to the task queues; compares the value of the first task priority and the value of the second task priority, and outputs the priority comparison result.
[0096] In the example, if the priority of the first task is first priority 11, and the priority of the second task is second priority 10 or third priority 01; or if the priority of the first task is first priority 11 or second priority 10, and the priority of the second task is third priority 01, then the priority comparison result is that the priority of the first task is higher than the priority of the second task, and a 2-bit priority comparison signal 10 can be output. The first bit (bit 0) of the priority comparison signal corresponds to the first task queue, and the second bit (bit 1) corresponds to the second task queue. 1 indicates high, 0 indicates low, and 10 means that the priority of the first task is higher than the priority of the second task, the first task queue is high priority, and the second task queue is low priority.
[0097] In the example, the first bit (bit 0) of the priority comparison signal can be fixed to correspond to the task queue of the graphics category, and the second bit (bit 1) to correspond to the task queue of the computation category. In this case, the priority determination module 31 may also include a computation task switcher, a vertex task switcher, and a pixel task switcher (not shown), which are used to input the priority comparison signal of the task queue of the corresponding category, respectively. For example, if the first task queue is the computation category and the second task queue is the vertex category, the computation task switcher outputs 1, and the vertex task switcher outputs 0. This disclosure does not limit the specific hardware structure of the priority determination module 31.
[0098] In the example, if the priority of the first task is the second priority 10 or the third priority 01, and the priority of the second task is the first priority 11; or if the priority of the first task is the third priority 01, and the priority of the second task is the first priority 11 or the second priority 10, then the priority comparison result is that the priority of the first task is lower than the priority of the second task. A 2-bit priority comparison signal 01 can be output, indicating that the priority of the first task is lower than the priority of the second task, the first task queue is low priority, and the second task queue is high priority.
[0099] In the example, if both the first task priority and the second task priority are high priority 11, normal priority 10, or low priority 01, the priority comparison result is that the first task priority is lower than the second task priority. A 2-bit priority comparison signal 11 can be output, indicating that the first task priority is equal to the second task priority, and the first task queue and the second task queue are of equal priority.
[0100] When the first task queue has a high priority or equal priority, the corresponding task scheduler can prioritize issuing tasks from the first task queue to improve system processing efficiency. When the first task queue has a low priority, the corresponding task scheduler can further determine whether to issue tasks from the first task queue based on other conditions.
[0101] In some possible implementations, the priority judgment module 31 outputs the priority comparison results to the timeout control module 32 and the output control module 34 respectively for subsequent processing.
[0102] This method enables priority comparison, allowing for subsequent task assignment and control based on priority, thereby improving task scheduling efficiency.
[0103] In some possible implementations, when a task queue has a high priority or equal priority, the corresponding task scheduler can prioritize assigning tasks from that queue to improve system processing efficiency. When a task queue has a low priority, the task scheduler can further determine whether to assign tasks based on other conditions. For example, if other high-priority task queues are empty and no load is assigned within a specified time (i.e., no new tasks arrive), tasks from low-priority task queues can be assigned to more effectively utilize the hardware resources and computing power of the processing device.
[0104] In some possible implementations, the timeout control module 32 is configured to execute step S12, determining the load timeout result of tasks in the target task queue based on the priority comparison result. The target task queue is a task queue with higher or equal priority from the first and second task queues.
[0105] In some possible implementations, step S12 may include:
[0106] The target task queue is determined based on the priority comparison results;
[0107] When the target task queue includes the first task queue and the first task queue is empty, the counter of the task scheduler is used to count at a preset counting period to determine the count value of the counter;
[0108] If the count value does not reach the count threshold, the load timeout result of the first task queue is determined to be that the task has not timed out;
[0109] If the count value reaches the count threshold, the load timeout result of the first task queue is determined to be a task timeout.
[0110] For example, based on the priority comparison results, the task queues with higher or equal priority in the first and second task queues can be determined first, and these are called the target task queues.
[0111] In some possible implementations, the step of determining the target task queue based on the priority comparison result includes:
[0112] If the priority comparison result shows that the priority of the first task is higher than the priority of the second task, it is determined that the target task queue includes the first task queue;
[0113] If the priority comparison result shows that the priority of the first task is lower than the priority of the second task, it is determined that the target task queue includes the second task queue;
[0114] If the priority comparison result is that the priority of the first task is equal to the priority of the second task, then the target task queue is determined to include the first task queue and the second task queue.
[0115] In other words, if the priority of the first task is higher than that of the second task, and the priority comparison result is 10, then the first task queue is used as the target task queue; if the priority of the first task is lower than that of the second task, and the priority comparison result is 01, then the second task queue is used as the target task queue; if the priority of the first task is equal to that of the second task, and the priority comparison result is 11, then both the first and second task queues are used as target task queues.
[0116] For the task scheduler of the first task queue, if the target task queue includes the first task queue (high priority or equal priority), it can be determined whether the first task queue is empty; if the first task queue is empty, the counter in the timeout control module 32 is started, and the counter counts at a preset counting period to determine the count value.
[0117] For the current round of scheduling, if the count value does not reach the preset counting threshold, the load timeout result of the first task queue is determined to be that the task has not timed out. For example, a 1-bit signal is output, where 0 indicates no timeout and 1 indicates timeout. In this case, the output load timeout result is 0, and it is necessary to continue waiting for new tasks in the first task queue. If the count value reaches the preset counting threshold, the load timeout result of the first task queue is determined to be that the task has timed out, and the output load timeout result is 1. Tasks in the lower priority second task queue can be assigned.
[0118] The counter's counting period can be a preset clock period, and the counting threshold is a configurable value, such as 1024. This disclosure does not impose any restrictions on the specific duration of the clock period or the specific value of the counting threshold.
[0119] Figure 4 This is a schematic diagram of a timeout control module provided in an embodiment of this disclosure. (Refer to...) Figure 4 The timeout control module may include a computation task input / output controller, a vertex task input / output controller, a pixel task input / output controller, and a counter. The computation task input / output controller, vertex task input / output controller, and pixel task input / output controller are respectively used to detect whether there are new tasks to be executed in the corresponding task queues. The priority comparison results are also input to the computation task input / output controller, vertex task input / output controller, and pixel task input / output controller respectively, enabling each task input / output controller to determine the priority of its corresponding task queue.
[0120] In some possible implementations, if the first task queue (high priority or equal priority) is a task queue of the computation category, then it is determined whether the first task queue is empty in each round of scheduling; if the first task queue changes from non-empty to empty in a certain determination, then a counter is started to count at a preset counting period to determine the count value; and the count value is compared with the configured counting threshold in the counter.
[0121] In the current round of scheduling, if the count value does not reach the count threshold, the load timeout result of the first task queue is determined to be that the task has not timed out. For example, a 1-bit signal 0 is output and broadcast to each task input / output controller, indicating that it is necessary to continue waiting for new tasks from the first task queue.
[0122] In the current round of scheduling, if the count value has not reached the count threshold, the load timeout result of the first task queue is determined as a task timeout. For example, a 1-bit signal 1 is output and broadcast to each task input / output controller, indicating that a task in the lower-priority second task queue can be assigned. When the higher-priority first task queue is a task queue for calculating categories, the lower-priority second task queue is a task queue for vertex categories or pixel categories, corresponding to... Figure 4 When the vertex task input / output controller or pixel task input / output controller receives the load timeout result, it can determine that the corresponding task queue has permission to pass. After determining whether the boundary has been reached, it needs to issue low-priority tasks in the corresponding task queue according to the scheduling result.
[0123] In this way, timeout checks can be performed on high-priority task queues, so that if a high-priority task queue times out, tasks in the low-priority task queue can be dispatched, thereby making more efficient use of the hardware resources and computing power of the processing device and improving the system's processing efficiency.
[0124] In some possible implementations, the task scheduling method according to embodiments of this disclosure further includes:
[0125] When the first task queue receives the second task, a reset signal is sent to the counter to turn off the counter and reset the counter's count value.
[0126] For example, for the first task queue (high priority or equal priority), if the first task queue receives a new second task, the corresponding task input / output controller receives the task input of the corresponding category. For example, if the first task queue is a computational task queue, the computational task input / output controller receives the task input and sends a reset signal to the counter. The counter determines, based on the priority comparison result, that the reset signal is for a high-priority or equal-priority task queue, then shuts down and resets the counter value, setting it to 0. Furthermore, the load timeout result will change to "task did not time out," the output signal will be pulled low to 0, and this will be broadcast to all task input / output controllers.
[0127] Similarly, for the second task queue (low priority), a reset signal is sent to the counter when the corresponding task input / output controller receives task input. For example, if the second task queue is a vertex-type task queue, the vertex task input / output controller receives task input and sends a reset signal to the counter. Here, an OR gate is set for both the vertex and pixel task input / output controllers, indicating that either the vertex or pixel task input / output controller can send a signal to the counter. The counter determines that the reset signal belongs to a low-priority task queue based on the priority comparison result, ignores the reset signal, and continues counting. In other words, the reset mechanism only applies to high-priority tasks.
[0128] In this way, the counter can be reset when a new task arrives in the high-priority task queue, and tasks in the high-priority task queue will continue to be issued with priority, thereby improving the issuance efficiency of high-priority tasks and thus improving the effect of task scheduling.
[0129] In some possible implementations, if the task execution process requires resources shared by multiple thread bundles from different tasks, such as shared memory for communication between multiple computation groups in a computation-type task, then switching to other task queues can only proceed after all tasks corresponding to the thread bundles using these shared resources have been deployed. In this case, additional checks are required.
[0130] In some possible implementations, the boundary determination module 33 is configured to execute step S13: determining the task boundary determination result of the first task queue based on the shared resource usage information of the first task queue. Step S13 includes:
[0131] If, based on the shared resource usage information, it is determined that there is a first shared resource and there is a third task that has not been issued, the task boundary judgment result is determined to be that the switching boundary has not been reached, and the third task is a task in the first task queue that needs to use the first shared resource.
[0132] If, based on the shared resource usage information, it is determined that there is no first shared resource, or if there is a first shared resource and there is no unissued third task, the task boundary judgment result is determined to be that a switching boundary has been reached.
[0133] For example, the boundary judgment module 33 is used to determine whether there are cross-thread bundles in the lifecycle of the allocated hardware resources (such as computing resources and memory resources). If the hardware resources need to be shared by multiple thread bundles, the corresponding task queue must be able to switch to other queues for priority switching only after all thread bundles using these shared resources have been issued.
[0134] In some possible implementations, the shared resource usage information of the first task queue can be determined based on the commands in the first task queue and the resource requirement information of the thread bundles in each task. It should be understood that those skilled in the art can set the specific method for determining the shared resource usage information according to actual circumstances, and this disclosure does not impose any limitations on this.
[0135] In some possible implementations, based on shared resource usage information, if it is determined that a first shared resource exists, and there are unissued tasks in the first task queue that require the use of the first shared resource, then the task boundary judgment result is that the switching boundary has not been reached. The tasks in the first task queue must continue execution, and the system cannot switch to tasks in other task queues. In this case, if a forced switch is attempted, dependency handling anomalies or resource allocation deadlocks may occur, leading to system crashes. The task boundary judgment result can be output as a 1-bit signal, where 0 indicates the switching boundary has not been reached and 1 indicates the switching boundary has been reached. In this case, the output task boundary judgment result is 0.
[0136] In some possible implementations, based on the shared resource usage information, if there is no first shared resource, or if there is a first shared resource but no unissued task that needs to use the first shared resource, then the task boundary judgment result is that the switching boundary has been reached and the task can be switched to other task queues. In this case, the output task boundary judgment result is 1.
[0137] In this way, a task boundary judgment mechanism can be introduced to reduce or even avoid deadlock problems that may occur in the thread bundle scheduling scheme of related technologies, improve the safety of task scheduling, and improve the stability of the processing device during operation.
[0138] In each round of scheduling, the priority comparison result of the priority judgment module 31, the load timeout result of the timeout control module 32, and the task boundary judgment result of the boundary judgment module 33 are all input to the output control module 34, which performs comprehensive judgment. The output control module 34 may include multiplexers and other devices, and this disclosure does not limit the specific structure of the output control module.
[0139] In some possible implementations, the output control module 34 is configured to execute step S14, which determines the scheduling result of the first task currently to be executed in the first task queue based on the priority comparison result, the load timeout result, and the task boundary judgment result. Step S14 includes:
[0140] If the target task queue does not include the first task queue, and the task boundary judgment result is that the switching boundary has not been reached, the scheduling result is determined to allow delivery.
[0141] If the target task queue does not include the first task queue, the load timeout result of the target task queue is that the task has not timed out, and the task boundary judgment result is that the switching boundary has been reached, then the scheduling result is determined to be that the task cannot be sent.
[0142] If the target task queue does not include the first task queue, and the load timeout result of the target task queue is a task timeout, then the scheduling result is determined to allow delivery.
[0143] If the target task queue includes the first task queue, the scheduling result is determined to be allowed to be issued.
[0144] Table 1 is a schematic diagram of the scheduling results in the task scheduling method according to an embodiment of the present disclosure.
[0145] Table 1
[0146] 0 - 0 1 0 0 1 0 0 1 - 1 1 - - 1
[0147] As shown in Table 1, in the priority comparison results, "0" indicates that the task queue (first task queue) is of low priority, and "1" indicates that the task queue is of high priority or equal priority; in the load timeout results, "0" indicates that the high priority or equal priority task queue has not timed out, and "1" indicates that the high priority or equal priority task queue has timed out; in the task boundary judgment results, "0" indicates that the task queue has not reached the switching boundary, "1" indicates that the task queue has reached the switching boundary, and "-" indicates either "0" or "1"; in the scheduling results, "1" indicates that the task can be sent, and "0" indicates that the task cannot be sent.
[0148] In some possible implementations, if the first task queue is of low priority (corresponding bit 0) and the task boundary judgment result is that the switching boundary has not been reached (corresponding bit 0), then regardless of whether the load of the high-priority task queue has timed out, the first task currently waiting to be executed in the first task queue should continue to be allowed to proceed, the scheduling result should be determined as allowed to proceed, and the output scheduling result should be 1.
[0149] In some possible implementations, if the first task queue is of low priority (corresponding bit 0) and the task boundary judgment result is that the switching boundary has been reached (corresponding bit 1), then it is necessary to determine whether the load of the high-priority task queue has timed out. If the load timeout result of the target task queue is that the task has not timed out (corresponding bit 0), it means that a new task has arrived in the target task queue within the specified time, and the task issuance of the first task queue needs to be stopped. In this case, the scheduling result is determined to be that issuance is not allowed, and the output scheduling result is 0.
[0150] In some possible implementations, if the first task queue is of low priority, the corresponding bit is 0; the load timeout result of the target task queue is task timeout, the corresponding bit is 1, indicating that no new task has arrived in the target task queue within the specified time. Then, regardless of whether the task boundary judgment result is whether the switching boundary has been reached, the first task currently waiting to be executed in the first task queue is allowed to proceed, the scheduling result is determined to be allowed to be issued, and the output scheduling result is 1.
[0151] In some possible implementations, if the first task queue is of high priority and then equal priority, and the corresponding bit is 1, then the first task to be executed in the first task queue must always be allowed to be sent down. In this case, the scheduling result is determined to be allowed to be sent down, and the output scheduling result is 1.
[0152] In this way, the logic for determining whether a task currently pending execution in the task queue is allowed to be issued can be implemented at a relatively low cost, enabling fine-grained priority scheduling and improving the efficiency of priority scheduling.
[0153] In some possible implementations, the output control module 34 is also configured to execute step S15, which, if the scheduling result is that the task can be sent, sends the first task so that the backend module allocates hardware resources for the thread bundle of the first task and executes it, thereby completing the scheduling process of the task scheduler in the current round corresponding to the first task queue.
[0154] In this way, the task scheduler corresponding to each task queue processes the task in the manner described above, thus completing the scheduling process for the current round of each task queue in the processing device, and the next round of scheduling can then begin.
[0155] According to the task scheduling method of this disclosure, a priority comparison, timeout management, and task boundary judgment mechanism are introduced into the task scheduler of the processing device. Based on the priority comparison results, load timeout results, and task boundary judgment results, the method can automatically manage and switch between high and low priority tasks, thereby fully utilizing the hardware computing power of the processing device. According to the embodiments of this disclosure, fine-grained priority scheduling at the thread bundle level can be achieved at a relatively low cost, improving the efficiency of priority scheduling and avoiding deadlock problems that may occur in thread bundle scheduling schemes of related technologies, thereby improving the overall performance and processing efficiency of the processing device.
[0156] It is understood that the various methods and apparatus embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the underlying principles and logic. Due to space limitations, these will not be elaborated upon further. Those skilled in the art will understand that the specific execution order of each step in the above methods of specific implementation should be determined by its function and possible internal logic.
[0157] In addition, this disclosure also provides electronic devices and computer-readable storage media, all of which can be used to implement any of the task scheduling methods provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section on methods and will not be repeated here.
[0158] Figure 5 This is a block diagram of an electronic device provided in an embodiment of the present disclosure.
[0159] Reference Figure 5 This disclosure provides an electronic device, which includes: at least one processor 501; at least one memory 502; and one or more I / O interfaces 503 connected between the processor 501 and the memory 502; wherein the memory 502 stores one or more computer programs that can be executed by the at least one processor 501, and the one or more computer programs are executed by the at least one processor 501 to enable the at least one processor 501 to perform the above-described task scheduling method.
[0160] This disclosure also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor / processor core, implements the task scheduling method described above. The computer-readable storage medium may be volatile or non-volatile.
[0161] This disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, wherein when the computer-readable code is run in a processor of an electronic device, the processor in the electronic device executes the above-described task scheduling method.
[0162] Those skilled in the art will understand that all or some of the steps, systems, and apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software can be distributed on a computer-readable storage medium, which may include computer storage media (or non-transitory media) and communication media (or transient media).
[0163] As is known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable program instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), static random access memory (SRAM), flash memory or other memory technologies, portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, it is known to those skilled in the art that communication media typically contain computer-readable program instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.
[0164] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.
[0165] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0166] The computer program product described herein can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0167] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0168] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0169] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0171] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.
Claims
1. A task scheduling method, characterized in that, A task scheduler applied in a processing device, the task scheduler corresponding to a first task queue, the tasks in the first task queue having the same processing category, the processing category including computation category and graphics category, and each task including multiple thread bundles; The method includes: The priority of the first task in the first task queue is compared with the priority of the second task in the second task queue to determine the priority comparison result. The processing categories of the tasks in the first task queue and the second task queue are different. Based on the priority comparison result, the load timeout result of the target task queue is determined, wherein the target task queue is the higher priority or equal priority task queue among the first task queue and the second task queue; Based on the shared resource usage information of the first task queue, determine the task boundary judgment result of the first task queue; Based on the priority comparison result, the load timeout result, and the task boundary judgment result, the scheduling result of the first task currently to be executed in the first task queue is determined; If the scheduling result allows the task to be sent, the first task is sent to the backend module so that the backend module allocates hardware resources for the thread bundle of the first task and executes it.
2. The method according to claim 1, characterized in that, Determining the load timeout result of the target task queue based on the priority comparison result includes: The target task queue is determined based on the priority comparison results; When the target task queue includes the first task queue and the first task queue is empty, the counter of the task scheduler is used to count at a preset counting period to determine the count value of the counter; If the count value does not reach the count threshold, the load timeout result of the first task queue is determined to be that the task has not timed out; If the count value reaches the count threshold, the load timeout result of the first task queue is determined to be a task timeout.
3. The method according to claim 2, characterized in that, The step of determining the target task queue based on the priority comparison result includes: If the priority comparison result shows that the priority of the first task is higher than the priority of the second task, it is determined that the target task queue includes the first task queue. If the priority comparison result shows that the priority of the first task is lower than the priority of the second task, it is determined that the target task queue includes the second task queue; If the priority comparison result is that the priority of the first task is equal to the priority of the second task, then the target task queue is determined to include the first task queue and the second task queue.
4. The method according to claim 2, characterized in that, The method further includes: When the first task queue receives the second task, a reset signal is sent to the counter to turn off the counter and reset the counter's count value.
5. The method according to claim 1, characterized in that, The step of determining the task boundary judgment result of the first task queue based on the shared resource usage information of the first task queue includes: If, based on the shared resource usage information, it is determined that there is a first shared resource and there is a third task that has not been issued, the task boundary judgment result is determined to be that the switching boundary has not been reached, and the third task is a task in the first task queue that needs to use the first shared resource. If, based on the shared resource usage information, it is determined that there is no first shared resource, or if there is a first shared resource and there is no unissued third task, the task boundary judgment result is determined to be that a switching boundary has been reached.
6. The method according to claim 1, characterized in that, The step of determining the scheduling result of the first task currently to be executed in the first task queue based on the priority comparison result, the load timeout result, and the task boundary judgment result includes: If the target task queue does not include the first task queue, and the task boundary judgment result is that the switching boundary has not been reached, the scheduling result is determined to allow delivery. If the target task queue does not include the first task queue, the load timeout result of the target task queue is that the task has not timed out, and the task boundary judgment result is that the switching boundary has been reached, then the scheduling result is determined to be that the task cannot be sent. If the target task queue does not include the first task queue, and the load timeout result of the target task queue is a task timeout, then the scheduling result is determined to allow delivery. If the target task queue includes the first task queue, the scheduling result is determined to be allowed to be issued.
7. The method according to claim 1, characterized in that, The processing device includes multiple task schedulers, each corresponding to a task queue. The graphics categories include vertex categories and pixel categories, and the task priority of the task queue for the vertex category is the same as the task priority of the task queue for the pixel category. If the processing category of a task in the first task queue is computation, then the processing category of a task in the second task queue is either vertex or pixel.
8. A task scheduler, characterized in that, The task scheduler corresponds to a first task queue, in which tasks have the same processing category, including computation and graphics categories, and each task includes multiple thread bundles. The task scheduler includes a priority judgment module, a timeout control module, a boundary judgment module, and an output control module. The priority determination module is connected to the timeout control module and the output control module, and is configured to: read the first task priority of the first task queue and the second task priority of the second task queue respectively; Compare the priorities of the first task and the second task, and output the priority comparison result; The timeout control module is connected to the output control module and is configured to: determine the load timeout result of the task in the target task queue based on the priority comparison result, wherein the target task queue is a higher priority or equal priority task queue among the first task queue and the second task queue; The boundary judgment module is connected to the output control module and is configured to: determine the task boundary judgment result of the first task queue based on the shared resource usage information of the first task queue; The output control module is configured to: determine the scheduling result of the first task currently to be executed in the first task queue based on the priority comparison result, the load timeout result, and the task boundary judgment result; and, if the scheduling result allows the first task to be executed, issue the first task so that the backend module allocates hardware resources to the thread bundle of the first task and executes it.
9. A processing apparatus, characterized in that, The device includes a front-end module, multiple task schedulers, and a back-end module, wherein the task scheduler includes the task scheduler according to claim 8, wherein... The front-end module is connected to multiple task schedulers and is configured to: package threads of the same processing category in the received work targets into thread bundles to obtain tasks including multiple thread bundles; and send the tasks to the corresponding task queues according to the processing category of the tasks. Multiple task schedulers are connected to the backend module, and each task scheduler corresponds to a task queue; The task scheduler is configured to: receive tasks from the corresponding first task queue and obtain the task priorities of multiple task queues; determine the scheduling result of the first task currently to be executed in the first task queue; and, if the scheduling result allows the task to be sent, send the first task to the backend module. The backend module is configured to allocate hardware resources for the thread bundle of the first task and execute the thread bundle of the first task.
10. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs that can be executed by the at least one processor, the one or more computer programs being executed by the at least one processor to enable the at least one processor to perform the task scheduling method as described in any one of claims 1-7.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the task scheduling method as described in any one of claims 1-7.
Citation Information
Patent Citations
Batch task execution method and device and electronic equipment
CN117519930A
Task scheduling method and device, computer equipment, storage medium and program product
CN119621330A