Two-stage scheduling method and device for thread blocks and medium

Through a two-level scheduling method of thread blocks, the limitations of existing GPGPU scheduling algorithm are solved through the determination of thread groups' grouping and scheduling priority, combined with the greedy scheduling algorithm, and more efficient resource utilization and task execution are achieved.

CN119938273APending Publication Date: 2025-05-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510010668.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing graphics processing unit GPGPU scheduling algorithms, such as polling and greedy algorithms, have limitations, which may cause thread bundles to encounter long delay operations at the same time or damage data locality, reducing cache hit rate.

Method used

A two-level scheduling method for thread blocks is proposed. By obtaining the total thread block in the graphics processor, dividing it into thread groups according to the preset grouping strategy, and determining its scheduling priority based on the thread group ID, and scheduling algorithms are used for scheduling.

Benefits of technology

By scheduling thread groups, thread blocks can be evenly allocated to the kernel, avoid kernel overload or resource idleness, ensure full utilization of computing resources, and effectively avoid unnecessary waiting and accelerate task execution when facing bottlenecks such as memory bandwidth limitations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938273A_ABST
    Figure CN119938273A_ABST
Patent Text Reader

Abstract

The invention discloses a two-stage scheduling method and device for thread blocks and a medium, and relates to the technical field of computers.The method comprises the steps that a started main thread block in a graphics processor is obtained, the main thread block is divided into a plurality of thread groups according to a preset grouping strategy, and each thread group comprises the thread blocks; determining scheduling priorities of the plurality of thread groups according to the IDs corresponding to the thread groups, and performing cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities; and determining the thread blocks contained in the thread groups corresponding to the execution thread scheduling, obtaining a plurality of thread bundles contained in the thread blocks, and scheduling the plurality of thread bundles through a greedy scheduling algorithm. Through a greedy scheduling algorithm, idle time, load imbalance and resource waste are reduced as much as possible in the execution process of the graphics processor, so that the computing throughput and the system performance are maximized, and the parallel execution efficiency is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a two-level scheduling method, device and medium for thread blocks. Background Art

[0002] With the rapid development of integrated circuit technology, the computing power of general-purpose graphics processing units (GPGPU) has been continuously improved, and it has now become a powerful high-performance computing platform. As a streaming multiprocessor, the graphics processing unit GPGPU has thousands of processors and a small number of control units, and has been widely used in the field of highly parallel general computing.

[0003] The scheduling algorithm of the graphics processing unit GPGPU has a great impact on the performance of the GPGPU. The purpose of the thread warp scheduling strategy is to select the appropriate thread warp to launch and execute. Typical scheduling algorithms, such as Round-Robin (RR) and Greedy, have their limitations. The round-robin algorithm may cause all thread warps to encounter long-latency operations at the same time without additional thread warps to cover up these delays. On the other hand, although the greedy algorithm can better hide long delays, it may damage data locality, reduce cache hit rate, and increase off-chip memory access. Summary of the invention

[0004] In order to solve the above problems, the present application proposes a two-level scheduling method based on a thread block, comprising: obtaining a total thread block started in a graphics processor, dividing the total thread block into a plurality of thread groups according to a preset grouping strategy, wherein the thread group includes a plurality of thread blocks;

[0005] Determine scheduling priorities of the plurality of thread groups according to IDs corresponding to the thread groups, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities;

[0006] Determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a number of thread warps included in the thread block, and schedule the number of thread warps through a greedy scheduling algorithm.

[0007] In one implementation of the present application, the cores included in the graphics processor are obtained to determine the maximum number of schedulable thread bundles corresponding to each core; the schedulable number of thread bundles of the graphics processor is determined based on the number of cores and the maximum number of schedulable thread bundles corresponding to each core; and the thread bundles based on the schedulable number of thread bundles are combined into a plurality of thread blocks.

[0008] In one implementation of the present application, the pipeline stage corresponding to the image processor is obtained; the thread group is initialized according to the initialization rule to obtain the initial number of thread groups; a determination rule is constructed according to the total number of thread blocks, the schedulable number of thread bundles and the pipeline stage, and the initial number of thread groups is judged based on the determination rule.

[0009] In one implementation of the present application, when the initial number of thread groups meets the determination rule, the total thread blocks are divided based on the initial number of thread groups; when the initial number of thread groups does not meet the determination rule, the initial number of thread groups is decremented based on a preset decrement amount; and the decremented number of thread groups is iteratively determined based on the determination rule until the number of thread groups converges.

[0010] In one implementation of the present application, the preset number of thread groups for determination is obtained, and the number of group thread blocks contained in each group after grouping is determined based on the total number of thread blocks and the preset number of thread groups; the number of group thread warps contained in each group after grouping is determined based on the schedulable number of thread warps, and it is determined whether the number of group thread warps in each group is not less than the length of the pipeline stage.

[0011] In one implementation of the present application, the thread groups are traversed in a loop according to the scheduling priority, the optimal thread group for priority execution scheduling is determined, and the thread bundles contained in the optimal thread group are scheduled; when it is determined that the optimal thread group is delayed in execution, the scheduling priority corresponding to the optimal thread group is reduced, and the thread group corresponding to the next highest scheduling priority is scheduled.

[0012] In one implementation of the present application, a greedy scheduling algorithm is recursively called; the optimal thread bundle in the optimal thread group is determined by the greedy scheduling algorithm, and the optimal thread bundle is scheduled; when the execution delay of the optimal thread bundle is determined, the suboptimal thread bundle in the optimal thread group is determined by the greedy scheduling algorithm, and the suboptimal thread bundle is scheduled.

[0013] In an implementation of the present application, the number of unscheduled thread warps in the optimal thread group is obtained; when the number of the unscheduled thread warps is 0, the execution delay of the optimal thread group is determined.

[0014] The two-level scheduling method of thread blocks proposed in this application can bring the following beneficial effects:

[0015] By scheduling thread groups, thread blocks can be evenly distributed to all cores according to the core load, avoiding overload or idle resources of some cores and ensuring that the computing resources of the graphics processor are fully utilized.

[0016] In addition, by scheduling within the thread group, the scheduling order of the thread bundles can be dynamically adjusted according to the current execution status. Especially when facing bottlenecks such as memory bandwidth limitations and data dependencies, it can effectively avoid unnecessary waiting and thus speed up the execution of tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A schematic diagram of a flow chart of a two-level scheduling method for thread blocks in an embodiment of the present application;

[0019] Figure 2 This is a schematic diagram of a model of a thread group in an embodiment of the present application;

[0020] Figure 3 This is a schematic diagram of the process of presetting the grouping strategy in the embodiment of the present application;

[0021] Figure 4 This is a flow chart of inter-group scheduling in an embodiment of the present application;

[0022] Figure 5 A schematic diagram of the process of intra-group scheduling in an embodiment of the present application;

[0023] Figure 6 This is a schematic diagram of a two-level scheduling device for thread blocks in an embodiment of the present application. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0025] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0026] like Figure 1 As shown, the embodiment of the present application provides a two-level scheduling method for thread blocks, including:

[0027] S101, obtaining a total number of thread blocks started in a graphics processor, and dividing the total number of thread blocks into a plurality of thread groups according to a preset grouping strategy, wherein the thread groups include a plurality of thread blocks.

[0028] Specifically, all total thread blocks started in the current graphics processor are counted, and according to thread scheduling requirements and resource constraints, the total thread blocks are divided into several thread groups through a preset grouping strategy, and each thread group consists of multiple thread blocks.

[0029] It should be noted that the factors affecting the resource requirements of each thread block include registers, shared memory, execution units, computing density or load, etc. According to the scheduling strategy, the thread blocks in the thread group can be executed together on certain specific hardware resources (such as shared memory, cache, etc.), which helps to reduce data access latency across thread blocks and improve performance. In this application specification, the way to dynamically make scheduling decisions based on the needs, dependencies, execution conditions and status of hardware resources of the thread blocks is achieved through thread block perception.

[0030] Before obtaining the total thread blocks started in the graphics processor, the cores included in the graphics processor are obtained, the maximum number of schedulable thread warps corresponding to each core is determined, and the schedulable number of thread warps of the graphics processor is determined according to the number of cores and the maximum number of schedulable thread warps corresponding to each core. The thread warps based on the schedulable number of thread warps are combined into a number of thread blocks. Figure 2 As shown, it is a thread block grouping model diagram, where Group is a thread group, the thread group contains several thread blocks (CTA), each thread block contains several thread warps (Warps), and each thread warp contains several instructions (Inst).

[0031] It should be noted that a graphics processor contains multiple cores, and one core can execute multiple thread bundles simultaneously. However, each core has a certain amount of resources (such as registers, shared memory, etc.), which limits the maximum number of thread bundles it can schedule. Thread scheduling is performed based on the maximum number of thread bundles, which can maximize the core computing resources.

[0032] Furthermore, the pipeline stage corresponding to the image processor is obtained, the thread group is initialized according to the initialization rule to obtain the initial number of thread groups, and a determination rule is constructed according to the total number of thread blocks, the number of schedulable thread warps and the pipeline stage, and the initial number of thread groups is determined based on the determination rule.

[0033] Among them, the process of formulating the judgment rules includes: obtaining the preset number of thread groups for judgment, determining the number of group thread blocks contained in each group after grouping according to the total number of thread blocks and the preset number of thread groups, determining the number of group thread warps contained in each group after grouping according to the schedulable number of thread warps, and judging whether the number of group thread warps is not less than the length of the pipeline stage.

[0034] Furthermore, when the initial number of thread groups meets the determination rules, the total thread blocks are divided based on the initial number of thread groups; when the initial number of thread groups does not meet the determination rules, the initial number of thread groups is decremented based on a preset decrement amount; and the decremented number of thread groups is iteratively determined based on the determination rules until the number of thread groups converges.

[0035] In the embodiment of the present application, the initial number of thread groups is consistent with the initial number of total thread blocks, and the preset decrement amount is 1. After the division based on the number of thread groups, each thread group must be judged. When the initial number of thread groups does not meet the judgment rule, the initial number of thread groups is reduced by 1, and the judgment is made again until the maximum number of thread groups that meet the judgment rule is obtained.

[0036] For example, Figure 3 As shown, assuming that the total number of thread blocks started in the GPU is N, and the GPU execution pipeline stage has p-stage pipelines, if all thread blocks are divided into M groups, the i-th group has n i thread blocks, where n i =N / M, which is equivalent to n i ×t thread bundles, when grouping, it is necessary to ensure that the number of thread bundles in the group is not less than the length of the pipeline to keep the pipeline busy, that is, p <n×t。

[0037] S102: Determine scheduling priorities of the plurality of thread groups according to IDs corresponding to the thread groups, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities.

[0038] Specifically, according to the order of the thread groups, corresponding IDs are added to the thread groups, and according to the IDs corresponding to the thread groups, the scheduling priorities of the thread groups are determined. In the embodiment of the present application, the scheduling priority of group 0 is initially the highest.

[0039] Furthermore, according to the scheduling priority, the thread groups are traversed in a loop to determine the optimal thread group for priority execution scheduling, and the thread warps contained in the optimal thread group are scheduled. When it is determined that all thread warps in the optimal thread group have executed long-delay instructions, the scheduling priority corresponding to the optimal thread group is reduced, and the thread group corresponding to the next highest scheduling priority is scheduled.

[0040] For example, Figure 4 As shown, assuming that the currently scheduled thread group is the i-th group Group[j], when all the thread warps in the group execute the long-delay instruction, the priority of the group is reduced to the lowest, and the next group Group[j+1] is given the highest priority to perform the next scheduling.

[0041] It should be noted that by reasonably allocating the execution order of thread groups, it is possible to ensure that hardware resources (such as computing units, memory, etc.) are fully utilized, thread groups with higher priorities are executed first, and thread groups with lower priorities can also get execution opportunities at reasonable times, thereby avoiding long-term blocking of certain thread groups and optimizing resource usage. By scheduling thread groups according to priority and resource requirements, the overall computing throughput can be maximized. Prioritizing the scheduling of computationally intensive tasks can ensure that these tasks are completed as early as possible, while subsequent tasks can continue to be executed when resources are idle, thereby improving the parallelism of task execution.

[0042] S103: Determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a number of thread warps included in the thread block, and schedule the number of thread warps by using a greedy scheduling algorithm.

[0043] Specifically, all thread blocks included in the thread group are obtained, and the execution order of the thread blocks is determined, the thread warps included in the thread blocks are obtained, and based on the execution order of the thread blocks, the thread warps included in the thread blocks are further scheduled.

[0044] Furthermore, the greedy scheduling algorithm is recursively called to determine the optimal thread bundle in the optimal thread group through the greedy scheduling algorithm, and the optimal thread bundle is scheduled. When the execution delay of the optimal thread bundle is determined, the suboptimal thread bundle in the optimal thread group is determined through the greedy scheduling algorithm, and the suboptimal thread bundle is scheduled.

[0045] It should be noted that the core idea of ​​the greedy scheduling algorithm is to select the current optimal decision each time, that is, to select the most favorable thread bundle for execution at each step in order to achieve global optimization. Through the greedy scheduling algorithm, the GPU can minimize idle time, load imbalance and resource waste during execution, thereby maximizing computing throughput and system performance and optimizing parallel execution efficiency.

[0046] Further, the number of unscheduled thread warps in the optimal thread group is obtained, and when the number of unscheduled thread warps is 0, the execution delay of the optimal thread group is determined.

[0047] For example, Figure 5 As shown in the figure, assuming that Warp[k] is currently scheduled, the scheduler will always schedule Warp[k] before Warp[k] encounters a long delay instruction. When Warp[k] is blocked and enters an inactive state due to encountering a long instruction, in order to hide the delay, Warp[k+1] is scheduled, and so on, until all Warps in the group are blocked due to long instructions, and the priority of the group will be reduced to the lowest.

[0048] For example, suppose the total number of thread blocks started in the graphics processor is N = 7, the number of pipeline stages is p = 3, and the number of thread warps in each thread block is t = 2. Set the number of thread blocks in the first thread group to 2, then the first thread group contains 2 × 2 = 4 thread warps, p = 3 < 4, which meets the judgment rule, and the size of the second group is similar, set to n = 2. In order to meet the minimum requirement of pipeline length, we set the size of the third group to n = 3, which contains the remaining thread blocks. We group threads while considering the locality of thread blocks, and adopt a two-level scheduling strategy based on the greedy algorithm. According to the grouping situation, thread group [0] is scheduled for execution first. Thread group [0] is scheduled according to the greedy scheduling algorithm. First, the thread bundle [0] instruction is scheduled. After issuing two short instructions Inst1 and Inst2, a long instruction Inst3 is encountered. In order to hide the delay, the instructions in thread bundle 1 are scheduled again. After two short delay instructions occur, a long delay instruction is also encountered. At this time, all thread bundles in thread group [0] are blocked due to the long instruction delay. Therefore, the thread bundle instructions in thread group [1] are scheduled, and thread bundles 2 and 3 are scheduled for execution in the same way.

[0049] By scheduling thread groups, thread blocks can be evenly distributed to all cores according to the core load, avoiding overload or idle resources of some cores and ensuring that the computing resources of the graphics processor are fully utilized.

[0050] In addition, by scheduling within the thread group, the scheduling order of the thread bundles can be dynamically adjusted according to the current execution status. Especially when facing bottlenecks such as memory bandwidth limitations and data dependencies, it can effectively avoid unnecessary waiting and thus speed up the execution of tasks.

[0051] like Figure 6 As shown, a two-level scheduling device for thread blocks is characterized by comprising:

[0052] at least one processor; and,

[0053] a memory communicatively connected to the at least one processor; wherein,

[0054] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the following operations: obtaining a total thread block started in a graphics processor, dividing the total thread block into a plurality of thread groups according to a preset grouping strategy, wherein the thread group includes a plurality of thread blocks; determining a scheduling priority of the plurality of thread groups according to an ID corresponding to the thread group, and performing cyclic thread scheduling on the plurality of thread groups based on the scheduling priority; determining a thread block included in a thread group corresponding to an execution thread scheduling, obtaining a plurality of thread bundles included in the thread block, and scheduling the plurality of thread bundles through a greedy scheduling algorithm.

[0055] The embodiment of the present application also provides a non-volatile computer storage medium, which stores computer executable instructions and is applied to a server, wherein the computer executable instructions are configured to: obtain a total thread block started in a graphics processor, divide the total thread block into a plurality of thread groups according to a preset grouping strategy, wherein the thread group includes a plurality of thread blocks; determine a scheduling priority of the plurality of thread groups according to an ID corresponding to the thread group, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priority; determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a plurality of thread bundles included in the thread block, and schedule the plurality of thread bundles through a greedy scheduling algorithm.

[0056] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0057] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0058] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0059] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0060] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0061] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0062] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0063] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0064] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0065] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0066] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A two-level scheduling method for thread blocks, characterized in that: include: Obtaining a total number of thread blocks started in a graphics processor, and dividing the total number of thread blocks into a plurality of thread groups according to a preset grouping strategy, wherein the thread groups include a plurality of thread blocks; Determine scheduling priorities of the plurality of thread groups according to IDs corresponding to the thread groups, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities; Determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a number of thread warps included in the thread block, and schedule the number of thread warps through a greedy scheduling algorithm.

2. A two-level scheduling method for thread blocks according to claim 1, characterized in that: Before obtaining the total thread blocks started in the graphics processor, the method further includes: Obtain the cores included in the graphics processor and determine the maximum number of schedulable thread warps corresponding to each core; Determining the schedulable number of thread warps of the graphics processor according to the number of the cores and the maximum number of schedulable thread warps corresponding to each core; A number of thread blocks are formed based on the schedulable number of thread warps.

3. The two-level scheduling method for thread blocks according to claim 2, characterized in that: The dividing the total thread block into a plurality of thread groups according to a preset grouping strategy specifically includes: Obtaining the pipeline stage corresponding to the image processor; Initialize the thread group according to the initialization rule to obtain the initial number of thread groups; A determination rule is constructed according to the total number of thread blocks, the schedulable number of thread warps and the pipeline stage, and the initial number of the thread groups is determined based on the determination rule.

4. The two-level scheduling method for thread blocks according to claim 3, characterized in that: After determining the initial number of the thread groups based on the determination rule, the method further includes: When the initial number of thread groups meets the determination rule, the total thread blocks are divided based on the initial number of thread groups; When the initial number of thread groups does not meet the determination rule, the initial number of thread groups is decremented based on a preset decrement amount; The number of thread groups after decreasing is iteratively determined based on the determination rule until the number of thread groups converges.

5. The two-level scheduling method for thread blocks according to claim 3, characterized in that: The constructing of a determination rule according to the total number of thread blocks, the schedulable number of thread warps and the pipeline stage specifically includes: Obtaining a preset number of thread groups for determination, and determining the number of group thread blocks contained in each group after grouping according to the total number of thread blocks and the preset number of thread groups; According to the schedulable number of thread warps, the number of group thread warps contained in each group after grouping is determined, and it is judged whether the number of group thread warps in each group is not less than the length of the pipeline stage.

6. The two-level scheduling method for thread blocks according to claim 1, characterized in that: The performing cyclic thread scheduling on the plurality of thread groups based on the scheduling priority specifically includes: According to the scheduling priority, the thread groups are traversed in a loop to determine an optimal thread group for priority execution scheduling, and the thread warps included in the optimal thread group are scheduled; When the execution delay of the optimal thread group is determined, the scheduling priority corresponding to the optimal thread group is reduced, and the thread group corresponding to the next highest scheduling priority is scheduled.

7. A two-level scheduling method for thread blocks according to claim 6, characterized in that: The step of scheduling the plurality of thread warps by using a greedy scheduling algorithm specifically includes: Recursively call the greedy scheduling algorithm; Determine the optimal thread warp in the optimal thread group by using the greedy scheduling algorithm, and schedule the optimal thread warp; When the execution delay of the optimal thread warp is determined, a suboptimal thread warp in the optimal thread group is determined by the greedy scheduling algorithm, and the suboptimal thread warp is scheduled.

8. The two-level scheduling method for thread blocks according to claim 7, characterized in that: After scheduling the plurality of thread warps by using the greedy scheduling algorithm, the method further includes: Obtaining the number of unscheduled thread warps in the optimal thread group; When the number of the unscheduled thread warps is 0, the optimal thread group execution delay is determined.

9. A two-level scheduling device for thread blocks, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform, for example: Obtaining a total number of thread blocks started in a graphics processor, and dividing the total number of thread blocks into a plurality of thread groups according to a preset grouping strategy, wherein the thread groups include a plurality of thread blocks; Determine scheduling priorities of the plurality of thread groups according to IDs corresponding to the thread groups, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities; Determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a number of thread warps included in the thread block, and schedule the number of thread warps through a greedy scheduling algorithm.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured to: Obtaining a total number of thread blocks started in a graphics processor, and dividing the total number of thread blocks into a plurality of thread groups according to a preset grouping strategy, wherein the thread groups include a plurality of thread blocks; Determine scheduling priorities of the plurality of thread groups according to IDs corresponding to the thread groups, and perform cyclic thread scheduling on the plurality of thread groups based on the scheduling priorities; Determine a thread block included in a thread group corresponding to the execution thread scheduling, obtain a number of thread warps included in the thread block, and schedule the number of thread warps through a greedy scheduling algorithm.

Citation Information

Cited By

  • Compiler-based thread bundle scheduling method, electronic equipment and storage medium

    CN122173143A