Thread bundle scheduling method and system for optimizing long delay operation
By dividing thread bundles into two categories: ready and pending, and grouping them according to long delay and short delay operations, long delay operations are given priority, and short delay operations are used to hide delays, the pipeline stagnation caused by long delay operations in GPGPU is solved, and performance is improved.
Patent Information
- Application Number
- CN202510171078.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-05-30
AI Technical Summary
The medium-long delay operation of the GPGPU of the general graphics processor causes pipeline stagnation and performance to degrade. The existing scheduling algorithms have failed to effectively solve the problem of pipeline stagnation.
By dividing the thread bundle into two categories: ready and pending, and grouping according to long delay and short delay operations, following the principle of priority of long delay operations, using short delay operations to hide the delay, and determining whether to send instructions to the execution unit based on the scoreboard signal.
It reduces pipeline stagnation, improves the ability of GPGPU to hide delays, improves thread scheduling efficiency and computing resource utilization, thereby enhancing the performance of GPGPU.
Smart Images

Figure CN120066585A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of general-purpose graphics processing units, and particularly to a warp scheduling method and system for optimizing long-latency operations. Background Art
[0002] The modern general-purpose graphics processing unit (GPGPU) architecture is based on the single-instruction multiple-thread (SIMT) computing model, where multiple threads form a warp. A warp is the smallest unit that can be scheduled and executed in parallel in a GPGPU, enabling all threads in the same warp to execute the same instruction simultaneously and process different data.
[0003] There are two main types of instruction operations in a GPGPU, namely long-latency instruction operations and short-latency instruction operations. Long-latency operations mainly refer to memory access operations. Short-latency operations are mainly some arithmetic and logical operations. However, long-latency instruction operations can lead to underutilization of on-chip computing resources and also cause pipeline stalls, resulting in a decline in the performance of the GPGPU.
[0004] Generally, short-latency operations can be used to hide the latency of long-latency operations and prevent pipeline stalls. However, when there are few short-latency operations, it is difficult to fully hide the latency. Also, since a GPGPU lacks structures such as complex pipelines and branch predictors that a CPU has, it relies on quickly switching warps to mask pipeline stalls. But this method of hiding long delays through quick switching is relatively inefficient, not only affecting the locality of memory access but also having a significant impact on the utilization rate of GPGPU computing resources, thereby significantly restricting the overall performance of the GPGPU.
[0005] Traditional LRR (Loose Round Robin) and GTO (Greedy Then Oldest) scheduling algorithms can effectively utilize data locality within and between warps, improve cache hit rates, reduce cache interference, and reduce off-chip memory access. However, they mainly focus on reducing the number of long-latency operations and do not directly address the pipeline stall problem caused by long-latency operations.
[0006] The present invention proposes a warp scheduling method and system for optimizing long-latency operations, aiming to reduce pipeline stalls, improve the ability of a GPGPU to hide latency, enhance thread scheduling efficiency and computing resource utilization rate, and thus enhance the performance of the GPGPU. Summary of the Invention
[0007] In order to make up for the defects of the prior art, the present invention provides a simple and efficient thread warp scheduling method and system for optimizing long-delay operations.
[0008] The present invention is achieved through the following technical solutions:
[0009] A thread warp scheduling method for optimizing long-delay operations, characterized in that it includes the following steps:
[0010] Step S1, dividing the thread warps into ready thread warps and pending thread warps based on the synchronization fence and the instruction buffer status flag full signal state, and distinguishing them by the flag bit S;
[0011] Step S2: group the ready thread warps according to the long-delay operation and the short-delay operation, and arrange the priorities according to the principle of giving priority to the long-delay operation and the principle of arranging the thread warps in the same group from small to large according to the thread warp ID;
[0012] Step S3: Send the thread warps to the execution unit in sequence, use short-delay operations to hide the delay, and decide whether to send instructions to the execution unit according to the scoreboard valid signal scoreboard_busy to ensure efficient use of computing resources.
[0013] In step S1, the thread warps are classified according to the synchronization barrier status of the thread warps and whether the status flag full signal from the instruction buffer is 0, as follows:
[0014] If the synchronization fence signal of the thread warp is 1 or the status flag full signal of the instruction buffer of the thread warp is 0, the thread warp is a pending thread warp; otherwise, it is a ready thread warp;
[0015] The flag bit S is used to identify whether the current thread warp is a ready thread warp. The flag bit S of a ready thread warp is 1, and the flag bit S of a pending thread warp is 0.
[0016] When the synchronization fence state of the thread warp and the status flag full signal from the instruction buffer change, the flag bit of the corresponding thread warp is updated synchronously.
[0017] In the step S2, the ready thread warps are grouped according to the instruction judgment signal long_short_judge from the instruction buffer module corresponding to the thread warp:
[0018] If the instruction judgment signal long_short_judge is at a high level, it is determined that the corresponding thread warp executes the long delay operation, and the thread warp is added to the long delay operation thread warp group;
[0019] If the instruction judgment signal long_short_judge is at a low level, it is determined that the corresponding warp executes a short delay operation, and the warp is placed in the short delay operation warp group.
[0020] In step S3, when the valid signal scoreboard_busy signal given by the scoreboard is 0, the corresponding warp and its instructions are sent to the execution unit; otherwise, the next ready warp is directly switched to emit the warp that meets this condition.
[0021] A warp scheduling system for optimizing long delay operations includes a decoding module, an instruction buffer module, a warp scheduler, an instruction cache module, a scoreboard, an emission module, and an execution unit;
[0022] The decoding module is responsible for reading instructions from the instruction cache module, obtaining decoding information after decoding, and distinguishing long delay operation instructions and short delay operation instructions in combination with RISC-V instructions, generating a corresponding instruction judgment signal long_short_judge, and sending the decoded instructions to the instruction buffer module;
[0023] The instruction buffer module is responsible for setting the corresponding instruction buffer depth and status flag full signal for each warp, and sending the instruction judgment signal long_short_judge and the status flag full signal to the warp scheduler to determine the warp status, distinguish between ready warps and pending warps, avoid fetching instructions from a full instruction buffer, achieve efficient warp switching and scheduling, and ensure the continuous operation of the pipeline;
[0024] The warp scheduler is the core module of the scheduling strategy, responsible for dividing warps into ready warps and pending warps based on the status of the synchronization fence and the status flag full signal of the instruction buffer, and distinguishing them through the flag bit S; at the same time, grouping the ready warps according to long delay operations and short delay operations, and following the principle of long delay operation priority and the principle that warps in the same group are arranged in ascending order of warp ID for priority arrangement;
[0025] The instruction cache module is responsible for reading and caching the operation instructions of the user, and at the same time, caching the instructions that have been grouped and arranged in priority in the warp scheduler;
[0026] The scoreboard is responsible for giving the valid signal scoreboard_busy signal;
[0027] The emission module determines whether to send the corresponding warp and its instructions to the execution unit according to the valid signal scoreboard_busy signal;
[0028] The execution unit is responsible for executing the operation instructions of the warp.
[0029] The decoding module uses the lower 7-bit instruction opcode in the RISC-V instruction to distinguish between long-delay operation instructions and short-delay operation instructions; after decoding, it sets the instruction judgment signal long_short_judge of the long-delay operation instruction to high level, and sets the instruction judgment signal long_short_judge of the short-delay operation instruction to low level.
[0030] When the buffer of the warp is not stored, the instruction buffer module sets the status flag full signal to 0; when the buffer of the warp is not full, the instruction buffer module sets the status flag full signal to 1; when the buffer of the warp is full, the instruction buffer module immediately sets the status flag full signal returned to the warp scheduler to 2, indicating that the warp cannot continue to fetch instructions, and then switches to the next fetching warp, so as to realize the fast switching of the fetching warp.
[0031] When the valid signal scoreboard_busy signal is 0, the emission module sends the corresponding warp and its instructions to the execution unit, otherwise it directly switches to the next ready warp.
[0032] A warp scheduler device for optimizing long-delay operations, characterized by including:
[0033] One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the above methods.
[0034] A readable storage medium, characterized in that: a computer program is stored on the readable storage medium, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0035] The beneficial effects of the present invention are: the warp scheduling method and system for optimizing long-delay operations reduce pipeline stalls, improve the ability of the general-purpose graphics processing unit (GPGPU) to hide delays, enhance the thread scheduling efficiency and computing resource utilization rate, and thus enhance the performance of the general-purpose graphics processing unit (GPGPU). BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0037] Appendix Figure 1 It is a schematic diagram of the thread bundle scheduling architecture of the present invention.
[0038] Appendix Figure 2 It is a schematic diagram of the logic flow of the thread bundle scheduling strategy of the present invention. Specific Embodiments
[0039] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0040] The thread bundle scheduling method for optimizing long-delay operations includes the following steps:
[0041] Step S1: Divide the thread bundles into ready thread bundles and pending thread bundles based on the synchronization fence and the status flag full signal status of the instruction buffer, and distinguish them through the flag bit S;
[0042] Step S2: Group the ready thread bundles according to long-delay operations and short-delay operations, and follow the principle of giving priority to long-delay operations, as well as the principle that the thread bundles in the same group are arranged in ascending order of the thread bundle ID for priority arrangement;
[0043] Step S3: Sequentially launch the thread bundles to the execution unit, use short-delay operations to hide the delay, and determine whether to send instructions to the execution unit according to the scoreboard valid signal scoreboard_busy to ensure efficient utilization of computing resources.
[0044] In the said step S1, the thread bundles are classified according to the synchronization fence status of the thread bundles and whether the status flag full signal from the instruction buffer is 0, specifically as follows:
[0045] If the synchronization fence signal of the thread bundle is 1 or the status flag full signal of the instruction buffer of the thread bundle is 0, then the thread bundle is a pending thread bundle; otherwise, it is a ready thread bundle;
[0046] The flag bit S is used to identify whether the current thread bundle is a ready thread bundle. The flag bit S of the ready thread bundle is 1, and the flag bit S of the pending thread bundle is 0;
[0047] When the synchronization fence status of the thread bundle and the status flag full signal from the instruction buffer change, the flag bit of the corresponding thread bundle is synchronously updated.
[0048] In step S2, the ready warps are grouped according to the instruction judgment signal long_short_judge from the instruction buffer module corresponding to the warp:
[0049] If the instruction judgment signal long_short_judge is at a high level, it is determined that the corresponding warp executes a long delay operation, and the warp is added to the long delay operation warp group;
[0050] If the instruction judgment signal long_short_judge is at a low level, it is determined that the corresponding warp executes a short delay operation, and the warp is placed in the short delay operation warp group.
[0051] In step S3, when the valid signal scoreboard_busy signal given by the scoreboard is 0, the corresponding warp and its instructions are sent to the execution unit; otherwise, the next ready warp is directly switched, so as to issue the warp that meets this condition.
[0052] Appendix Figure 2 The specific implementation process of the optimization strategy described in detail. First, determine the state of the warp. As Figure 2 shown, there are eight warps w1, w2, w3, w4, w5, w6, w7, and w8 waiting for scheduling. Among them, the ready warps are w1, w2, w3, w4, w5, w6, and the pending warps are w7, w8.
[0053] The ready warps are grouped according to long delay operations and short delay operations. The warp queues w1, w2, w5, w6 are long delay operation warps, and the warp queues w3, w4 are short delay operation warps.
[0054] According to the idea that long delay operations take precedence over short delay operations, the warp that executes the long delay instruction is placed first. For multiple warps with long delay operations in the long delay operation warp group, they are arranged in ascending order according to the warp ID; for warps that execute short delay instructions, they are arranged in the same way. The arranged warp queues w1, w2, w5, w6, w3, w4 can be obtained as shown in the appendix Figure 2 shown.
[0055] When there are no executable long delay operation warps in each loop, the short delay operation warps will be executed, so as to use short delay operations to hide the delay.
[0056] The warp scheduling system for optimizing long delay operations includes a decoding module, an instruction buffer module, a warp scheduler, an instruction cache module, a scoreboard, an issue module, and an execution unit;
[0057] The decoding module is responsible for reading instructions from the instruction cache module, obtaining decoding information after decoding, differentiating long-delay operation instructions and short-delay operation instructions in combination with RISC-V instructions, generating a corresponding instruction judgment signal long_short_judge, and sending the decoded instructions to the instruction buffer module;
[0058] The instruction buffer module is responsible for setting the corresponding instruction buffer depth and status flag full signal under each warp, and sending the instruction judgment signal long_short_judge and the status flag full signal to the warp scheduler to determine the warp status, differentiate ready warps and pending warps, avoid fetching instructions to a full instruction buffer, realize efficient warp switching and scheduling, and ensure continuous operation of the pipeline;
[0059] The warp scheduler is the core module of the scheduling strategy, responsible for dividing warps into ready warps and pending warps based on the status of the synchronization barrier and the status flag full signal of the instruction buffer, and differentiating them through the flag bit S; at the same time, grouping the ready warps according to long-delay operations and short-delay operations, following the principle of long-delay operation priority, and the principle that warps in the same group are arranged in ascending order of warp ID for priority arrangement;
[0060] The instruction cache module is responsible for reading and caching the user's operation instructions, and at the same time, caching the instructions grouped and arranged in priority in the warp scheduler;
[0061] The scoreboard is responsible for giving a valid signal scoreboard_busy signal;
[0062] The issuing module determines whether to send the corresponding warp and its instructions to the execution unit according to the valid signal scoreboard_busy signal;
[0063] The execution unit is responsible for executing the operation instructions of the warp.
[0064] The decoding module uses the lower 7-bit instruction opcode in the RISC-V instruction to differentiate long-delay operation instructions and short-delay operation instructions; when encountering a memory access operation instruction, which is a long-delay operation instruction type, after decoding, set the instruction judgment signal long_short_judge signal of the long-delay operation instruction to high level; when encountering an integer operation, which is a short-delay operation instruction type, after decoding, set the instruction judgment signal long_short_judge signal of the short-delay operation instruction to low level.
[0065] When the buffer of a warp is not storing a signal, the instruction buffer module sets the status flag full signal to 0; when the buffer of a warp is not full, the instruction buffer module sets the status flag full signal to 1; when the buffer of a warp is full, the instruction buffer module immediately sets the status flag full signal returned to the warp scheduler to 2, indicating that this warp cannot continue fetching instructions, and then switches to the next fetching warp, thereby realizing the fast switching of fetching warps.
[0066] When the valid signal scoreboard_busy signal is 0, the emission module sends the corresponding warp and its instructions to the execution unit, otherwise directly switches to the next ready warp.
[0067] The warp scheduler device for optimizing long-latency operations includes:
[0068] One or more processors, one or more memories, and one or more programs, where one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any one of the above methods.
[0069] The readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.
[0070] In summary, the warp scheduling method and system for optimizing long-latency operations preferentially schedule warps to perform long-latency operations, enabling long-latency operations to hide part of the latency from each other, then using short-latency operations to hide the latency, and combined with the warp scheduling architecture, enhancing the warp scheduling efficiency, thereby realizing the optimization of long-latency operations.
[0071] The above-described embodiments are only one of the specific implementation manners of the present invention, and the common changes and substitutions made by those skilled in the art within the scope of the technical solution of the present invention should be included in the protection scope of the present invention.
Claims
1. A thread warp scheduling method for optimizing long-latency operations, characterized in that: The following steps are involved: Step S1, dividing the thread warps into ready thread warps and pending thread warps based on the synchronization fence and the instruction buffer status flag full signal state, and distinguishing them by the flag bit S; Step S2: group the ready thread warps according to the long-delay operation and the short-delay operation, and arrange the priorities according to the principle of giving priority to the long-delay operation and the principle of arranging the thread warps in the same group from small to large according to the thread warp ID; Step S3: Send the thread warps to the execution unit in sequence, use the short-delay operation to hide the delay, and decide whether to send the instruction to the execution unit according to the scoreboard valid signal scoreboard_busy.
2. The thread warp scheduling method for optimizing long-latency operations according to claim 1, characterized in that: In step S1, the thread warps are classified according to the synchronization barrier status of the thread warps and whether the status flag full signal from the instruction buffer is 0, as follows: If the synchronization fence signal of the thread warp is 1 or the status flag full signal of the instruction buffer of the thread warp is 0, the thread warp is a pending thread warp; otherwise, it is a ready thread warp; The flag bit S is used to identify whether the current thread warp is a ready thread warp. The flag bit S of a ready thread warp is 1, and the flag bit S of a pending thread warp is 0. When the synchronization fence state of the thread warp and the status flag full signal from the instruction buffer change, the flag bit of the corresponding thread warp is updated synchronously.
3. The thread warp scheduling method for optimizing long-latency operations according to claim 1, characterized in that: In the step S2, the ready thread warps are grouped according to the instruction judgment signal long_short_judge from the instruction buffer module corresponding to the thread warp: If the instruction judgment signal long_short_judge is at a high level, it is determined that the corresponding thread warp executes the long delay operation, and the thread warp is added to the long delay operation thread warp group; If the instruction judgment signal long_short_judge is at a low level, it is determined that the corresponding thread warp performs the short-delay operation, and the thread warp is placed in the short-delay operation thread warp group.
4. The thread warp scheduling method for optimizing long-latency operations according to claim 1, characterized in that: In step S3, when the scoreboard gives a valid signal scoreboard_busy signal of 0, the corresponding thread warp and its instructions are sent to the execution unit, otherwise the next ready thread warp is directly switched to transmit the thread warp that meets the condition.
5. A thread warp scheduling system for optimizing long-latency operations, characterized in that: It includes a decoding module, an instruction buffer module, a thread warp scheduler, an instruction cache module, a scoreboard, a launch module and an execution unit; The decoding module is responsible for reading instructions from the instruction cache module, obtaining decoding information after decoding, distinguishing long-delay operation instructions from short-delay operation instructions in combination with RISC-V instructions, generating a corresponding instruction judgment signal long_short_judge, and sending the decoded instructions to the instruction buffer module; The instruction buffer module is responsible for setting the instruction buffer depth and status flag full signal corresponding to each thread warp, and sending the instruction judgment signal long_short_judge and the status flag full signal to the thread warp scheduler to determine the thread warp status, distinguish the ready thread warp from the pending thread warp, avoid continuing to fetch instructions from the full instruction buffer, realize thread warp switching and scheduling, and ensure the continuous operation of the pipeline; The thread warp scheduler is responsible for dividing the thread warps into ready thread warps and pending thread warps based on the synchronization fence and the instruction buffer status flag full signal state, and distinguishing them through the flag bit S; at the same time, the ready thread warps are grouped according to the long delay operation and the short delay operation, and the priority is arranged according to the principle of long delay operation priority and the principle of arranging the thread warps in the same group from small to large according to the thread warp ID; The instruction cache module is responsible for reading and caching the user's operation instructions, and at the same time, caching the instructions in the thread warp scheduler that have been grouped and arranged by priority; The scoreboard is responsible for giving a valid signal scoreboard_busy signal; The transmitting module determines whether to send the corresponding thread warp and its instructions to the execution unit according to the valid signal scoreboard_busy; The execution unit is responsible for executing the operation instructions of the thread warp.
6. The thread warp scheduling system for optimizing long-latency operations according to claim 5, characterized in that: The decoding module uses the lower 7-bit instruction opcode in the RISC-V instruction to distinguish between long-delay operation instructions and short-delay operation instructions; after decoding, the instruction judgment signal long_short_judge signal of the long-delay operation instruction is set to a high level, and the instruction judgment signal long_short_judge signal of the short-delay operation instruction is set to a low level.
7. The thread warp scheduling system for optimizing long-latency operations according to claim 5, characterized in that: When the buffer of the thread warp does not store any signal, the instruction buffer module sets the status flag full signal to 0; when the buffer of the thread warp is not full, the instruction buffer module sets the status flag full signal to 1; when the buffer of the thread warp is full, the instruction buffer module immediately sets the status flag full signal returned to the thread warp scheduler to 2, indicating that the thread warp cannot continue to fetch instructions, and then switches to the next instruction fetch thread warp, thereby realizing fast switching of instruction fetch thread warps.
8. The thread warp scheduling system for optimizing long-latency operations according to claim 5, characterized in that: When the valid signal scoreboard_busy is 0, the emission module sends the corresponding thread warp and its instructions to the execution unit, otherwise directly switches to the next ready thread warp.
9. A thread warp scheduling device for optimizing long-latency operations, characterized in that: include: One or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and are configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods according to claims 1 to 4.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.