A task scheduling system and method based on warp reorganization
Through the task scheduling system based on thread bundle reorganization, the problem of computing power waste in GPU during branch jump calculation is solved, and higher computing power utilization and effective utilization of computing units are achieved.
Patent Information
- Application Number
- CN202510020088.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-07
AI Technical Summary
When GPU processes branch jump calculation, there is a huge waste of computing power, resulting in extremely low utilization of overall computing power, which seriously restricts the expansion of GPU's application field.
The task scheduling system based on thread bundle reorganization is adopted, and the thread block scheduling module, instruction execution pipeline, operand collector and write back module are used to realize the reorganization and cache of thread bundles, and branch execution and data prefetching are optimized.
By reducing the number of repeated read and write times of thread bundles, the read and write efficiency of register files and operand collectors are improved, the effective utilization rate of the computing unit is enhanced, the zero-delay switching of thread bundles is realized, and the computing power utilization rate of the GPU is improved.
Smart Images

Figure CN119415241B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of control and scheduling of GPU computing resources, and particularly relates to a task scheduling system and method based on warp recombination. Background Art
[0002] With the booming development of artificial intelligence and big data computing, GPUs are playing an increasingly important role in multiple industries, and computing power performance has become one of the main technical indicators for evaluating computing processors such as CPUs, GPUs, and FPGAs.
[0003] Currently, the total computing power of a single chip of GPUs or AI computing devices from domestic and foreign manufacturers is continuously increasing. To maximize the computing power of computing devices, manufacturers have optimized from multiple aspects such as computing power scheduling, storage access methods, and cache bandwidth, and achieved certain results.
[0004] However, to ensure the high parallel computing ability of GPUs, GPUs have reduced functions such as data prefetching and complex logic calculations. When dealing with branch jump calculations, there is generally a huge waste of computing power, resulting in an extremely low overall computing power utilization rate of GPUs in some application scenarios, seriously restricting the expansion of the application fields of GPUs. Summary of the Invention
[0005] Aiming at the problem that there is generally a huge waste of computing power when dealing with branch jump calculations, resulting in an extremely low overall computing power utilization rate of GPUs in some application scenarios and seriously restricting the expansion of the application fields of GPUs, the present invention provides a task scheduling system and method based on warp recombination.
[0006] In a first aspect, the technical solution of the present invention further provides a task scheduling system based on warp recombination, including a thread block scheduling module, an instruction execution pipeline, a first-level operand collector, a second-level operand collector, a distribution module, an execution module, and a write-back module; the second-level operand collector is provided with a recombined warp cache module;
[0007] The thread block scheduling module caches all thread blocks that need to be calculated; each thread block contains several warps; the thread block scheduling module controls the instruction execution pipeline to fetch instructions corresponding to the same program counter of the current branch of the thread block for decoding; the instructions involve read and write operations on registers in the warp;
[0008] The first-level operand collector reads and stores the register file data corresponding to each warp in the thread block, and sends all the operands of the warp to the second-level operand collector; among them, the register file data is the operand required for executing the instruction;
[0009] When polling different warps of a thread block, the second-level operand collector extracts the active threads in the warp according to the active flag of the warp mask information, reorganizes the extracted active threads into a warp, and caches them in the reorganized warp cache module.
[0010] The distribution module monitors the reorganized warp cache module in real time. When the reorganized warp cache module is not empty, it reads the reorganized warp in the reorganized warp cache module and sends it to the subsequent execution module.
[0011] After the execution of the reorganized warp ends and when writing back to the register file, the write-back module splits the operand of the calculation result of the reorganized warp according to the mask information of the original warp and writes it back to the corresponding register file of the original warp.
[0012] As a further limitation of the technical solution of the present invention, after all branches are executed and during the execution of non-branch operations, the thread block scheduling module controls the bypass of the second-level operand collector, so that the distribution module directly reads the operands of the first-level operand collector and sends them to the execution module.
[0013] As a further limitation of the technical solution of the present invention, the second-level operand collector is provided with an active thread extraction module.
[0014] The active thread extraction module extracts the active warps in the warp according to the active flag of the warp mask information, rearranges the active threads in the warp in ascending order, and generates reorganized threads to be stored in the reorganized warp cache module.
[0015] As a further limitation of the technical solution of the present invention, the second-level operand collector is also provided with a warp mask cache module.
[0016] The mask information corresponding to the warp in all the operands of the warp issued by the first-level operand collector is sequentially stored in the warp mask cache module.
[0017] As a further limitation of the technical solution of the present invention, the instruction execution pipeline is provided with an instruction decoding module, and the second-level operand collector is provided with a branch reorganized warp cache module.
[0018] During branch calculation, when polling different warps of a thread block, there are idle cycles in the instruction decoding module. The thread block scheduling module uses the idle cycles to control the instruction execution pipeline to read the calculation instructions of other branches for decoding. The first-level operand collector sequentially reads the warps of the thread block, reads the data of the register file corresponding to the warp. After the first-level operand collector completes the operand collection, it sends them to the active thread extraction module to extract active threads, reorganize them into a new warp, and send them to the branch reorganized warp cache module.
[0019] As a further limitation of the technical solution of the present invention, the reorganized thread bundle cache module monitors the barrier or merge instruction of branch execution according to the execution status of the instruction, marks the last group of valid reorganized thread bundles, and sends them to the distribution module in the form of status marks; at the same time, after the instruction execution pipeline obtains the branch barrier or merge instruction, it sends a status mark to the thread block scheduling module for control. The thread block scheduling module controls and issues the calculation instructions of other branches to the instruction execution pipeline to start executing the calculations of the remaining other branches.
[0020] As a further limitation of the technical solution of the present invention, after the distribution module reads the last group of reorganized thread bundles in the reorganized thread bundle cache module, it monitors the branch reorganized thread bundle cache module and executes the thread bundles of the new branch; when the branch reorganized thread bundle cache module is in a non-empty state, it reads the reorganized thread bundles in the branch reorganized thread bundle cache module, and after the branch reorganized thread bundle cache module is read empty, it switches to the reorganized thread bundle cache module to read the subsequent reorganized thread bundles of the new branch of the reorganized thread bundle cache module; the distribution module marks the empty state of the branch reorganized thread bundle cache module and sends it to the thread block scheduling module. The thread block scheduling module controls the execution of data prefetching of other branches, or waits for all branches to finish execution and executes the post-branch merge instructions.
[0021] As a further limitation of the technical solution of the present invention, after the reorganized thread bundle finishes execution and writes back to the register file, the write-back module reads the thread bundle mask cache module of the second-level operand collector, splits the calculation result operands of the reorganized thread bundle according to the mask information of the original thread bundle, and writes them back to the corresponding register file of the original thread bundle.
[0022] In a second aspect, the technical solution of the present invention provides a task scheduling method based on thread bundle reorganization, including the following steps:
[0023] Cache all thread blocks that need to be calculated; each thread block contains several thread bundles;
[0024] Take out the instructions corresponding to the same program counter of the current branch of each thread block for decoding, and read and store the register file data corresponding to each thread bundle in the thread block;
[0025] When polling different thread bundles of the thread block, according to the active mark of the thread bundle mask information, extract the active threads in the thread bundle, reorganize the extracted active threads into thread bundles; and cache them in the reorganized thread bundle cache module;
[0026] Monitor the reorganized thread bundle cache module in real time. When the reorganized thread bundle cache module is in a non-empty state, read the reorganized thread bundles in the reorganized thread bundle cache module and send them to the subsequent execution unit;
[0027] After the execution of the reorganized warp ends and when writing back to the register file, according to the mask information of the original warp, split the operand of the calculation result of the reorganized warp and write it back to the corresponding register file of the original warp.
[0028] As a further limitation of the technical solution of the present invention, the step of performing warp reorganization on the extracted active threads includes:
[0029] Extract the active threads in each warp in sequence according to the warp number and store them in the reorganized warp. When the reorganized warp is full, automatically store them in the next reorganized warp; wherein, the number of threads in the reorganized warp is the same as the number of threads in the warp before reorganization.
[0030] As a further limitation of the technical solution of the present invention, the method further includes:
[0031] The reorganized warp cache module monitors the barrier or merge instruction of branch execution according to the execution status of the instruction, marks the last group of valid reorganized warps, and sends them to the distribution module in the form of status marks; at the same time, after the instruction execution pipeline obtains the barrier or merge instruction of branch execution, it sends a status mark;
[0032] Based on the status mark, control the instruction corresponding to the starting program counter of other branches to be issued to the instruction execution pipeline to start the calculation of other branches.
[0033] As a further limitation of the technical solution of the present invention, the method further includes:
[0034] During branch calculation, poll different warps of the thread block. When there is an idle cycle, read the instructions for calculating other branches, and sequentially read the warps of the corresponding thread block, and read the register file corresponding to the warp for storage.
[0035] Extract the active threads in the warp according to the active mark of the warp mask information, perform warp reorganization on the extracted active threads; and cache them in the branch reorganized warp cache module.
[0036] As a further limitation of the technical solution of the present invention, the method further includes:
[0037] After reading the last group of reorganized warps in the reorganized warp cache module, monitor the branch reorganized cache module and execute the warps of the new branch; when the branch reorganized cache module is in a non-empty state, read the reorganized warps in the branch reorganized cache module, and after reading it empty, switch to the reorganized warp cache module to read the subsequent reorganized warps of the new branch.
[0038] As a further limitation of the technical solution of the present invention, the method further includes:
[0039] After all branches have finished execution, during the non-branch operation, directly read the data stored in the register file and send it to the subsequent execution unit.
[0040] The beneficial effects of the technical solution of the present invention are as follows: By means of a scheduling method based on warp polling within a thread block, sequentially execute all warps with branch operations, reduce the number of repeated reads and writes of warps, and improve the read and write efficiency of the register file and operand collector; Collect threads of the same type of branch operation through the second-level operand collector to form a full-load warp, improving the effective utilization rate of the computing unit; Prefetch the operands of other branch operations through the second-level operand collector to achieve zero-latency switching of warps, further improving the effective utilization rate of the computing unit. Description of the Drawings
[0041] In order to more clearly illustrate the technical solution of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 It is a schematic diagram of the system provided by the embodiment of the present invention.
[0043] Figure 2 It is a schematic diagram of the reorganization method of the branch operation warp in the embodiment of the present invention.
[0044] Figure 3 It is a schematic flowchart of the method provided by the embodiment of the present invention. Detailed Embodiments
[0045] The present invention mainly relates to a warp scheduling management method for GPU hardware resources. Software generates instruction data for GPU execution and issues it, and sets a special warp barrier instruction, so that during the execution of the GPU, it stops executing after reaching the thread branch instruction preferentially. After ensuring that all warps in the same thread block reach the barrier instruction, it enters the branch execution process. In order to make the purpose, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below in conjunction with the drawings in the specific embodiments of the present invention. Obviously, the embodiments described below are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.
[0046] Such as Figure 1As shown, the technical solution of the present invention further provides a task scheduling system based on warp recombination, including a thread block scheduling module, an instruction execution pipeline, a first-level operand collector, a second-level operand collector, a distribution module, an execution module, and a write-back module; the second-level operand collector is provided with a recombined warp cache module;
[0047] The thread block scheduling module caches all the thread blocks that need to be calculated; each thread block contains several warps; the thread block scheduling module controls the instruction execution pipeline to fetch the instructions corresponding to the same program counter of the current branch of the thread block for decoding; the instructions involve read and write operations on the registers in the warp.
[0048] The first-level operand collector reads the register file data corresponding to each warp in the thread block for storage, and sends all the operands of the warp to the second-level operand collector; among them, the register file data is the operands required for executing the instructions.
[0049] When polling different warps of the thread block, the second-level operand collector extracts the active threads in the warp according to the active flag of the warp mask information, recombines the extracted active threads, and caches them in the recombined warp cache module.
[0050] The distribution module monitors the recombined warp cache module in real time. When the recombined warp cache module is in a non-empty state, it reads the recombined warp in the recombined warp cache module and sends it to the subsequent execution module.
[0051] After the recombined warp execution ends and writes back to the register file, the write-back module splits the calculation result operands of the recombined warp according to the mask information of the original warp and writes them back to the corresponding register file of the original warp.
[0052] After all branches are executed and during the execution of non-branch operations, the thread block scheduling module controls the bypass of the second-level operand collector, so that the distribution module directly reads the operands of the first-level operand collector and sends them to the execution module.
[0053] In some embodiments, the second-level operand collector is provided with an active thread extraction module;
[0054] The active thread extraction module extracts the active warps in the warp according to the active flag of the warp mask information, and rearranges the active threads in the warp in ascending order to generate a recombined thread number and stores it in the recombined warp cache module.
[0055] In some embodiments, the second-level operand collector is further provided with a warp mask cache module;
[0056] The mask information corresponding to the warp in all the operands of the warp issued by the first-level operand collector is sequentially stored in the warp mask cache module.
[0057] In some embodiments, an instruction decoding module is provided in the instruction execution pipeline, and a branch recombined warp cache module is provided in the second-level operand collector;
[0058] During branch calculation, when polling different warps of a thread block, there are idle cycles in the instruction decoding module. The thread block scheduling module uses these idle cycles to control the instruction execution pipeline to read the calculation instructions of other branches for decoding. The first-level operand collector sequentially reads the warps of the thread block, reads the data of the register file corresponding to the warp. After the first-level operand collector completes operand collection, it sends the data to the active thread extraction module to extract active threads, perform warp recombination to form new warps, and send them to the branch recombined warp cache module.
[0059] In some embodiments, the recombined warp cache module monitors the barrier or merge instruction of branch execution according to the execution status of the instruction, marks the last set of valid recombined warps, and sends them to the distribution module in the form of status marking; at the same time, after the instruction execution pipeline obtains the branch barrier or merge instruction, it sends the status marking to the thread block scheduling module for control. The thread block scheduling module controls the issuance of calculation instructions of other branches to the instruction execution pipeline to start executing the calculation of the remaining other branches.
[0060] In some embodiments, after the distribution module reads the last set of recombined warps in the recombined warp cache module, it monitors the branch recombined warp cache module and executes the warps of the new branch; when the branch recombined warp cache module is in a non-empty state, it reads the recombined warps in the branch recombined warp cache module, and after the branch recombined warp cache module is read empty, it switches to the recombined warp cache module to read the subsequent recombined warps of the new branch in the recombined warp cache module; the distribution module marks the read-empty state of the branch recombined warp cache module and sends it to the thread block scheduling module. The thread block scheduling module controls the execution of data prefetching of other branches, or waits for all branches to finish execution and executes the post-branch merge instructions.
[0061] In some embodiments, after the recombined warp finishes execution and writes back to the register file, the write-back module reads the warp mask cache module of the second-level operand collector, splits the calculation result operands of the recombined warp according to the mask information of the original warp, and writes them back to the corresponding register file of the original warp.
[0062] Based on the system provided by the embodiments of the present invention, the steps for task scheduling include:
[0063] S1. The thread block scheduling module caches all the thread blocks that need to be calculated. When making a branch condition determination, it fixedly executes a one-way conditional branch. The thread block scheduling module controls the instruction pipeline, fetches the instruction corresponding to the PC of the current branch of a thread block for decoding, and reads the register file data corresponding to the warp of that thread block. After the first-level operand collector completes operand collection, it directly sends all the operands of the warp to the second-level operand collector.
[0064] Among them, the active thread extraction module extracts the active warps in the warp according to the active flags of the warp masks. The active thread extraction module rearranges the active threads in the warp in ascending order. As Figure 2 shown, a single warp contains 8 threads. The solid lines are active threads, and the dashed lines are inactive threads. Warp 1 sent by the first-level operand collector contains four active threads. After being extracted by the active thread extraction module, they are stored in the first four positions of the reorganized warp 1. In the same way, in the next cycle, the active threads ②, ③, ⑥, ⑦ of warp 2 are extracted and stored in the last four positions of the reorganized warp 1. In the third cycle, the active thread ⑧ of warp 2 is stored in the first position of the reorganized warp 2; the active threads ②, ③ of warp 3 are extracted and stored in the second and third positions of the reorganized warp 2; the active threads ①, ②, ③, ⑥, ⑦ of warp 4 are extracted and stored in the last five positions of the reorganized warp 2.
[0065] The reorganized warps are stored in the reorganized warp cache module. The cache depth is set to 3, which not only maintains the reading and writing of the reorganized warps but also enables the remaining active warps to continue to be stored after the warps sent by the first-level operand collector fill a reorganized warp, reducing the power consumption of re-reading the warps.
[0066] At the same time, the masks corresponding to the warps sent by the first-level operand collector are sequentially stored in the warp mask cache module for data write-back operations; the depth of the warp mask cache module is N (the maximum number of warps included in a thread block allowed by the hardware), where N is the depth of the branch reorganized warp cache module; when the thread block scheduling module processes the warp branch calculation, it enables the warp polling scheduling within the thread block to provide active threads with the same program counter (PC) for the two-level operand collectors.
[0067] After a single branch instruction jumps, before the branch directly ends, the masks corresponding to different PCs of the warp are exactly the same. Therefore, the warp mask cache module only stores the warp mask when executing the first PC of the branch operation.
[0068] In the embodiment of the present invention, the depth of the reorganized warp cache module is adjustable; the total number of threads included in a single warp is adjustable. The 8 threads shown in the present invention are only for illustrative purposes;
[0069] S2. During branch calculation, when polling different thread bundles of the thread block, there is an idle cycle in the instruction decoding module; the thread block scheduling module controls the reading of calculation instructions of other branches, and sequentially reads the low-order thread bundles of the thread block, reads the register file corresponding to the thread bundle, and after the first-level operand collector completes the operand collection, sends it to the active thread extraction module, extracts the active thread, reorganizes it into a new thread bundle according to the thread bundle reorganization mode of S1, and sends it to the branch reorganization thread bundle cache module; preferably, the active thread extraction module can be merged into one, and the thread bundle reorganization of the reorganization thread bundle cache module is executed first;
[0070] S3, the distribution module monitors the reorganized thread bundle cache module in real time, and when the reorganized thread bundle cache module is in a non-empty state, reads the reorganized thread bundle and sends it to the subsequent execution module;
[0071] The reorganized thread warp cache module monitors the barrier or merge instructions of branch execution according to the execution status of the instruction, marks the last group of valid reorganized thread warps, and sends them to the distribution module through the status mark; at the same time, after the instruction execution pipeline obtains the branch barrier or merge instruction, it sends the status mark to the thread block scheduling module for control, and the thread block scheduling module controls the sending of the starting PC of other branches to the instruction pipeline to start executing other branches;
[0072] After the distribution module reads the last set of reorganized thread bundles, it monitors the branch reorganized thread bundle cache module and executes the thread bundles of the new branch; when the branch reorganized thread bundle cache module is in a non-empty state, it reads the reorganized thread bundles therein, and after reading empty, switches to the reorganized thread bundle cache module to read the subsequent reorganized thread bundles of the new branch; the distribution module marks the branch reorganized thread bundle cache module as being in an empty state and sends it to the thread block scheduling module, which controls the data prefetching of other branches, or waits for the execution of all branches to be completed and executes the instructions after the branch merge.
[0073] S4. After the reorganized thread warp is executed, when writing back to the register file, the write-back module reads the thread warp mask cache module of the second-level operand collector, splits the calculation result operands of the reorganized thread warp according to the mask information of the original thread warp, and writes them back to the corresponding register file of the basic thread warp.
[0074] After all branches are executed, during the execution of non-branch operations, the thread block scheduling module controls the bypass function of the second-level operand collector, so that the distribution module directly reads the operands of the first-level operand collector and sends them to the subsequent execution module for execution.
[0075] The second-level operand collector collects the active branch threads emitted by the first-level operand collector and reorganizes them into a complete warp adapted to the computing unit; the thread block scheduling module, when processing the warp branch calculation, enables the warp polling scheduling within the thread block, provides active threads with the same program counter (PC) for the two-level operand collector, and when the decoding is idle, sends instructions of another branch PC to control the operand collector to perform data prefetching of the corresponding branch warp; the second-level operand collector bypass function bypasses the second-level operand collector during non-branch calculations and directly sends the data of the first-level operand collector to the distribution module to execute parallel computing tasks. The present invention maximizes the reduction of the frequency of idle computing units and improves the computing power utilization rate of the GPU through the scheduling control method of warp reorganization.
[0076] As Figure 3 shown, an embodiment of the present invention provides a task scheduling method based on warp reorganization, including the following steps:
[0077] Step 1: Cache all thread blocks that need to be calculated; each thread block contains several warps;
[0078] Step 2: Take out the instructions corresponding to the same program counter of the current branch of each thread block for decoding, and read and store the register file data corresponding to each warp in the thread block;
[0079] Step 3: When polling different warps of the thread block, extract the active threads in the warp according to the active flag of the warp mask information, reorganize the extracted active threads into a warp; and cache them in the reorganized warp cache module;
[0080] Step 4: Monitor the reorganized warp cache module in real time. When the reorganized warp cache module is in a non-empty state, read the reorganized warp in the reorganized warp cache module and send it to the subsequent execution unit;
[0081] Step 5: After the execution of the reorganized warp ends, when writing back to the register file, split the calculation result operands of the reorganized warp according to the mask information of the original warp and write them back to the corresponding register file of the original warp.
[0082] In some embodiments, the step of reorganizing the extracted active threads into a warp includes:
[0083] Sequentially extract the active threads in each warp according to the warp number and store them in the reorganized warp. When the reorganized warp is full, it is automatically stored in the next reorganized warp; wherein, the number of threads in the reorganized warp is the same as the number of threads in the warp before reorganization.
[0084] Extract the active warps in a warp according to the active flags of the warp mask, and rearrange the active threads in the warp in ascending order. As Figure 2 shown, a single warp contains 8 threads. The solid lines are active threads and the dashed lines are inactive threads. Warp 1 issued by the first-level operand collector contains four active threads, and the active threads are extracted and stored in the first four positions of the reorganized warp 1. In the same way, in the next cycle, the active threads ②, ③, ⑥, ⑦ of warp 2 are extracted and stored in the last four positions of the reorganized warp 1. In the third cycle, the active thread ⑧ of warp 2 is stored in the first position of the reorganized warp 2; the active threads ②, ③ of warp 3 are extracted and stored in the second and third positions of the reorganized warp 2; the active threads ①, ②, ③, ⑥, ⑦ of warp 4 are extracted and stored in the last five positions of the reorganized warp 2.
[0085] In some embodiments, the method further includes:
[0086] The reorganized warp cache module monitors the barrier or merge instruction of branch execution according to the execution status of the instruction, marks the last group of valid reorganized warps, and sends them to the distribution module in the form of status flags; at the same time, after the instruction execution pipeline obtains the barrier or merge instruction of branch execution, it sends a status flag; based on the status flag, it controls the issuance of instructions corresponding to the starting program counters of other branches to the instruction execution pipeline to start the calculation of other branches.
[0087] During branch calculation, poll different warps of a thread block. When there is an idle cycle, read the instructions for other branch calculations, sequentially read the warps of the corresponding thread block, and read the register file corresponding to the warp for storage;
[0088] Extract the active threads in the warp according to the active flags of the warp mask information, reorganize the extracted active threads into warps; and cache them in the branch reorganized warp cache module.
[0089] In some embodiments, the method further includes:
[0090] After reading the last group of reorganized warps in the reorganized warp cache module, monitor the branch reorganized cache module and execute the warps of the new branch; when the branch reorganized cache module is in a non-empty state, read the reorganized warps in the branch reorganized cache module, and after reading them empty, switch to the reorganized warp cache module to read the subsequent reorganized warps of the new branch.
[0091] After all branches have been executed, during non-branch operations, directly read the stored register file data and send it to the subsequent execution unit.
[0092] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task scheduling system based on warp reorganization, characterized in that: It includes a thread block scheduling module, an instruction execution pipeline, a first-level operand collector, a second-level operand collector, a distribution module, an execution module and a write-back module; the second-level operand collector is provided with a reorganized thread warp cache module and a thread warp mask cache module; the mask information corresponding to the thread warp in all operands of the thread warp issued by the first-level operand collector is sequentially stored in the thread warp mask cache module; The thread block scheduling module caches all thread blocks that need to be calculated; each thread block contains a number of thread warps; the thread block scheduling module controls the instruction execution pipeline to extract the instructions corresponding to the same program counter of the current branch of the thread block for decoding; the instructions involve read and write operations on registers in the thread warp; The first-level operand collector reads the register file data corresponding to each thread warp in the thread block for storage, and sends all operands of the thread warp to the second-level operand collector; wherein the register file data is the operand required to execute the instruction; When polling different thread warps of a thread block, the second-level operand collector extracts active threads in the thread warp according to the active mark of the thread warp mask information, reorganizes the extracted active threads into thread warps, and caches them into the reorganized thread warp cache module; The distribution module monitors the reorganized thread bundle cache module in real time, and when the reorganized thread bundle cache module is in a non-empty state, reads the reorganized thread bundle in the reorganized thread bundle cache module and sends it to the subsequent execution module; After the reorganized thread warp is executed, when writing back to the register file, the write-back module reads the thread warp mask cache module of the second-level operand collector, splits the calculation result operands of the reorganized thread warp according to the mask information of the original thread warp, and writes them back to the corresponding register file of the original thread warp.
2. The task scheduling system based on thread warp reorganization according to claim 1, characterized in that: After all branches are executed, during the execution of non-branch operations, the thread block scheduling module controls the second-level operand collector to bypass, so that the distribution module directly reads the operands of the first-level operand collector and sends them to the execution module.
3. The task scheduling system based on thread warp reorganization according to claim 2, characterized in that: The second-level operand collector is provided with an active thread extraction module; The active thread extraction module extracts active thread warps in the thread warps according to the active mark of the thread warp mask information, and rearranges the active threads in the thread warps in order from low to high to generate a reorganized thread number and stores it in the reorganized thread warp cache module.
4. The task scheduling system based on thread warp reorganization according to claim 3, characterized in that: The instruction execution pipeline is provided with an instruction decoding module, and the second-level operand collector is provided with a branch reorganization thread warp cache module; During branch calculation, when polling different thread bundles of the thread block, there is an idle cycle in the instruction decoding module. The thread block scheduling module uses the idle cycle to control the instruction execution pipeline to read the calculation instructions of other branches for decoding. The first-level operand collector sequentially reads the thread bundles of the thread block and reads the data of the register file corresponding to the thread bundle. After the first-level operand collector completes the operand collection, it is sent to the active thread extraction module to extract the active thread, reorganize the thread bundle into a new thread bundle, and send it to the branch reorganization thread bundle cache module.
5. The task scheduling system based on thread warp reorganization according to claim 4, characterized in that: The reorganized thread warp cache module monitors the barrier or merge instructions of branch execution according to the execution status of the instructions, marks the last set of valid reorganized thread warps, and sends them to the distribution module through status marking; at the same time, after the instruction execution pipeline obtains the branch barrier or merge instruction, it sends the status mark to the thread block scheduling module control, and the thread block scheduling module controls the issuance of calculation instructions of other branches to the instruction execution pipeline and starts to execute the calculation of the remaining other branches.
6. The task scheduling system based on warp reorganization according to claim 5, characterized in that: After the distribution module reads the last group of reorganized thread bundles in the reorganized thread bundle cache module, it monitors the branch reorganized thread bundle cache module and executes the thread bundles of the new branch; when the branch reorganized thread bundle cache module is in a non-empty state, it reads the reorganized thread bundles in the branch reorganized thread bundle cache module, and after the branch reorganized thread bundle cache module is read empty, it switches to the reorganized thread bundle cache module and reads the subsequent reorganized thread bundles of the new branch of the reorganized thread bundle cache module; The distribution module marks the branch reorganization thread bundle cache module read empty state and sends it to the thread block scheduling module. The thread block scheduling module controls the data prefetching of other branches, or waits for the execution of all branches to end and executes the instructions after the branch merge.
7. A task scheduling method based on thread warp reorganization, characterized in that: The steps include: Cache all thread blocks that need to be calculated; each thread block contains several thread warps; The instructions corresponding to the same program counter of the current branch of each thread block are taken out for decoding, and the register file data corresponding to each thread warp in the thread block is read for storage; the mask information corresponding to the thread warp is sequentially stored in the thread warp mask cache module; When polling different thread warps of a thread block, active threads in the thread warp are extracted according to the active mark of the thread warp mask information, and the extracted active threads are reorganized into thread warps; And cache it to the reorganized thread warp cache module; Monitor the reorganized thread bundle cache module in real time, and when the reorganized thread bundle cache module is in a non-empty state, read the reorganized thread bundle in the reorganized thread bundle cache module and send it to the subsequent execution unit; After the reorganized thread warp is executed, when writing back to the register file, the thread warp mask cache module of the second-level operand collector is read, and the calculation result operands of the reorganized thread warp are split according to the mask information of the original thread warp, and written back to the corresponding register file of the original thread warp.
8. The task scheduling method based on warp reorganization according to claim 7, characterized in that: The steps of reorganizing the extracted active threads into thread warps include: Active threads in each warp are extracted in sequence according to the warp number and stored in the reorganized warp. When the reorganized warp is fully loaded, the threads are automatically stored in the next reorganized warp. The number of threads in the reorganized warp is the same as the number of threads in the warp before reorganization.
Citation Information
Patent Citations
Thread recombination method and device based on GPGPU (General Purpose Graphics Processing Unit) and medium
CN118819868A