GPGPU (General Purpose Graphics Processing Unit) thread bundle scheduling method and device, equipment and medium

By optimizing the thread beam scheduling strategy in GPGPU, and using cache access result signals to manage the active and suspended state of the thread beam, the problem of the impact of the access memory delay of the GPGPU computing performance is solved, and pipeline efficiency and computing performance are improved.

CN120066726APending Publication Date: 2025-05-30SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510202425.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The computing performance of GPGPU is severely limited by external memory access latency, resulting in reduced instruction pipeline efficiency.

Method used

By controlling the current thread beam of the instruction pipeline to enter the decoding stage, determine whether it is a fetch type thread beam, and add the thread beam to the active or suspended table based on the cache access result signal, optimizing the scheduling strategy of the thread beam.

Benefits of technology

It effectively avoids the memory fetch type thread bundle blocking the instruction pipeline for a long time, concealing the memory fetch delay, improving the efficiency and throughput of the pipeline, and improving the computing performance of the GPGPU.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066726A_ABST
    Figure CN120066726A_ABST
Patent Text Reader

Abstract

The invention discloses a GPGPU (General Purpose Graphics Processing Unit) thread bundle scheduling method, device and equipment and a medium, relates to the technical field of computers, is applied to a GPGPU, and comprises the following steps: if a current thread bundle is determined not to be a memory access type thread bundle according to a decoded instruction of the current thread bundle in an instruction pipeline, controlling the current thread bundle to enter an execution stage of the instruction pipeline; if the thread bundle is determined to be a memory access type thread bundle, acquiring a cache access result signal based on the decoded instruction; if the cache access result signal meets an active condition, adding the current thread bundle into an active thread bundle table to enter an execution stage; if the cache access result signal meets the hanging condition, adding the current thread bundle into a hanging thread bundle table, and scheduling the next thread bundle to carry out the process; and if it is monitored that a target thread bundle meeting an active condition exists in the suspended thread bundle table, adding the target thread bundle into an active thread bundle table to enter an execution stage. A GPGPU pipeline scheduling mechanism is optimized to improve pipeline efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and particularly to a warp scheduling method, apparatus, device and medium for GPGPU. Background Art

[0002] GPGPU (General-Purpose computing on Graphics Processing Units) is a technology that extends the computing power of a graphics processing unit to general computing tasks. Through this technology, GPGPU is not only used for graphics rendering, but can also efficiently execute various general computing tasks, including scientific computing, data analysis, deep learning, etc.

[0003] GPGPU uses an instruction pipeline to implement instruction fetching, decoding, memory access, arithmetic operation, and write-back. Among them, LSU (Load Store Unit) is an important module in the arithmetic operation stage of the pipeline, which is used to implement external memory access instructions. In the overall computing process of GPGPU, the external memory access for computing data severely limits the computing performance of GPGPU. Compared with the computing process, the external memory access process has a greater delay. When a warp executes an external memory access instruction, the waiting for external data will block the operation of the instruction pipeline and reduce the efficiency of the pipeline. Although currently LSU is made independent of the instruction pipeline, additional logic resources are still required to maintain the operation of the independent LSU, and the adjacent arithmetic instructions of the memory access instruction still need to wait for the memory access to complete, and the actual improvement effect is not significant, aiming at the problem of large memory access latency of memory access instructions.

[0004] In summary, how to optimize the GPGPU pipeline scheduling mechanism to improve the pipeline efficiency is a problem to be solved in this field. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a warp scheduling method, apparatus, device and medium for GPGPU, which optimizes the GPGPU pipeline scheduling mechanism to improve the pipeline efficiency. The specific solutions are as follows:

[0006] In a first aspect, the present application discloses a warp scheduling method for GPGPU, which is applied to GPGPU and includes:

[0007] Controlling the current warp of the instruction pipeline to enter the decoding stage to obtain the decoded instructions, and determining whether the current warp is a memory access type warp according to the decoded instructions;

[0008] If the current warp is not a memory access type warp, controlling the current warp to enter the execution stage of the instruction pipeline;

[0009] If the current warp is a memory access type warp, obtain a cache access result signal returned by the L1 Cache based on the decoded instruction;

[0010] If the cache access result signal meets a preset active condition, add the current warp to the active warp table, and control each warp in the active warp table to enter the execution stage of the instruction pipeline;

[0011] If the cache access result signal meets a preset suspension condition, add the current warp to the suspended warp table, update the next warp in the instruction pipeline to the new current warp, and re-jump to the step of controlling the current warp in the instruction pipeline to enter the decoding stage;

[0012] When it is monitored that there is a target warp in the suspended warp table that meets the preset active condition, add the target warp to the active warp table, and control each warp in the active warp table to enter the execution stage of the instruction pipeline.

[0013] Optionally, the GPGPU includes a warp management unit, an address calculation unit, and a Cache pre-memory access unit;

[0014] The obtaining the cache access result signal returned by the L1 Cache based on the decoded instruction includes:

[0015] Control the warp management unit to send the decoded instruction to the address calculation unit;

[0016] Control the address calculation unit to obtain the current memory access address according to the decoded instruction, and send the current memory access address to the Cache pre-memory access unit;

[0017] Control the Cache pre-memory access unit to send a memory access request carrying the current memory access address to the L1 Cache to obtain the cache access result signal returned by the L1 Cache based on the memory access request.

[0018] Optionally, obtaining the cache access result signal returned by the L1 Cache based on the memory access request includes:

[0019] Match the current memory access address in the memory access request in the cached data of the L1 Cache through the L1 Cache to generate a cache access result signal;

[0020] Obtain the cache access result signal returned by the L1 Cache.

[0021] Optionally, the GPGPU thread bundle scheduling method further includes:

[0022] If the cache access result signal is a hit signal, it is determined that the cache access result signal meets a preset active condition; wherein, the hit signal indicates that there is data in the cached data that matches the current memory access address;

[0023] If the cache access result signal is a miss signal, it is determined that the cache access result signal meets a preset suspension condition; wherein, the miss signal indicates that there is no data in the cached data that matches the current memory access address.

[0024] Optionally, during the process of adding the current thread bundle to the suspended thread bundle table, it further includes:

[0025] Store the current memory access address in the MSHR register, so that when the MSHR register meets the preset merged memory access condition, obtain target data that matches each memory access address stored in the MSHR register from the external memory, and save the target data to the L1 Cache to update the cached data in the L1 Cache.

[0026] Optionally, the GPGPU thread bundle scheduling method further includes:

[0027] Monitor the space occupied by the access addresses in the MSHR register and the first quantity of the thread bundles corresponding to the access addresses;

[0028] If the space occupied by the access addresses in the MSHR register is greater than a first preset threshold and / or the first quantity is greater than a second preset threshold, it is determined that the MSHR register meets the preset merged memory access condition.

[0029] Optionally, controlling each thread bundle in the active thread bundle table to enter the execution stage of the instruction pipeline includes:

[0030] Determine the second quantity of the thread bundles to be executed; wherein, the thread bundles to be executed are the thread bundles that currently need to enter the execution stage of the instruction pipeline;

[0031] If the second quantity is greater than a third preset threshold, control each thread bundle in the active thread bundle table to enter the execution stage of the instruction pipeline in sequence according to the priorities of the thread bundles to be executed.

[0032] In a second aspect, the present application discloses a GPGPU thread bundle scheduling device, which is applied to a GPGPU and includes:

[0033] A type judgment module, which is used to control the current warp of the instruction pipeline to enter the decoding stage to obtain the decoded instruction, and judge whether the current warp is a memory access type warp according to the decoded instruction;

[0034] A first execution module, which is used to control the current warp to enter the execution stage of the instruction pipeline if the current warp is not a memory access type warp;

[0035] A signal acquisition module, which is used to obtain the cache access result signal returned by the L1 Cache based on the decoded instruction if the current warp is a memory access type warp;

[0036] A second execution module, which is used to add the current warp to the active warp table if the cache access result signal meets the preset active condition, and control each warp in the active warp table to enter the execution stage of the instruction pipeline;

[0037] A thread suspension module, which is used to add the current warp to the suspended warp table if the cache access result signal meets the preset suspension condition, update the next warp of the instruction pipeline to a new current warp, and re-jump to the step of controlling the current warp of the instruction pipeline to enter the decoding stage;

[0038] A third execution module, which is used to add the target warp that meets the preset active condition in the suspended warp table to the active warp table when it is monitored, and control each warp in the active warp table to enter the execution stage of the instruction pipeline.

[0039] In a third aspect, the present application discloses an electronic device, including:

[0040] A memory, which is used to store a computer program;

[0041] A processor, which is used to execute the computer program to implement the steps of the GPGPU warp scheduling method disclosed above.

[0042] In a fourth aspect, the present application discloses a computer-readable storage medium, which is used to store a computer program; wherein, when the computer program is executed by a processor, the steps of the GPGPU warp scheduling method disclosed above are implemented.

[0043] The beneficial effects of this application are as follows: This application is applied to a GPGPU and includes: controlling the current warp of the instruction pipeline to enter the decoding stage to obtain decoded instructions, and judging whether the current warp is a memory access type warp according to the decoded instructions; if the current warp is not a memory access type warp, controlling the current warp to enter the execution stage of the instruction pipeline; if the current warp is a memory access type warp, obtaining a cache access result signal returned by the L1 Cache based on the decoded instructions; if the cache access result signal meets a preset active condition, adding the current warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline; if the cache access result signal meets a preset suspension condition, adding the current warp to the suspended warp table, updating the next warp of the instruction pipeline to a new current warp, and re-jumping to the step of controlling the current warp of the instruction pipeline to enter the decoding stage; when it is monitored that there is a target warp in the suspended warp table that meets the preset active condition, adding the target warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline. It can be seen that this application uses the Cache pre-memory access mechanism to implement the scheduling of warps. By identifying the instruction types of each warp, that is, judging whether the current warp is a memory access type warp according to the decoded instructions; performing Cache memory access on the memory access type warp during the warp scheduling stage. If the cache access result signal obtained based on the decoded instructions meets the preset suspension condition, the current warp is suspended, and the next warp of the instruction pipeline is updated to a new current warp, and re-jumped to the step of controlling the current warp of the instruction pipeline to enter the decoding stage, avoiding the memory access type warp from blocking the instruction pipeline for a long time. Further, when the suspended memory access type warp meets the preset active condition, it is activated and scheduled into the instruction pipeline. This application optimizes the warp scheduling strategy through the Cache pre-memory access and suspension mechanism of the memory access type warp, solves the problem that the memory access type warp blocks the instruction pipeline for a long time, hides the memory access latency, improves the efficiency and throughput of the pipeline, and enhances the computing performance of the GPGPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0045] Figure 1Flowchart of a GPGPU warp scheduling method disclosed in this application;

[0046] Figure 2 Specific GPGPU architecture diagram disclosed in this application;

[0047] Figure 3 Schematic diagram of a specific scheduling strategy pipeline disclosed in this application;

[0048] Figure 4 Structural diagram of a specific prefetch mechanism disclosed in this application;

[0049] Figure 5 Schematic diagram of a specific warp scheduling mechanism disclosed in this application;

[0050] Figure 6 Schematic diagram of a specific warp scheduling pipeline disclosed in this application;

[0051] Figure 7 Structural diagram of a GPGPU warp scheduling device disclosed in this application;

[0052] Figure 8 Structural diagram of an electronic device disclosed in this application. Detailed implementation manners

[0053] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] GPGPU is a technology that extends the computing power of a graphics processing unit to general computing tasks. Through this technology, GPGPU is not only used for graphics rendering, but also can efficiently execute various general computing tasks, including scientific computing, data analysis, deep learning, etc.

[0055] GPGPUs use an instruction pipeline to fetch, decode, access memory, perform operations, and write back instructions. Among them, the LSU is an important module in the pipeline operation stage and is used to implement external memory access instructions. In the overall computing process of GPGPUs, the external memory access for computing data severely limits the computing performance of GPGPUs. Compared with the computing process, the external memory access process has a greater delay. When a warp executes an external memory access instruction, the waiting for external data will block the operation of the instruction pipeline and reduce the efficiency of the pipeline. Although the LSU is currently made independent from the instruction pipeline, additional logical resources are still required to maintain the operation of the independent LSU, and the adjacent arithmetic instructions of the memory access instruction still need to wait for the memory access to complete, resulting in little actual improvement in solving the problem of large memory access latency for memory access instructions.

[0056] Therefore, this application correspondingly provides a GPGPU warp scheduling scheme to optimize the GPGPU pipeline scheduling mechanism to improve the pipeline efficiency.

[0057] See Figure 1 As shown in

[0058] Step S11: Control the current warp of the instruction pipeline to enter the decoding stage to obtain the decoded instructions, and determine whether the current warp is a memory access type warp according to the decoded instructions.

[0059] For example Figure 2 A specific GPGPU architecture diagram as shown is divided into two parts: Host (main processor) and GPGPU. Among them, the Host consists of a CPU (Central Processing Unit), memory, and a compiler. The CPU is used to complete the task establishment, allocation, scheduling, and release of the main processor. Programs written by developers on the CPU are translated into hardware instructions that can be parsed by the GPGPU through the compiler.

[0060] As a coprocessor, the GPGPU processes compute-intensive tasks that the CPU is not very good at. Generally, it is typically controlled by PCIE (peripheral component interconnect express, a high-speed serial computer expansion bus standard), thread block scheduling, global memory, L2 Cache, constant memory, and multiple SMs.

[0061] The PCIE control module instantiates the PCIE interface, controls the communication between the CPU and the GPGPU, and establishes a connection between the GPGPU memory and the HostMemory to achieve the transfer of data between the HOST memory and the GPGPU memory. The GPGPU configuration module initializes the GPGPU and configures the control parameters of the GPGPU; the L2 Cache is a cache that prefetches data in the global memory by merging memory accesses to reduce the stalls caused by data memory access in the pipeline and improve the computing efficiency; the thread block scheduling module realizes the scheduling and allocation of thread blocks and allocates the thread blocks to each SM (streaming multiprocessor).

[0062] The GPGPU memory is divided into global memory and constant parameter memory. The global memory stores the data to be processed and the hardware control instructions, and the constant parameter memory is used as a read-only memory during the operation of the GPGPU to store the constant parameters during the operation process.

[0063] The SM is the computing unit of the GPGPU and is internally composed of a pipeline, a SIMT stack, a register file, an L1 Cache, a shared memory, and a warp scheduler. The SITM stack processes the jumps of branch instructions; the register file is used to cache data and intermediate calculation results during the operation process; the shared memory is used for data exchange between warps and between threads; the L1 Cache prefetches data and instructions from the L2 Cache to further reduce the memory access time of the pipeline.

[0064] The instruction pipeline is the key to the SM to achieve the computing function. It completes instruction fetching, decoding, issuing, execution, and result writing back through a five-stage pipeline. The efficiency of the pipeline directly determines the computing performance of the GPGPU. Among them, in the execution stage of the pipeline, several threads in a warp are allocated to each SP (streaming processor) for execution.

[0065] The SP is the smallest computing unit of the GPGPU and completes operations such as the ALU (arithmetic and logic unit), SFU (Special Function Units), and FPU (floating point unit). The number of SPs determines the parallelism of thread execution in a warp. Generally speaking, when the number of SPs is the same as the number of threads in a warp, the parallelism is the highest, and all threads in the warp can be fully parallel in the execution stage.

[0066] The GPGPU realizes instruction fetching, decoding, issuing, operation, and result writing back through an instruction pipeline. After decoding, the GPGPU can identify the type of the current instruction, such as memory access type, operation type, branch type, etc. Usually, the warp scheduling strategy of the GPGPU is polling, greedy, or a strategy combining greedy and resource scheduling. These strategies hardly consider the instruction type of the current warp. However, when the warp type is memory access, when the warp is in the operation (memory access) stage of the pipeline, due to the too long memory access latency, it will block the flow of the pipeline. At the same time, because it occupies the process of the pipeline, it also affects the overall scheduling of the warp, reduces the pipeline efficiency, and affects the computing performance.

[0067] For example Figure 3 As shown in a specific schematic diagram of the scheduling strategy pipeline, under the traditional scheduling strategy, warp1 is a memory access type warp, and warp2 and warp3 are non-memory access type warps. When warp1 accesses the Cache (cache) during the execution stage of the pipeline, if the Cache misses, it needs to access external memory. Usually, this time takes from more than a dozen to dozens of clock cycles. During this period, warp2 needs to wait after completing the issue, and warp3 also needs to wait after completing the decoding, seriously hindering the normal operation of the pipeline and reducing the instruction execution efficiency and throughput.

[0068] For example Figure 4 As shown in a specific structural diagram of the pre-memory access mechanism, the warp scheduling module consists of four parts: a warp management unit, an address calculation unit, a Cache pre-memory access unit, and a warp table. The warp management unit first controls the current warp of the instruction pipeline to enter the decoding stage. After obtaining the decoded instruction, it determines whether the current warp is a memory access type warp according to the decoded instruction.

[0069] Step S12: If the current warp is not a memory access type warp, control the current warp to enter the execution stage of the instruction pipeline.

[0070] If the current warp is not a memory access type warp, it means that the current warp does not need to access the L1 Cache. Then the warp management unit controls the current warp to directly enter the execution stage of the instruction pipeline.

[0071] Step S13: If the current warp is a memory access type warp, obtain the cache access result signal returned by the L1 Cache based on the decoded instruction.

[0072] If the current warp is a memory access type warp, it means that the current warp needs to access the L1 Cache. Then obtain the cache access result signal returned by the L1 Cache based on the decoded instruction.

[0073] In this embodiment, the GPGPU includes a warp management unit, an address calculation unit, and a Cache prefetch unit; obtaining the cache access result signal returned by the L1 Cache based on the decoded instruction includes: controlling the warp management unit to send the decoded instruction to the address calculation unit; controlling the address calculation unit to obtain the current memory access address according to the decoded instruction, and sending the current memory access address to the Cache prefetch unit; controlling the Cache prefetch unit to send a memory access request carrying the current memory access address to the L1 Cache, so as to obtain the cache access result signal returned by the L1 Cache based on the memory access request.

[0074] In the warp scheduling module, control the warp management unit to send the decoded instruction to the address calculation unit; control the address calculation unit to obtain the current memory access address according to the decoded instruction. Specifically, control the address calculation unit to read the source operand from the register file according to the source operand register address in the decoded instruction, and add the source operand to the immediate number in the decoded instruction to obtain the Cache memory access address, that is, obtain the current memory access address, and send the current memory access address to the Cache prefetch unit; the Cache prefetch unit is used for the memory access management of the Cache. After receiving the current memory access address from the address calculation unit, the Cache prefetch unit initiates a memory access request to the L1 Cache. That is to say, control the Cache prefetch unit to send a memory access request carrying the current memory access address to the L1 Cache, so as to obtain the cache access result signal returned by the L1 Cache based on the memory access request.

[0075] In this embodiment, obtaining the cache access result signal returned by the L1 Cache based on the memory access request includes: the L1 Cache matches according to the current memory access address in the memory access request in the cached data of the L1 Cache to generate a cache access result signal; obtaining the cache access result signal returned by the L1 Cache. After receiving the memory access request, the L1 Cache will match according to the current memory access address and the cached data of the L1 Cache to generate a cache access result signal.

[0076] In this embodiment, it further includes: if the cache access result signal is a hit signal, it is determined that the cache access result signal meets the preset active condition; wherein, the hit signal indicates that there is data in the cached data that matches the current memory access address; if the cache access result signal is a miss signal, it is determined that the cache access result signal meets the preset suspension condition; wherein, the miss signal indicates that there is no data in the cached data that matches the current memory access address.

[0077] The warp table is divided into an active warp table and a suspended warp table. The active warp table stores warps that meet the preset active conditions, and the suspended warp table stores warps that meet the preset suspension conditions. Among them, if the cache access result signal is a hit signal, that is, there is data in the cached data that matches the current memory access address, it is determined that the cache access result signal meets the preset active conditions; if the cache access result signal is a miss signal, that is, there is no data in the cached data that matches the current memory access address, it is determined that the cache access result signal meets the preset suspension conditions.

[0078] Step S14: If the cache access result signal meets the preset active conditions, add the current warp to the active warp table and control each warp in the active warp table to enter the execution stage of the instruction pipeline.

[0079] It can be understood that if the cache access result signal meets the preset active conditions, that is, there is data in the cached data that matches the current memory access address, add the current warp to the active warp table and return the matching data to enter the execution stage of the instruction pipeline.

[0080] Step S15: If the cache access result signal meets the preset suspension conditions, add the current warp to the suspended warp table, update the next warp in the instruction pipeline to the new current warp, and then jump back to the step of controlling the current warp in the instruction pipeline to enter the decoding stage.

[0081] If the cache access result signal meets the preset suspension conditions, that is, there is no data in the cached data that matches the current memory access address, add the current warp to the suspended warp table, and the current warp needs to wait to obtain the matching data. Further, update the next warp in the instruction pipeline to the new current warp, and then jump back to the step of controlling the current warp in the instruction pipeline to enter the decoding stage. That is to say, it is necessary to repeat the scheduling steps for the current warp for the next warp to avoid the next warp being blocked in the emission stage. In other words, avoid subsequent warps from having to wait for the current warp to obtain data before continuing to execute. For example, if the current warp is Warp 1 in the suspended stage and the next warp is Warp 2, at this time, Warp 2 does not need to wait for Warp 1 to resume the active state, and directly perform the processing of Warp 1 on Warp 2, that is, determine whether Warp 2 is a memory access type warp according to the decoded instructions of Warp 2, and subsequent corresponding steps.

[0082] In this embodiment, during the process of adding the current warp to the suspended warp table, the following steps are further included: storing the current memory access address in the MSHR register, so that when the MSHR register meets the preset combined memory access condition, target data matching the memory access addresses stored in the MSHR register is fetched from the external memory, and the target data is saved to the L1 Cache to update the cached data in the L1 Cache. If there is no data in the cached data that matches the current memory access address, a miss signal is returned, and the memory access address is written into the MSHR (Miss Status Handling Register) register, waiting for the next combined memory access. That is to say, when the MSHR register meets the preset combined memory access condition, target data matching the memory access addresses stored in the MSHR register is fetched from the external memory. It can be understood that these target data are the data that each warp in the suspended warp table needs to match, and the target data is saved to the L1 Cache to update the cached data in the L1 Cache. In this way, the updated cached data stores the data that each warp in the suspended warp table needs to match, so that the warp can obtain the required data, thereby meeting the preset active condition, and can jump to the step of adding the current warp to the active warp table to enter the execution stage of the instruction pipeline.

[0083] In this embodiment, the following steps are further included: monitoring the space occupied by the access addresses in the MSHR register and the first quantity of the warps corresponding to the access addresses; if the space occupied by the access addresses in the MSHR register is greater than a first preset threshold and / or the first quantity is greater than a second preset threshold, it is determined that the MSHR register meets the preset combined memory access condition. It is necessary to monitor whether the MSHR register meets the preset combined memory access condition. Specifically, the space occupied by the access addresses in the MSHR register and the first quantity of the warps corresponding to the access addresses are monitored. If the space occupied by the access addresses in the MSHR register is greater than the first preset threshold, it is determined that the MSHR register meets the preset combined memory access condition, or if the first quantity is greater than the second preset threshold, it is determined that the MSHR register meets the preset combined memory access condition. That is to say, if there is a large amount of data to be fetched in the MSHR register, the preset combined memory access condition is met. In this embodiment, it is not that every time a warp needs to fetch data, corresponding data is fetched from the external memory, but combined memory access is performed to avoid excessive memory access frequency and prevent too many warps from being stored in the suspended warp table, resulting in too many warps being in the waiting stage.

[0084] In this embodiment, controlling each warp in the active warp table to enter the execution stage of the instruction pipeline includes: determining a second quantity of warps to be executed, where the warps to be executed are the warps that currently need to enter the execution stage of the instruction pipeline; if the second quantity is greater than a third preset threshold, then controlling each warp in the active warp table to enter the execution stage of the instruction pipeline in sequence according to the priorities of the warps to be executed.

[0085] It should be noted that when controlling each warp in the active warp table to enter the execution stage of the instruction pipeline, if there are multiple warps entering the execution stage of the instruction pipeline, that is, the second quantity of warps to be executed is greater than the third preset threshold, it is important to avoid the instruction pipeline delaying the execution of relatively important warps due to too many warps. In this embodiment, the priorities of each warp are also determined. When the quantity of warps to be executed is large, the warps to be executed can enter the execution stage of the instruction pipeline in sequence according to their priorities. For example, currently there are warps to be executed, Warp 1, Warp 5, and Warp 6, where the priorities from high to low are Warp 1, Warp 5, and Warp 6. Then Warp 1 can be preferentially made to enter the execution stage of the instruction pipeline, followed by Warp 5 entering the execution stage of the instruction pipeline, and finally Warp 6 entering the execution stage of the instruction pipeline, avoiding the warps with higher priorities waiting too long to enter the execution stage and achieving a more reasonable scheduling of warps.

[0086] Step S16: When it is monitored that there is a target warp in the suspended warp table that meets the preset active condition, the target warp is added to the active warp table, and each warp in the active warp table is controlled to enter the execution stage of the instruction pipeline.

[0087] It can be understood that, during the process of adding the current thread warp to the suspended thread warp table, the current memory access address is also stored in the MSHR register. So that when the MSHR register meets the preset merged memory access condition, the target data required by each thread warp in the suspended thread warp table can be obtained, and the cached data in the L1 Cache can be updated. That is to say, the cached data in the L1 Cache is continuously updated, so that the cached data stores the data required by each thread warp in the suspended thread warp table, so that each thread warp in the suspended thread warp table can meet the preset active condition, and then the target thread warp can also enter the execution stage of the instruction pipeline. For example, if there are thread warps Warp 1 and Warp 3 in the suspended thread warp table, when it is detected that the data required by Warp 1 exists in the cached data of the L1 Cache, that is, when there is data in the cached data that matches the memory access address of Warp 1, then Warp 1 can be removed from the suspended thread warp table and added to the active thread warp table, so that Warp 1 enters the execution stage of the instruction pipeline.

[0088] In a specific embodiment, when it is detected that there is a target thread warp in the suspended thread warp table that meets the preset active condition, the number of target thread warps is determined. If the number of target thread warps is greater than the fourth preset threshold, then each target thread warp can be divided into a preset number of batches according to the priority of each target thread warp, the transfer level of each batch is determined, based on the preset transfer interval between batches, and in the order from high to low of the transfer level, each batch of target thread warps is transferred from the suspended thread warp table to the active thread warp table in turn. In this way, it is avoided that the number of thread warps in the active thread warp table suddenly increases in a short time, where the preset transfer interval between batches represents the transfer time interval between adjacent batches.

[0089] The beneficial effects of this application are as follows: This application is applied to a GPGPU and includes: controlling the current warp of the instruction pipeline to enter the decoding stage to obtain the decoded instructions, and determining whether the current warp is a memory access type warp according to the decoded instructions; if the current warp is not a memory access type warp, controlling the current warp to enter the execution stage of the instruction pipeline; if the current warp is a memory access type warp, obtaining the cache access result signal returned by the L1 Cache based on the decoded instructions; if the cache access result signal meets the preset active condition, adding the current warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline; if the cache access result signal meets the preset suspension condition, adding the current warp to the suspended warp table, updating the next warp of the instruction pipeline to a new current warp, and re-jumping to the step of controlling the current warp of the instruction pipeline to enter the decoding stage; when it is monitored that there is a target warp in the suspended warp table that meets the preset active condition, adding the target warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline. It can be seen that this application uses the Cache pre-memory access mechanism to implement the scheduling of warps. By identifying the instruction types of each warp, that is, determining whether the current warp is a memory access type warp according to the decoded instructions; performing Cache memory access on the memory access type warps during the warp scheduling stage. If the cache access result signal obtained based on the decoded instructions meets the preset suspension condition, suspending the current warp, updating the next warp of the instruction pipeline to a new current warp, and re-jumping to the step of controlling the current warp of the instruction pipeline to enter the decoding stage, avoiding the memory access type warps from blocking the instruction pipeline for a long time. Further, when the suspended memory access type warps meet the preset active condition, activating and scheduling them into the instruction pipeline. This application optimizes the warp scheduling strategy through the Cache pre-memory access and suspension mechanism of the memory access type warps, solves the problem that the memory access type warps block the instruction pipeline for a long time, hides the memory access latency, improves the efficiency and throughput of the pipeline, and enhances the computing performance of the GPGPU.

[0090] For example Figure 5 As shown in a schematic diagram of a specific warp scheduling mechanism, the thread scheduling module of the GPGPU includes a warp management unit, an address calculation unit, a Cache pre-memory access unit, and a warp table. The mechanism is specifically as follows:

[0091] (1) The thread bundle table maintains two tables, namely the active thread bundle table and the suspended thread bundle table; the active thread bundle table maintains the active thread bundles being scheduled. When the thread bundle completes the pipeline write-back stage, the thread bundle is released from the active thread bundle table and waits for the next scheduling by the thread bundle management unit; the suspended thread bundle table maintains the suspended thread bundles waiting for cache hits. When the suspended thread bundle receives a hit signal, the thread bundle state changes from suspended to active, and is released from the suspended thread bundle table and enters the active thread bundle table.

[0092] (2) The thread bundle management unit identifies the thread bundle instruction type and manages the thread bundle status; when the thread bundle completes decoding, the thread bundle management unit identifies the thread bundle instruction type. If the thread bundle instruction type is a non-memory access instruction, the thread bundle is placed in the active thread bundle table, that is, the current thread bundle is not a memory access type thread bundle, and it is emitted into the pipeline execution stage; if the thread bundle instruction type is a memory access instruction, that is, the current thread bundle is a memory access type thread bundle, the decoded instruction is sent to the address calculation unit to calculate the memory access address and perform cache pre-access, waiting for the cache response. If the cache pre-access unit returns a cache hit signal, the thread bundle is placed in the active thread bundle table and emitted into the pipeline execution stage; if the cache pre-access unit returns a cache miss signal, the thread bundle is suspended and placed in the suspended thread bundle table, and the cache response signal of the thread bundle is continuously detected. When a cache hit is detected, the thread bundle is moved from the suspended thread bundle table to the active thread bundle table and emitted into the pipeline execution stage.

[0093] (3) The address calculation unit is used to calculate the cache memory access address of the memory access type thread bundle instruction. After the thread bundle management unit determines that the current thread bundle is a memory access type thread bundle, it will send the thread bundle's decoded instruction to the address calculation unit. The address calculation unit reads the source operand from the register file according to the source operand register address decoded from the instruction, and superimposes the source operand with the immediate number in the decoded instruction to calculate the cache memory access address, that is, the current memory access address, and sends the memory access address to the cache pre-access unit.

[0094] (4) The cache pre-access unit is used to manage the memory access of the cache. After receiving the memory access address from the address calculation unit, the cache pre-access unit initiates a memory access request to the L1 cache so that after receiving the memory access request, the L1 cache will match it with the data in the cache. If the access data already exists in the cache, a cache hit signal is returned. If not, a cache miss signal is returned and the memory access address is written into the MSHR register, waiting for the next merged memory access. When the cache pre-access signal receives the cache hit or miss signal returned by the L1 cache, it is forwarded to the thread warp management unit.

[0095] For example Figure 6 The diagram shows a specific thread warp scheduling pipeline. Warp 1 is a memory access type thread warp, and the other thread warps are non-memory access type thread warps. When Warp 1 is decoded and identified as a memory access type thread warp, and the cache misses, Warp 1 will be suspended and not scheduled to be emitted, and the next non-memory access type thread warp will be scheduled instead, which is Warp 2 here. After that, Warp 3~Warp 6 are scheduled in sequence. In the 8th cycle, Warp 1 receives a cache hit signal. At this time, Warp 1 is activated and emitted to the execution stage. At this time, Warp 5 also enters the execution stage, but with a lower priority. It needs to wait until Warp 1 is executed before it can be executed. It can be seen that the thread warp scheduling strategy based on cache pre-access can greatly improve the efficiency of the pipeline compared with the traditional strategy, hide the memory access delay of memory access instructions, improve the instruction execution efficiency, and thus enhance computing performance.

[0096] The present application is described below by taking Warp 1 as an example.

[0097] 1) After GPGPU is started, the warp management unit starts the pipeline, Warp 1 begins to fetch instructions and decode, and enters the launch phase;

[0098] 2) The warp management unit determines whether Warp 1 is a memory access warp based on the decoded instructions of Warp 1;

[0099] 3) If Warp 1 is a memory access type warp, the decoded instruction is sent to the address calculation unit;

[0100] 4) The address calculation unit reads the source operand from the register file according to the source operand register address in the decoded instruction, and superimposes the source operand with the immediate value in the decoded instruction to obtain the current memory access address of Warp 1;

[0101] 5) The address calculation unit sends the current memory access address to the cache pre-access unit, and the cache pre-access unit initiates a memory access request to the L1 cache, that is, controls the cache pre-access unit to send the memory access request carrying the current memory access address to the L1 cache;

[0102] 6) After receiving the memory access request, the L1 Cache matches the cached data in the L1 Cache according to the current memory access address. If the match is successful, a hit signal is returned to the cache pre-access unit;

[0103] 7) If the matching is unsuccessful, the memory access address is stored in the MSHR register, waiting for the next combined memory access. A cache miss signal is returned to the cache prefetch unit, and after a successful match at the end of a combined memory access, a cache hit signal for Warp 1 is returned to the cache prefetch unit;

[0104] 8) The cache prefetch unit forwards the received cache access result signal to the warp management unit, that is, forwards the hit signal or the miss signal to the warp management unit;

[0105] 9) If the warp management unit receives a hit signal for Warp 1, the cache access result signal of Warp 1 meets the preset active condition. Warp 1 is placed in the active warp table and launched into the pipeline execution stage, that is, Warp 1 enters the execution stage of the instruction pipeline;

[0106] 10) If the warp management unit receives a miss signal for Warp 1, Warp 1 is suspended and placed in the suspended warp table, and then Warp 2 is scheduled. The cache access result signal of Warp 1 is continuously monitored;

[0107] 11) After Warp 2 reaches the launch stage, it is scheduled according to the above steps 2) to 10);

[0108] 12) When the warp management unit detects a cache hit signal for a certain warp in the suspended warp table, for example, the hit signal for Warp 1, Warp 1 is moved from the suspended warp table to the active warp table and launched into the pipeline execution stage;

[0109] 13) At this time, if another warp, Warp 5, needs to enter the execution stage and the priority of Warp 1 is higher than that of Warp 5, then Warp 5 is temporarily suspended to give priority to the memory access type warp for execution. After the memory access type warp finishes execution, Warp 5 is launched into the pipeline execution stage to continue execution;

[0110] 14) All warps are scheduled according to steps 2) to 13).

[0111] It can be seen that this application optimizes the warp scheduling strategy based on the cache prefetch mechanism. Through the prefetch and suspension of the memory access type warp, the efficient scheduling of warps is realized, the execution efficiency of instructions is improved, the latency of memory access instructions is hidden, and the throughput is increased.

[0112] See Figure 7 As shown, an embodiment of this application discloses a GPGPU warp scheduling device, including: applied to GPGPU, including:

[0113] A type judgment module 11, configured to control the current warp of the instruction pipeline to enter the decoding stage to obtain decoded instructions, and determine whether the current warp is a memory access type warp according to the decoded instructions;

[0114] A first execution module 12, configured to control the current warp to enter the execution stage of the instruction pipeline if the current warp is not a memory access type warp;

[0115] A signal acquisition module 13, configured to obtain a cache access result signal returned by the L1 Cache based on the decoded instructions if the current warp is a memory access type warp;

[0116] A second execution module 14, configured to add the current warp to an active warp table and control each warp in the active warp table to enter the execution stage of the instruction pipeline if the cache access result signal meets a preset active condition;

[0117] A thread suspension module 15, configured to add the current warp to a suspended warp table, update the next warp of the instruction pipeline to a new current warp, and re-jump to the step of controlling the current warp of the instruction pipeline to enter the decoding stage if the cache access result signal meets a preset suspension condition;

[0118] A third execution module 16, configured to add the target warp that meets the preset active condition in the suspended warp table to the active warp table and control each warp in the active warp table to enter the execution stage of the instruction pipeline when it is monitored that there is a target warp that meets the preset active condition in the suspended warp table.

[0119] The beneficial effects of this application are as follows: This application is applied to a GPGPU and includes: controlling the current warp of the instruction pipeline to enter the decoding stage to obtain the decoded instructions, and determining whether the current warp is a memory access type warp according to the decoded instructions; if the current warp is not a memory access type warp, controlling the current warp to enter the execution stage of the instruction pipeline; if the current warp is a memory access type warp, obtaining the cache access result signal returned by the L1 Cache based on the decoded instructions; if the cache access result signal meets the preset active condition, adding the current warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline; if the cache access result signal meets the preset suspension condition, adding the current warp to the suspended warp table, updating the next warp of the instruction pipeline to a new current warp, and re-jumping to the step of controlling the current warp of the instruction pipeline to enter the decoding stage; when it is monitored that there is a target warp in the suspended warp table that meets the preset active condition, adding the target warp to the active warp table and controlling each warp in the active warp table to enter the execution stage of the instruction pipeline. It can be seen that this application uses the Cache pre-memory access mechanism to implement the scheduling of warps. By identifying the instruction types of each warp, that is, determining whether the current warp is a memory access type warp according to the decoded instructions; performing Cache memory access on the memory access type warps during the warp scheduling stage. If the cache access result signal obtained based on the decoded instructions meets the preset suspension condition, the current warp is suspended, and the next warp of the instruction pipeline is updated to a new current warp, and re-jump to the step of controlling the current warp of the instruction pipeline to enter the decoding stage, avoiding the memory access type warps from blocking the instruction pipeline for a long time. Further, when the suspended memory access type warps meet the preset active condition, they are activated and scheduled into the instruction pipeline. This application optimizes the warp scheduling strategy through the Cache pre-memory access and suspension mechanisms of the memory access type warps, solves the problem that the memory access type warps block the instruction pipeline for a long time, hides the memory access latency, improves the efficiency and throughput of the pipeline, and enhances the computing performance of the GPGPU.

[0120] Furthermore, an embodiment of this application also provides an electronic device. Figure 8 It is a structural diagram of the electronic device 20 shown according to an exemplary embodiment, and the content in the figure cannot be considered as any limitation on the scope of use of this application.

[0121] Figure 8A structural schematic diagram of an electronic device provided by an embodiment of the present application. Specifically, it may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the GPGPU thread bundle scheduling method executed by the electronic device disclosed in any of the foregoing embodiments.

[0122] In this embodiment, the power supply 23 is used to provide operating voltages for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed on it here; the input / output interface 25 is used to obtain external input data or output data to the outside, and the specific interface type thereof can be selected according to specific application needs, and no specific limitation is made here.

[0123] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 21 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0124] In addition, as a carrier for resource storage, the memory 22 may be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon include an operating system 221, a computer program 222, and data 223, etc., and the storage method may be temporary storage or permanent storage.

[0125] Among them, the operating system 221 is used to manage and control each hardware device and computer program 222 on the electronic device, so as to implement the operation and processing of the massive data 223 in the memory 22 by the processor 21. It can be Windows, Unix, Linux, etc. In addition to the computer program capable of implementing the GPGPU warp scheduling method executed by the electronic device disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of completing other specific tasks. In addition to the data transmitted by external devices received by the electronic device, the data 223 may also include data collected by its own input / output interface 25, etc.

[0126] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the foregoing disclosed GPGPU warp scheduling method is implemented. For the specific steps of this method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0127] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts between the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and reference can be made to the description in the method part for related parts.

[0128] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application. The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of both. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable EPROM (Erasable Programmable Read Only Memory), electrically erasable programmable EEPROM (Electrically Erasable Programmable read only memory), registers, hard disks, removable disks, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.

[0129] Finally, it should also be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0130] The above has introduced in detail a GPGPU warp scheduling method, device, equipment and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A GPGPU warp scheduling method, characterized in that: Applied to GPGPU, including: Controlling a current thread warp of the instruction pipeline to enter a decoding stage to obtain a decoded instruction, and determining whether the current thread warp is a memory access type thread warp according to the decoded instruction; If the current thread warp is not a memory access type thread warp, controlling the current thread warp to enter the execution phase of the instruction pipeline; If the current thread warp is a memory access type thread warp, obtaining a cache access result signal returned by the L1 Cache based on the decoded instruction; If the cache access result signal satisfies a preset active condition, the current thread warp is added to an active thread warp table, and each thread warp in the active thread warp table is controlled to enter the execution phase of the instruction pipeline; If the cache access result signal satisfies a preset suspension condition, the current thread warp is added to a suspended thread warp table, the next thread warp of the instruction pipeline is updated as a new current thread warp, and the current thread warp of the control instruction pipeline is jumped again to enter a decoding stage; When it is detected that there is a target warp in the suspended warp table that meets the preset active condition, the target warp is added to the active warp table, and each warp in the active warp table is controlled to enter the execution phase of the instruction pipeline.

2. The GPGPU warp scheduling method according to claim 1, characterized in that: The GPGPU includes a thread warp management unit, an address calculation unit and a cache pre-access unit; The obtaining the cache access result signal returned by the L1 Cache based on the decoded instruction includes: Controlling the thread warp management unit to send the decoded instruction to the address calculation unit; Control the address calculation unit to obtain a current memory access address according to the decoded instruction, and send the current memory access address to the Cache pre-access unit; The cache pre-access unit is controlled to send a memory access request carrying the current memory access address to the L1 cache, so as to obtain a cache access result signal returned by the L1 cache based on the memory access request.

3. The GPGPU warp scheduling method according to claim 2, characterized in that: Obtaining a cache access result signal returned by the L1 Cache based on the memory access request, comprising: Matching, by the L1 Cache, the cached data in the L1 Cache according to the current memory access address in the memory access request to generate a cache access result signal; Obtain a cache access result signal returned by the L1 Cache.

4. The GPGPU warp scheduling method according to claim 3, characterized in that: Also includes: If the cache access result signal is a hit signal, determining that the cache access result signal satisfies a preset active condition; wherein the hit signal indicates that there is data in the cached data that matches the current memory access address; If the cache access result signal is a miss signal, it is determined that the cache access result signal meets a preset suspension condition; wherein the miss signal indicates that there is no data matching the current memory access address in the cached data.

5. The GPGPU warp scheduling method according to claim 3, characterized in that: The process of adding the current thread warp to the suspended thread warp table further includes: The current memory access address is stored in the MSHR register, so that when the MSHR register meets the preset merged memory access condition, target data matching each memory access address stored in the MSHR register is obtained from the external memory, and the target data is saved to the L1 Cache to update the cached data in the L1 Cache.

6. The GPGPU warp scheduling method according to claim 5, characterized in that: Also includes: Monitoring the space occupied by the access address in the MSHR register and the first number of thread warps corresponding to the access address; If the space occupied by the access address in the MSHR register is greater than a first preset threshold and / or the first number is greater than a second preset threshold, it is determined that the MSHR register meets a preset merged memory access condition.

7. The GPGPU warp scheduling method according to any one of claims 1 to 6, characterized in that: The controlling each thread warp in the active thread warp table to enter the execution phase of the instruction pipeline includes: Determine a second number of thread warps to be executed; wherein the thread warps to be executed are thread warps that currently need to enter the execution phase of the instruction pipeline; If the second number is greater than a third preset threshold, each warp in the active warp table is controlled in sequence to enter the execution phase of the instruction pipeline according to the priority of the warp to be executed.

8. A GPGPU warp scheduling device, characterized in that: Applied to GPGPU, including: A type determination module, used for controlling the current thread warp of the instruction pipeline to enter the decoding stage to obtain a decoded instruction, and determining whether the current thread warp is a memory access type thread warp according to the decoded instruction; A first execution module, configured to control the current warp to enter the execution phase of the instruction pipeline if the current warp is not a memory access type warp; a signal acquisition module, configured to acquire a cache access result signal returned by the L1 Cache based on the decoded instruction if the current thread warp is a memory access type thread warp; a second execution module, configured to add the current thread warp to an active thread warp table if the cache access result signal satisfies a preset active condition, and control each thread warp in the active thread warp table to enter an execution phase of the instruction pipeline; A thread suspension module, configured to add the current thread bundle to a suspended thread bundle table, update the next thread bundle of the instruction pipeline to a new current thread bundle, and jump to the current thread bundle of the control instruction pipeline to enter a decoding stage again if the cache access result signal satisfies a preset suspension condition; The third execution module is used to add the target thread warp to the active thread warp table when it is detected that there is a target thread warp that meets the preset active condition in the suspended thread warp table, and control each thread warp in the active thread warp table to enter the execution stage of the instruction pipeline.

9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the GPGPU warp scheduling method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein, when the computer program is executed by a processor, the steps of the GPGPU warp scheduling method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Instruction prefetching method and device and artificial intelligence chip

    CN121187652A

  • Memory access request processing method and device for universal graphics processor

    CN121901119A

  • Memory access request processing method and apparatus for general purpose graphics processor

    CN121901119B