Out-of-order scheduling method, system and device for GPGPU instruction pre-analysis and medium

By introducing a pre-analysis stage into the GPGPU pipeline, the instructions are checked and the instructions are transmitted out of order, the pipeline stagnation caused by long delay instructions is solved in the traditional GPGPU architecture, which improves execution efficiency and reduces hardware resource overhead.

CN120029674APending Publication Date: 2025-05-23SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510023109.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When traditional GPGPU architecture encounters long delay instructions, it is easy to cause pipeline stagnation, resulting in reduced execution efficiency, and it is difficult for the existing technology to completely eliminate overall stagnation.

Method used

The out-of-order scheduling method of pre-analysis of GPGPU instruction is adopted. By adding the pre-analysis stage in the initial stage of the pipeline, the instructions are checked for correlation, and the instructions without data correlation are transmitted in an out-of-order to avoid pipeline stagnation caused by long delay instructions.

Benefits of technology

It effectively avoids pipeline stagnation caused by long delay instructions, improves the execution efficiency of GPGPU, reduces hardware resource overhead, and ensures the correctness of instruction execution and the stability of pipeline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029674A_ABST
    Figure CN120029674A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of GPGPU instruction scheduling execution, for example, relates to an out-of-order scheduling method, system and device for GPGPU instruction pre-analysis and a medium. The method comprises the following steps of: an instruction fetching and decoding stage: fetching an instruction to be executed from an instruction cache under the guidance of an instruction fetching scheduler, and storing the instruction into the instruction cache; in the pre-analysis stage, the instructions entering the same thread bundle in sequence in the instruction buffer and the instructions entering the same thread bundle in the sending buffer are subjected to correlation check and are sent out of order, a thread bundle scheduler schedules and switches the thread bundles, the instructions are sent to an execution unit to be executed, and the execution unit executes the instructions. The transmitted instruction information is sent to the scoreboard module; in the execution stage, the transmitted instruction is executed, and an execution result is written back to the register file and fed back to the SIMT stack and the instruction fetching scheduler. The method has the remarkable beneficial effects of reducing hardware resource overhead, improving execution efficiency, reducing pipeline pause and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of GPGPU instruction scheduling and execution, for example, to a GPGPU instruction pre-analysis out-of-order scheduling method and system, device, and medium. Background Art

[0002] Currently, in the GPGPU (general purpose graphics processing unit) architecture, the traditional sequential execution mode easily causes pipeline stalls when encountering long-latency instructions (such as memory access instructions), thereby reducing the execution efficiency of the entire GPU. Although GPGPU uses large-scale multi-threading combined with a scheduler to hide instruction delays, the latency changes of memory access instructions are still inevitable during the execution of complex programs. Many technologies focus on improving thread-level parallelism to deal with GPU stall cycles, such as increasing the number of warps (thread bundles) to increase the probability of finding non-stalled warps and switching, or reordering the priorities of warps to increase parallelism. However, these methods still have limitations because they cannot completely eliminate the overall stall caused by all warps encountering long-latency instructions at the same time.

[0003] In order to more effectively avoid pipeline stalls, borrowing the out-of-order execution technology from the CPU is considered a potential solution. Out-of-order execution allows the processor to continue to issue and execute other ready instructions while waiting for data for certain instructions, thereby improving execution efficiency. However, implementing out-of-order execution in GPGPU faces high hardware resource overhead, especially when dealing with GPGPUs with a large number of registers, reordering loads and stores and register renaming techniques will result in huge resource consumption.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present application, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] The embodiments of the present disclosure provide a GPGPU instruction pre-analysis out-of-order scheduling method, system, device, and medium to solve the problem that traditional methods and technologies still have limitations when dealing with GPGPU pipeline stall problems.

[0007] In some embodiments, the method comprises:

[0008] Instruction fetch and decode stage: under the guidance of the instruction fetch scheduler, the instruction to be executed is fetched from the instruction cache and stored in the instruction buffer; when a branch instruction is encountered, the instruction is fetched according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when a PC instruction jump is performed, the instruction is fetched according to the jump address calculated by the execution unit;

[0009] Pre-analysis stage: Check the correlation between the instructions of the same thread bundle in the instruction buffer and the instructions of the same thread bundle in the sending buffer, and store the new instructions without data correlation in the sending buffer; at the same time, transmit the instruction information that has been issued but not written back from the scoreboard module to the sending buffer module, and check the correlation between the instructions in the sending buffer and the instruction information that has been issued but not written back; for instructions with correlation or special instructions, set control signals at the corresponding conflicting positions in the sending buffer to control the sending; select instructions without correlation for out-of-order sending;

[0010] Emission phase: The warp scheduler schedules and switches warps, sends instructions to the execution unit for execution, and sends the information of the instructions that have been sent to the scoreboard module;

[0011] Execution stage: execute the issued instructions, write the execution results back to the register file, and feed them back to the SIMT stack and instruction fetch scheduler.

[0012] Preferably, in the pre-analysis stage, the special instructions include synchronization instructions, memory access instructions and branch instructions.

[0013] Preferably, the scoreboard module is used to record information of instructions that have been issued but not written back, and supports correlation checking in the pre-analysis stage.

[0014] Preferably, the correlation check in the pre-analysis stage is to perform a register correlation check on each new Warp instruction in the received instruction buffer, specifically in the following manner:

[0015] Compare the source 1, source 2, source 3 register index values ​​and the destination register index value of the new instruction with the destination register index values ​​of all instructions stored in the same thread warp in the current send buffer to see if they are the same; if so, determine that the new instruction is a data conflict instruction;

[0016] Compare the destination register index value of the new instruction with the source 1, source 2, and source 3 register index values ​​of all instructions stored in the current sending buffer to see if they are equal. If they are equal, determine the new instruction as a data conflict instruction;

[0017] Instructions that are determined to be data conflicting are prevented from entering the send buffer.

[0018] Preferably, the instruction determined to be a data conflict is prevented from entering the sending buffer, specifically in the following manner:

[0019] When it is determined that the new instruction is not a data conflict instruction, it will be stored in the sending buffer, and the Valid position of the corresponding sending check list item will be set to 1. The sending buffer contains instructions that have no data correlation with each other. When an instruction detects that there is no scoreboard conflict, that is, the instruction in the sending buffer has no data correlation with the currently executed instruction, the Ready position of the corresponding instruction entry will be set high. If there is a scoreboard conflict, the conflict position of the corresponding instruction entry will be set high, and the Ready bit will not be set high.

[0020] Preferably, the specific manner of out-of-order transmission is as follows: select instructions with both the Ready bit and the Valid bit set high for out-of-order transmission, and then perform out-of-order execution and out-of-order write-back.

[0021] Preferably, the memory consistency model is followed when executing the load instruction; out-of-order transmission of the load instruction and the storage instruction is not allowed;

[0022] For load instructions and store instructions, the load instruction bit and the store instruction bit are set high to control the sending of the instruction;

[0023] For synchronous instructions, before entering the pre-analysis stage from the instruction buffer in sequence, the synchronous instruction position in the send check list item is high, the instruction buffer is stopped, and the instruction is sent to the pre-analysis stage again, until all instructions except the synchronous instruction in the send buffer are executed, and after the synchronous instruction is also executed, the subsequent instruction pre-analysis continues;

[0024] For branch instructions, after entering the pre-analysis stage, the branch instruction position of the corresponding instruction will be high, and the instructions after the branch instruction will not be input into the pre-analysis stage before the branch result is determined. The instructions after the branch will wait for the final result of the branch, and then select the instructions of the corresponding branch path to enter the pre-analysis stage for correlation checking.

[0025] In some embodiments, the system comprises:

[0026] The instruction fetching and decoding module is configured to fetch the instruction to be executed from the instruction cache and store it in the instruction buffer; when encountering a branch instruction, fetch the instruction according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when performing a PC instruction jump, fetch the instruction according to the jump address calculated by the execution unit;

[0027] The emission module is configured as a warp scheduler to schedule and switch warps, emit instructions to the execution unit for execution, and send the emitted instruction information to the scoreboard module;

[0028] Execution module: is configured to execute the issued instructions, write the execution results back to the register file, and feed back to the SIMT stack and instruction fetch scheduler;

[0029] The pre-analysis module is configured to perform dependency checking and out-of-order sending of instructions in the instruction buffer; including:

[0030] An instruction receiving module, used for receiving an instruction buffer including a plurality of Warp instructions;

[0031] A correlation checking module, connected to the instruction receiving module, for performing a register correlation check on each new Warp instruction received;

[0032] A sending buffer module, used for storing non-data conflicting instructions that pass the correlation check, and setting a corresponding Valid bit for each stored instruction;

[0033] The sending check module is used to monitor the scoreboard conflicts between the instructions in the sending buffer and the currently executed instructions. For instructions without scoreboard conflicts, the corresponding Ready bit is set high;

[0034] An out-of-order transmission module, connected to the transmission check module, for selecting instructions with both the Ready bit and the Valid bit set high for out-of-order transmission;

[0035] The execution and write-back module is used to execute the instructions issued out of order and write the results back to the corresponding registers out of order.

[0036] In some embodiments, the device includes: a processor and a memory storing program instructions, and the processor is configured to execute the GPGPU instruction pre-analysis out-of-order scheduling method when running the program instructions.

[0037] In some embodiments, the storage medium stores program instructions, and when the program instructions are run, they execute the GPGPU instruction pre-analysis out-of-order scheduling method.

[0038] The embodiment of the present disclosure provides a GPGPU instruction pre-analysis out-of-order scheduling method, which can achieve the following technical effects:

[0039] By adding a pre-analysis stage to the initial stage of the pipeline, checking the dependencies of instructions and issuing instructions without data dependencies out of order, pipeline stalls caused by long-delay instructions (such as memory access instructions) are effectively avoided, thereby improving the execution efficiency of GPGPU.

[0040] The method of the present invention adopts special operation strategies when encountering special instructions (such as memory access instructions, synchronization instructions, and branch instructions), such as prohibiting the out-of-order sending of load instructions and storage instructions to maintain memory consistency, and pausing the sending of instructions through control bits until the synchronization instructions are executed. These measures further ensure the correctness of instruction execution and the stability of the pipeline.

[0041] Since the method of the present invention avoids complex and resource-intensive operations such as reordering load and store techniques and register renaming techniques, the overhead of GPGPU hardware resources is reduced. Specifically, through the correlation check in the pre-analysis stage, it is ensured that there is no data correlation between the instructions sent, which simplifies the hardware design, reduces the pipeline pause time, and improves the overall execution efficiency.

[0042] In summary, the GPGPU instruction pre-analysis out-of-order scheduling method proposed in the present invention shows significant beneficial effects in reducing hardware resource overhead, improving execution efficiency, reducing pipeline pauses, etc., and provides a new and more efficient solution for GPGPU instruction scheduling and execution.

[0043] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0045] Figure 1 It is a schematic flow chart of the method of the present invention;

[0046] Figure 2 This is a schematic diagram of the pre-analysis pipeline architecture of the present invention;

[0047] Figure 3 This is a schematic diagram of the pre-analysis function of the present invention;

[0048] Figure 4 It is a schematic diagram of a device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0049] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0050] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0051] Unless otherwise stated, the term "plurality" means two or more.

[0052] like Figure 1 As shown, a GPGPU instruction pre-analysis out-of-order scheduling method includes:

[0053] Instruction fetch and decode stage: under the guidance of the instruction fetch scheduler, the instruction to be executed is fetched from the instruction cache and stored in the instruction buffer; when a branch instruction is encountered, the instruction is fetched according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when a PC instruction jump is performed, the instruction is fetched according to the jump address calculated by the execution unit;

[0054] Pre-analysis stage: Check the correlation between the instructions of the same thread bundle in the instruction buffer and the instructions of the same thread bundle in the sending buffer, and store the new instructions without data correlation in the sending buffer; at the same time, transmit the instruction information that has been issued but not written back from the scoreboard module to the sending buffer module, and check the correlation between the instructions in the sending buffer and the instruction information that has been issued but not written back; for instructions with correlation or special instructions, set control signals at the corresponding conflicting positions in the sending buffer to control the sending; select instructions without correlation for out-of-order sending;

[0055] Emission phase: Schedule and switch the thread bundles through the thread bundle scheduler, send the instructions to the execution unit for execution, and send the issued instruction information to the scoreboard module;

[0056] Execution stage: execute the issued instructions, write the execution results back to the register file, and feed them back to the SIMT stack and instruction fetch scheduler.

[0057] As a refinement of the above embodiment, in the pre-analysis stage, the special instructions include synchronization instructions, memory access instructions and branch instructions.

[0058] As a refinement of the above embodiment, the scoreboard module is used to record information of instructions that have been issued but not written back, and supports dependency checking in the pre-analysis stage.

[0059] As a refinement of the above embodiment, the correlation check in the pre-analysis stage is to perform a register correlation check on each new Warp instruction in the received instruction buffer, and the specific method is as follows:

[0060] Compare the source 1, source 2, source 3 register index values ​​and the destination register index value of the new instruction with the destination register index values ​​of all instructions stored in the same thread warp in the current send buffer to see if they are the same; if so, determine that the new instruction is a data conflict instruction;

[0061] Compare the destination register index value of the new instruction with the source 1, source 2, and source 3 register index values ​​of all instructions stored in the current sending buffer to see if they are equal. If they are equal, determine the new instruction as a data conflict instruction;

[0062] Instructions that are determined to be data conflicting are prevented from entering the send buffer.

[0063] By adding a pre-analysis stage to the initial stage of the pipeline, checking the dependencies of instructions and issuing instructions without data dependencies out of order, pipeline stalls caused by long-delay instructions (such as memory access instructions) are effectively avoided, thereby improving the execution efficiency of GPGPU.

[0064] As a refinement of the above embodiment, the instruction determined to be a data conflict is prevented from entering the sending buffer, and the specific method is as follows:

[0065] When it is determined that the new instruction is not a data conflict instruction, it will be stored in the sending buffer, and the Valid position of the corresponding sending check list item will be set to 1. The sending buffer contains instructions that have no data correlation with each other. When an instruction detects that there is no scoreboard conflict, that is, the instruction in the sending buffer has no data correlation with the currently executed instruction, the Ready position of the corresponding instruction entry will be set high. If there is a scoreboard conflict, the conflict position of the corresponding instruction entry will be set high, and the Ready bit will not be set high.

[0066] As a refinement of the above embodiment, the specific method of out-of-order transmission is as follows: select instructions with both the Ready bit and the Valid bit set high for out-of-order transmission, and then perform out-of-order execution and out-of-order write-back. Because the instructions sent are not related, there is no need to perform reordering operations and register renaming operations, which greatly reduces resource consumption and reduces pipeline pauses.

[0067] As a refinement of the above embodiment, the memory consistency model is followed when executing the load instruction; out-of-order transmission between the load instruction and the storage instruction is not allowed;

[0068] For load instructions and store instructions, the load instruction bit and the store instruction bit are set high to control the sending of the instruction;

[0069] For synchronous instructions, before entering the pre-analysis stage from the instruction buffer in sequence, the synchronous instruction position in the send check list item is high, the instruction buffer is stopped, and the instruction is sent to the pre-analysis stage again, until all instructions except the synchronous instruction in the send buffer are executed, and after the synchronous instruction is also executed, the subsequent instruction pre-analysis continues;

[0070] For branch instructions, after entering the pre-analysis stage, the branch instruction position of the corresponding instruction will be high, and the instructions after the branch instruction will not be input into the pre-analysis stage before the branch result is determined. The instructions after the branch will wait for the final result of the branch, and then select the instructions of the corresponding branch path to enter the pre-analysis stage for correlation checking.

[0071] The method of the present invention adopts special operation strategies when encountering special instructions, such as prohibiting the out-of-order sending of load instructions and storage instructions to maintain memory consistency, and pausing the sending of instructions through control bits until the synchronization instructions are executed. These measures further ensure the correctness of instruction execution and the stability of the pipeline.

[0072] A GPGPU instruction pre-analysis out-of-order scheduling system, comprising:

[0073] The instruction fetching and decoding module is configured to fetch the instruction to be executed from the instruction cache and store it in the instruction buffer; when encountering a branch instruction, fetch the instruction according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when performing a PC instruction jump, fetch the instruction according to the jump address calculated by the execution unit;

[0074] Transmitter module: configured as a warp scheduler to schedule and switch warps, transmit instructions to execution units for execution, and send transmitted instruction information to the scoreboard module;

[0075] Execution module: is configured to execute the issued instructions, write the execution results back to the register file, and feed back to the SIMT stack and instruction fetch scheduler;

[0076] The pre-analysis module is configured to perform dependency checking and out-of-order issuance on the instructions in the instruction buffer.

[0077] As a refinement of the above embodiment, the analysis module includes:

[0078] An instruction receiving module, used for receiving an instruction buffer including a plurality of Warp instructions;

[0079] A correlation checking module, connected to the instruction receiving module, for performing a register correlation check on each new Warp instruction received;

[0080] A sending buffer module, used for storing non-data conflicting instructions that pass the correlation check, and setting a corresponding Valid bit for each stored instruction;

[0081] The sending check module is used to monitor the scoreboard conflicts between the instructions in the sending buffer and the currently executed instructions. For instructions without scoreboard conflicts, the corresponding Ready bit is set high;

[0082] An out-of-order transmission module, connected to the transmission check module, for selecting instructions with both the Ready bit and the Valid bit set high for out-of-order transmission;

[0083] The execution and write-back module is used to execute the instructions issued out of order and write the results back to the corresponding registers out of order.

[0084] Combination Figure 4 As shown, the embodiment of the present disclosure provides a GPGPU instruction pre-analysis out-of-order scheduling device 300, including a processor (processor) 304 and a memory (memory) 301. Optionally, the device may also include a communication interface (Communication Interface) 302 and a bus 303. Among them, the processor 304, the communication interface 302, and the memory 301 can communicate with each other through the bus 303. The communication interface 302 can be used for information transmission. The processor 304 can call the logic instructions in the memory 301 to execute the GPGPU instruction pre-analysis out-of-order scheduling method of the above embodiment.

[0085] In addition, the logic instructions in the memory 301 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0086] The memory 301 is a computer-readable storage medium that can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 304 executes the functional application and data processing by running the program instructions / modules stored in the memory 301, that is, the out-of-order scheduling method for GPGPU instruction pre-analysis in the above embodiment is implemented.

[0087] The memory 301 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 301 may include a high-speed random access memory and may also include a non-volatile memory.

[0088] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned GPGPU instruction pre-analysis out-of-order scheduling method.

[0089] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0090] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, or a transient storage medium.

[0091] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent only possible changes. Unless explicitly required, separate components and functions are optional, and the order of operation may vary. The parts and features of some embodiments may be included in or replace the parts and features of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or groups thereof. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0092] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. The technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. The technicians may clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here.

[0093] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0094] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A GPGPU instruction pre-analysis out-of-order scheduling method, characterized in that: include: Instruction fetch and decode stage: under the guidance of the instruction fetch scheduler, the instruction to be executed is fetched from the instruction cache and stored in the instruction buffer; When a branch instruction is encountered, the instruction is fetched according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when a PC instruction jump is performed, the instruction is fetched according to the jump address calculated by the execution unit; Pre-analysis stage: Check the correlation between the instructions of the same thread bundle in the instruction buffer and the instructions of the same thread bundle in the sending buffer, and store the new instructions without data correlation in the sending buffer; at the same time, transmit the instruction information that has been issued but not written back from the scoreboard module to the sending buffer module, and check the correlation between the instructions in the sending buffer and the instruction information that has been issued but not written back; for instructions with correlation or special instructions, set control signals at the corresponding conflicting positions in the sending buffer to control the sending; select instructions without correlation for out-of-order sending; Emission phase: The warp scheduler schedules and switches warps, sends instructions to the execution unit for execution, and sends the information of the instructions that have been sent to the scoreboard module; Execution stage: execute the issued instructions, write the execution results back to the register file, and feed them back to the SIMT stack and instruction fetch scheduler.

2. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 1, characterized in that: The special instructions in the pre-analysis stage include synchronization instructions, memory access instructions and branch instructions.

3. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 1, characterized in that: The scoreboard module is used to record the information of instructions that have been issued but not written back, and supports the correlation check in the pre-analysis stage.

4. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 1, characterized in that: The pre-analysis phase correlation check is to perform a register correlation check on each new Warp instruction in the received instruction buffer. The specific method is as follows: Compare the source 1, source 2, source 3 register index values ​​and the destination register index value of the new instruction with the destination register index values ​​of all instructions stored in the same thread warp in the current send buffer to see if they are the same; if so, determine that the new instruction is a data conflict instruction; Compare the destination register index value of the new instruction with the source 1, source 2, and source 3 register index values ​​of all instructions stored in the current sending buffer to see if they are equal. If they are equal, determine the new instruction as a data conflict instruction; Instructions that are determined to be data conflicting are prevented from entering the send buffer.

5. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 4, characterized in that: The instructions that are determined to be data conflicting are prevented from entering the send buffer. The specific methods are as follows: When it is determined that the new instruction is not a data conflict instruction, it will be stored in the sending buffer, and the Valid position of the corresponding sending check list item will be set to 1. The sending buffer contains instructions that have no data correlation with each other. When an instruction detects that there is no scoreboard conflict, that is, the instruction in the sending buffer has no data correlation with the currently executed instruction, the Ready position of the corresponding instruction entry will be set high. If there is a scoreboard conflict, the conflict position of the corresponding instruction entry will be set high, and the Ready bit will not be set high.

6. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 5, characterized in that: The specific method of out-of-order transmission is as follows: select instructions with both the Ready bit and the Valid bit set high for out-of-order transmission, and then perform out-of-order execution and out-of-order write back.

7. The out-of-order scheduling method for GPGPU instruction pre-analysis according to claim 6, characterized in that: The memory consistency model is followed when executing load instructions; out-of-order transmission of load instructions and storage instructions is not allowed; For load instructions and store instructions, the load instruction bit and the store instruction bit are set high to control the sending of the instruction; For synchronous instructions, before entering the pre-analysis stage from the instruction buffer in sequence, the synchronous instruction position in the send check list item is high, the instruction buffer is stopped, and the instruction is sent to the pre-analysis stage again, until all instructions except the synchronous instruction in the send buffer are executed, and after the synchronous instruction is also executed, the subsequent instruction pre-analysis continues; For branch instructions, after entering the pre-analysis stage, the branch instruction position of the corresponding instruction will be high, and the instructions after the branch instruction will not be input into the pre-analysis stage before the branch result is determined. The instructions after the branch will wait for the final result of the branch, and then select the instructions of the corresponding branch path to enter the pre-analysis stage for correlation checking.

8. An out-of-order scheduling system for GPGPU instruction pre-analysis that executes the method of any one of claims 1-7, characterized in that: include: An instruction fetch and decode module is configured to fetch instructions to be executed from the instruction cache and store them in the instruction buffer; When a branch instruction is encountered, the instruction is fetched according to the PC value of the branch to be executed popped out by the SIMT stack management unit; when a PC instruction jump is performed, the instruction is fetched according to the jump address calculated by the execution unit; The emission module is configured as a warp scheduler to schedule and switch warps, emit instructions to the execution unit for execution, and send the emitted instruction information to the scoreboard module; Execution module: is configured to execute the issued instructions, write the execution results back to the register file, and feed back to the SIMT stack and instruction fetch scheduler; The pre-analysis module is configured to perform dependency checking and out-of-order sending of instructions in the instruction buffer; including: An instruction receiving module, used for receiving an instruction buffer including a plurality of Warp instructions; A correlation checking module, connected to the instruction receiving module, for performing a register correlation check on each new Warp instruction received; A sending buffer module, used for storing non-data conflicting instructions that pass the correlation check, and setting a corresponding Valid bit for each stored instruction; The sending check module is used to monitor the scoreboard conflicts between the instructions in the sending buffer and the currently executed instructions. For instructions without scoreboard conflicts, the corresponding Ready bit is set high; An out-of-order transmission module, connected to the transmission check module, for selecting instructions with both the Ready bit and the Valid bit set high for out-of-order transmission; The execution and write-back module is used to execute the instructions issued out of order and write the results back to the corresponding registers out of order.

9. A GPGPU instruction pre-analysis out-of-order scheduling device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the out-of-order scheduling method for GPGPU instruction pre-analysis as described in any one of claims 1 to 7 when running the program instructions.

10. A storage medium storing program instructions, characterized in that: When the program instructions are run, the out-of-order scheduling method for pre-analysis of GPGPU instructions as described in any one of claims 1 to 7 is executed.