Instruction scheduling system and method and electronic equipment
By building a dependency linked list in a superscalar out-of-order processor, the instruction scheduling sequence is optimized, and the storage space waste and delay caused by read-write dependencies is solved, and the processor's instruction scheduling efficiency and execution efficiency are improved.
Patent Information
- Application Number
- CN202510386712.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
When a superscalar out-of-order processor has a write-after-read dependency (RAW dependency) during instruction scheduling, it leads to waste of reserved station storage space and delayed instruction execution.
The dispatch module, dependency inspection module and transmission module are used to write the dispatch queue by combining the dispatch instructions to be dispatched and the renamed register index, perform dependency inspection, build a dependency linked list, and send instructions to the transmission queue in the order of the dependency linked list characterization, and finally send them to the execution unit.
Improve the processing efficiency of RAW dependencies in superscalar out-of-order processors, reduce instruction waiting time, avoid waste of pipeline resources, and shorten instruction execution delay.
Smart Images

Figure CN120276771A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an instruction scheduling system, method and electronic device. Background Art
[0002] A superscalar out-of-order processor can issue multiple instructions to different execution units in the same clock cycle. These execution units can parallelly process various instructions such as integer operations, floating-point operations, and memory accesses, and can also dynamically rearrange the instruction execution order according to the data dependencies and resource availabilities among the instructions.
[0003] In the related art, during the instruction scheduling process of the current superscalar out-of-order processor, the order of instruction scheduling is usually determined according to the waiting time of the instructions in the reservation station. However, when there is a read after write (RAW) dependency between instructions, that is, the result of the previous instruction is the operand of the next instruction, and when the execution result of the previous instruction is not available, the instruction can only wait until the operand is available before it can be executed, which not only wastes the storage space of the reservation station but also increases the execution delay of the instructions. Summary of the Invention
[0004] This application provides an instruction scheduling system, method and electronic device to at least solve the problem in the related art that not only wastes the storage space of the reservation station but also increases the execution delay of the instructions.
[0005] This application provides an instruction scheduling system, including: a dispatching module, a dependency checking module, and an issuing module;
[0006] The dispatching module is configured to, for each obtained to-be-scheduled instruction and the renaming register index corresponding to the to-be-scheduled instruction, use the combined field of the to-be-scheduled instruction and the renaming register index as a dispatch table entry and write it into the corresponding dispatch queue;
[0007] The dependency checking module is configured to perform dependency checking on the dispatch table entries in the dispatch queue according to a preset dependency checking period, construct a dependency linked list according to the dependency checking results, and send the to-be-scheduled instructions to the corresponding issue queue in the order of instruction execution represented by the dependency linked list;
[0008] The issuing module is configured to send the to-be-scheduled instructions in the issue queue to the corresponding execution unit and receive the instruction execution results returned by the execution unit.
[0009] This application also provides an instruction scheduling method, including:
[0010] Obtain any to-be-scheduled instruction and the renaming register index corresponding to the to-be-scheduled instruction;
[0011] Use the combined field of the instruction to be scheduled and the renamed register index as a dispatch table entry and write it into the corresponding dispatch queue;
[0012] Perform dependency checks on the dispatch table entries in the dispatch queue according to a preset dependency check period;
[0013] Construct a dependency linked list based on the dependency check results;
[0014] Send the instructions to be scheduled to the corresponding issue queue in the order of instruction execution represented by the dependency linked list;
[0015] Send the instructions to be scheduled in the issue queue to the corresponding execution unit and receive the instruction execution results returned by the execution unit.
[0016] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above instruction scheduling methods when executing the computer program.
[0017] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of any of the above instruction scheduling methods when executed by a processor.
[0018] This application also provides a computer program product including a computer program, and the computer program implements the steps of any of the above instruction scheduling methods when executed by a processor.
[0019] With this application, since the dispatch module writes the combination of the instruction to be scheduled and the renamed register index into the dispatch queue, and the dependency check module constructs a dependency linked list based on the dispatch queue to determine the instruction execution order and send instructions accordingly, ensuring that the sending order of instructions with RAW dependencies matches the execution order, it is possible to solve the technical problems of low RAW dependency processing efficiency, long instruction waiting time, and waste of pipeline instruction storage resources in superscalar out-of-order processor instruction scheduling, and reduce the instruction execution delay. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] To illustrate the embodiments of the present application more clearly, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0021] Figure 1 It is a schematic interaction flow diagram of the instruction scheduling system provided by the embodiment of the present application;
[0022] Figure 2 It is a schematic overall structure diagram of the instruction scheduling system provided by the embodiment of the present application;
[0023] Figure 3 Schematic diagram of the dispatch module provided by the embodiment of the present application;
[0024] Figure 4 Schematic diagram of the structure of an exemplary dependency list provided by the embodiment of the present application;
[0025] Figure 5 Schematic diagram of the operation process of instruction routing provided by the embodiment of the present application;
[0026] Figure 6 Schematic diagram of the structure of the emission module provided by the embodiment of the present application;
[0027] Figure 7 Schematic diagram of the operation process of the buffer unit provided by the embodiment of the present application;
[0028] Figure 8 Schematic diagram of the process of the instruction scheduling method provided by the embodiment of the present application;
[0029] Figure 9 Schematic diagram of the structure of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0031] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a specific order or sequence.
[0032] Most modern processors adopt the design of superscalar out-of-order processors. The main feature of this processor architecture is the ability to execute multiple instructions simultaneously and dynamically reorder these instructions during execution to maximize the utilization of processor resources and performance. A superscalar processor allows multiple instructions to be issued simultaneously to different execution units within one clock cycle. These execution units can execute different types of instructions in parallel, thus accelerating the overall instruction processing ability. Out-of-order execution is a key feature of superscalar processors, which allows the processor to dynamically rearrange the execution order of instructions based on data dependencies and resource availability between instructions. The superscalar out-of-order technology enables the processor to effectively increase the instruction parallelism and execution efficiency without changing the program semantics.
[0033] In the related art, the basic idea of instruction scheduling for superscalar out-of-order processors is that the reservation station provides the register renaming function. While caching the instructions waiting to be sent, the reservation station also caches the instruction operands waiting to be issued. The instructions waiting to be executed specify the reservation station to provide inputs for themselves. Currently, during the instruction scheduling process of superscalar out-of-order processors, the order of instruction scheduling is usually determined according to the waiting time of the instructions in the reservation station. However, when there is a read after write (RAW) dependency, that is, the result of the previous instruction is the operand of the next instruction, the instruction can only wait in the reservation station when the execution result of the previous instruction is not available. This design overly relies on the judgment of operand availability. The instructions in the reservation station can only wait before the judgment is passed. Only after the operand is judged to be available can the next step be taken, wasting the storage space of the reservation station and delaying the execution of the instructions.
[0034] Embodiments of the present application are provided to solve the above technical problems, and provide an instruction scheduling system, method and electronic device. The system includes: a dispatching module, a dependency check module, and an issuing module; the dispatching module is configured to, for each obtained to-be-scheduled instruction and the renamed register index corresponding to the to-be-scheduled instruction, use the combined field of the to-be-scheduled instruction and the renamed register index as a dispatch entry, and write it into the corresponding dispatch queue; the dependency check module is configured to perform dependency checks on the dispatch entries in the dispatch queue according to a preset dependency check period, construct a dependency list according to the dependency check results, and send the to-be-scheduled instructions to the corresponding issue queue in the order of instruction execution represented by the dependency list; the issuing module is configured to send the to-be-scheduled instructions in the issue queue to the corresponding execution unit and receive the instruction execution results returned by the execution unit. The method provided by the above solution improves the RAW dependency processing efficiency in the instruction scheduling of superscalar out-of-order processors, shortens the waiting time of instructions, avoids wasting the pipeline instruction storage resources (the storage space of the reservation station), and reduces the instruction execution delay by ensuring that the sending order and execution order of instructions with RAW dependencies match.
[0035] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] Embodiments of the present application provide an instruction scheduling system for performing instruction scheduling for a superscalar out-of-order processor.
[0037] As Figure 1 shown, it is a schematic diagram of the interaction process of the instruction scheduling system provided by embodiments of the present application. The system includes: a dispatching module, a dependency check module, and an issuing module.
[0038] Among them, the dispatching module is configured to, for each obtained to-be-scheduled instruction and the renamed register index corresponding to the to-be-scheduled instruction, use the combined field of the to-be-scheduled instruction and the renamed register index as a dispatch entry, and write it into the corresponding dispatch queue; the dependency check module is configured to perform dependency checks on the dispatch entries in the dispatch queue according to a preset dependency check period, construct a dependency list according to the dependency check results, and send the to-be-scheduled instructions to the corresponding issue queue in the order of instruction execution represented by the dependency list; the issuing module is configured to send the to-be-scheduled instructions in the issue queue to the corresponding execution unit and receive the instruction execution results returned by the execution unit.
[0039] It should be noted that the to-be-scheduled instructions obtained by the dispatch module are microinstructions sent by a multi-way decoder. The multi-way decoder can be an 8-way decoder or a 16-way decoder, etc. The multi-way decoder receives the original instructions sent from the host side and decodes one original instruction into multiple microinstructions. The instruction scheduling system provided by the embodiments of the present application can perform parallel scheduling processing on multiple microinstructions (to-be-scheduled instructions) within one cycle to improve the instruction scheduling efficiency. Based on the physical register index (renamed register index) generated by the rename module for renaming the logical register index in the microinstruction, the renamed register index corresponding to the to-be-scheduled instruction is obtained.
[0040] Specifically, after the dispatch module writes the combined field of the to-be-scheduled instruction and the renamed register index as a dispatch table entry into the corresponding dispatch queue, the dependency check module performs a dependency check on all existing dispatch table entries in the dispatch queue according to a preset dependency check period. During the check process, the dependency relationships between instructions are analyzed, such as which instruction execution results are the source operands required for the execution of other instructions, etc., so as to construct a dependency linked list. The dependency linked list clarifies the execution order of instructions. The dependency check module sends the to-be-scheduled instructions to the corresponding issue queue in turn according to this order. Finally, the issue module sends the to-be-scheduled instructions in the issue queue to the corresponding execution unit.
[0041] Specifically, in one embodiment, the dispatch module is used to obtain the operation code of the to-be-scheduled instruction; determine the instruction type of the to-be-scheduled instruction according to the operation code of the to-be-scheduled instruction; and write the combined field of the to-be-scheduled instruction and the renamed register index as a dispatch table entry into the corresponding dispatch queue according to the instruction type of the to-be-scheduled instruction.
[0042] Among them, the dispatch queue corresponds to the instruction type one by one.
[0043] It should be noted that the instruction types are at least divided into memory access instructions, vector instructions, fixed-point instructions, and floating-point instructions. The dispatch queues correspondingly include a memory access dispatch queue, a vector dispatch queue, a fixed-point dispatch queue, and a floating-point dispatch queue.
[0044] Specifically, the dispatch module can determine the instruction type of the to-be-scheduled instruction by parsing the operation code carried in the to-be-scheduled instruction, and then write the dispatch table entry corresponding to the to-be-scheduled instruction into the corresponding dispatch queue according to its instruction type.
[0045] Exemplarily, such as Figure 2As shown in the figure, it is a schematic diagram of the overall structure of the instruction scheduling system provided by the embodiment of the present application. After receiving the instruction to be scheduled and the renamed register index, the dispatch module first analyzes its instruction category, judges the status of the dispatch queue, and judges the status of the Reorder Buffer (ROB for short). When the status of the dispatch queue and the status of the ROB indicate that there are empty entries in the dispatch queue corresponding to the instruction and there are also allocable empty entries in the ROB, the dispatch entry of the instruction is written into the corresponding dispatch queue. When writing the dispatch entry into the dispatch queue, the corresponding entry flag bit in the BusyTable (resource status record table) is also set to pre-occupy the position of the instruction in the physical destination register. After judging the instruction category, if it is a memory access instruction, it is sent to the LSQ (load store queue); if it is of other categories, such as vector instructions, scalar instructions, or floating-point instructions, it is sent to the corresponding dispatch queue. The dependency check module will redistribute the dispatch entries in the dispatch queue, and distribute them to the vector issue queue, atomic issue queue, scalar issue queue, cryptographic issue queue, floating-point issue queue, etc. according to the instruction category.
[0046] Furthermore, as Figure 3 shown in the figure, it is a schematic diagram of the structure of the dispatch module provided by the embodiment of the present application. After the instruction enters the dispatch module, it first judges the instruction type according to the opcode of the instruction, then obtains the corresponding dispatch queue address according to the judgment result, and allocates the instruction to the corresponding dispatch queue; the dispatch module also obtains the physical destination register index in the renamed register index, allocates it to the resource status record table, and sets the corresponding status bit. In addition, before the instruction is allocated to the corresponding dispatch queue, it will also be combined with the renamed register index. The combined field of the instruction to be scheduled and the renamed register index enters the dispatch queue as a dispatch entry.
[0047] Specifically, in one embodiment, the dependency check module is used to perform dependency checks on each dispatch entry in any dispatch queue according to the composition fields of each dispatch entry in the dispatch queue, and determine the instruction execution order, source register, and target register of each instruction to be scheduled in the dispatch queue; according to the instruction execution order, source register, and target register of each instruction to be scheduled in each dispatch queue, a dependency linked list is constructed.
[0048] Among them, the dependency check result includes the instruction execution order, source register, and target register of each instruction to be scheduled in each dispatch queue; the dependency linked list uses the instruction to be scheduled as a node, and each node is connected according to the instruction execution order. Each node is connected through the source register and the destination register. The source register is used as the input end of the node, and the destination register is used as the output end of the node.
[0049] It should be noted that the dependencies to be solved in the embodiments of the present application mainly refer to RAW dependencies. Since there is no RAW dependency for memory access instructions, the dependency check module mainly performs dependency checks on vector instructions, fixed-point instructions, and floating-point instructions. Vector instructions are used to request a superscalar out-of-order processor to perform vector calculations, fixed-point instructions are used to request a superscalar out-of-order processor to perform fixed-point calculations, and floating-point instructions are used to request a superscalar out-of-order processor to perform floating-point calculations.
[0050] Among them, the dependency check module is also used to send the dispatch table entries in the memory access dispatch queue to the load / store queue.
[0051] Among them, the composition fields of the dispatch table entries are shown in Table 1 below:
[0052] Table 1
[0053]
[0054] Specifically, the source register of the instruction to be scheduled can be determined according to Rj and Rk, the destination register of the instruction to be scheduled can be determined according to Ds, and the instruction execution order can be determined according to the PC. The larger the PC value, the later the instruction execution order.
[0055] Specifically, first check whether there are the same source registers or destination registers in the fields of the dispatch table entries of the current dispatch queues where each micro-instruction is located. For example, x1 is one of the source operand registers or destination registers of instruction A, and x1 is also one of the source operand registers or destination registers of instruction B. If there are instructions with the same source operand register or destination register, further determine the field positions where these instructions' same registers are located. If there is a register that is both the source operand register of one instruction and the destination register of another instruction, then these two instructions are very likely to be adjacent two instructions in the same RAW chain (dependency linked list). However, sometimes there are more than two instructions that meet the above conditions, and there may be a phenomenon where the same register repeatedly appears in a RAW chain. Therefore, the instruction order in the RAW chain can be further determined according to the PC values of the instructions.
[0056] Exemplarily, as Figure 4 shown, it is a schematic structural diagram of an exemplary dependency linked list provided by the embodiments of the present application. This linked list uses instructions as nodes, and each node is connected by source registers and target registers, indicating the RAW dependency relationship of the dispatched instructions in the current cycle. The target register (Rd) of instruction 1 is the source register (Rs1) of instruction 2. After the dependency check module constructs all RAW chains in the current preset dependency check cycle, it also determines the instruction dispatch order in the RAW chain.
[0057] It should be further noted that in the embodiments of the present application, by constructing a dependency linked list, the implicit register dependencies are converted into an intuitive linked list structure, making the data flow between instructions traceable, and determining the instruction execution order in advance in the dependency check module, reducing the dynamic scheduling overhead in the issue queue.
[0058] Specifically, in one embodiment, the dependency check module is further configured to, for any instruction to be scheduled, determine the linked list number of the instruction to be scheduled and the dependency sequence number in the dependency linked list according to the dependency linked list of the instruction to be scheduled; add the linked list number of the instruction to be scheduled and the dependency sequence number in the dependency linked list to the dispatch table entry corresponding to the instruction to be scheduled.
[0059] Among them, the linked list number and the dependency sequence number are used to represent the execution order of instructions. Each dependency linked list constructed by the dependency check module has its own unique number (linked list number), and the instructions in the dependency linked list also have their own sequence numbers in the dependency linked list (dependency sequence numbers), and these numbers can be added to the table entry fields in the dispatch queue.
[0060] Specifically, in one embodiment, the dependency check module is configured to determine the issue priority of each instruction to be scheduled according to the waiting duration of the instruction represented by the dispatch table entry; among them, the instruction waiting duration and the issue priority are positively correlated; according to the execution order of instructions represented by the dependency linked list, the issue priority of each instruction to be scheduled, and the instruction type of each instruction to be scheduled, send each instruction to be scheduled to the corresponding issue queue in sequence.
[0061] Among them, each instruction type corresponds to at least one issue queue, and multiple instructions to be scheduled of the same instruction type belonging to the same dependency linked list will be sent to the same issue queue. Among them, fixed-point instructions can be further divided into atomic instructions, scalar instructions, and cryptographic instructions. Atomic instructions are used to request atomic operation operations, scalar instructions are used to request scalar operation operations, and cryptographic instructions are used to request cryptographic operation operations. As Figure 2 shown, the issue queues include a vector issue queue, an atomic issue queue, a scalar issue queue, a cryptographic issue queue, and a floating-point issue queue.
[0062] Specifically, the dependency check module includes an instruction router, as Figure 5As shown in the figure, it is a schematic diagram of the operation process of instruction routing provided by an embodiment of the present application. The instruction routing reads the dispatch table entries in each dispatch queue, first sorts the dispatch table entries according to the instruction waiting time to determine the emission priorities of each instruction to be scheduled. Further, it classifies them according to the linked list numbers of the dispatch table entries to classify multiple instructions belonging to the same dependency linked list into the same instruction set. Finally, for each instruction set corresponding to a dependency linked list, according to the dependency sequence numbers of the instructions, the instructions to be scheduled within the instruction set are sorted accordingly (sorting within the RAW chain). According to the principle of the higher the emission priority and the smaller the linked list number, the more preferential the scheduling. When there are enough empty table entries in the emission queue, the instructions to be scheduled are sent to the corresponding emission queue in sequence.
[0063] When dispatching to the emission queue, the instruction routing tries to dispatch instructions of the same instruction type in the same RAW chain to the same emission queue, which is convenient for them to enter the same execution queue subsequently, so as to improve the usage efficiency of the forwarded operands mentioned later as much as possible. And, since some instructions to be scheduled in actual applications have no RAW dependency with other instructions, that is, they do not belong to any dependency linked list, such instructions to be scheduled are scheduled according to their priorities. While ensuring that the instructions in the dependency linked list can be executed orderly, it avoids the blocking of instructions that do not belong to any dependency linked list.
[0064] On the basis of the above embodiment, as Figure 6 shown, it is a schematic diagram of the structure of the emission module provided by an embodiment of the present application. As an implementable way, in one embodiment, the emission module includes:
[0065] A buffer unit ( Figure 6 not shown in the figure), which is used to screen a corresponding number of instructions to be scheduled in each emission queue as instructions to be emitted according to the communication bandwidth of each emission queue, cache the instructions to be emitted into a preset emission buffer area, and at the same time send the instructions to be emitted to the structural conflict check unit;
[0066] The structural conflict check unit is used to perform a structural conflict check on the received instructions to be emitted. When the structural conflict check passes, it sends a read data signal to the independent forwarding unit to obtain the source operands of the instructions to be emitted based on the independent forwarding unit in the preset register file and write the source operands into the preset data buffer area;
[0067] The dynamic scheduler is used to determine the target execution unit of any instruction to be emitted in the preset emission buffer area, obtain the current occupancy of the target execution unit, and determine the best issuing time of the instruction to be emitted according to the current occupancy of the target execution unit, so as to send the instruction to be emitted to the target execution unit at the best issuing time;
[0068] An independent pre-forwarding unit for reading the source operands of the instruction to be issued from a preset data buffer and sending the source operands of the instruction to be issued to the target execution unit while the dynamic scheduler sends the instruction to be issued to the target execution unit.
[0069] Specifically, the communication bandwidth determines the amount of data that can be transmitted per unit time. Therefore, the buffer unit determines how many instructions can be taken out from the issue queue according to this bandwidth. After screening out the instructions to be issued, they are cached in a preset issue buffer to facilitate the operation of subsequent units. At the same time, the buffer unit also sends these instructions to be issued to the structural conflict check unit for the next step of inspection. After receiving the instructions to be issued sent by the buffer unit, the structural conflict check unit performs a structural conflict check. A structural conflict refers to a conflict that occurs when multiple instructions compete for the same hardware resource simultaneously. For example, multiple instructions may need to use the same execution unit at the same time. If it is checked that there is no structural conflict among the instructions to be issued, the structural conflict check unit sends a read data signal to the independent pre-forwarding unit.
[0070] Among them, after receiving the read data signal, the independent pre-forwarding unit obtains the source operands required for the instruction to be issued in the preset register file and writes these source operands into the preset data buffer to prepare the data for the execution of the instruction. The dynamic scheduler first determines the target execution unit of the instruction to be issued, that is, on which execution unit this instruction will finally be executed, such as a vector operation execution unit and a floating-point operation execution unit, etc. Then, the dynamic scheduler obtains the current occupancy of the target execution unit to understand whether this execution unit is being used by other instructions. According to the occupancy of the target execution unit, the dynamic scheduler determines the best issue timing of the instruction to be issued. For example, if the target execution unit is currently busy, the dynamic scheduler will send the instruction to be issued when it is about to enter the idle state to ensure that the instruction can be executed smoothly, and at the same time ensure that the execution unit is fully utilized and reduce pipeline bubbles.
[0071] Among them, the independent pre-forwarding unit is used to read the source operands of the instruction to be issued from the preset data buffer. When the dynamic scheduler sends the instruction to be issued to the target execution unit, the independent pre-forwarding unit will simultaneously send the source operands of the instruction to be issued to the target execution unit to reduce the time for the instruction to wait for the operands and improve the efficiency of instruction execution.
[0072] Specifically, such as Figure 6As shown in the figure, the emission module (emission stage) of the embodiment of the present application has a two-stage pipeline. The first-stage pipeline is responsible for selecting instructions from each emission queue. In this stage, the corresponding number of instructions is selected according to the width of each emission queue and enters the instruction emission buffer (preset emission buffer); the second-stage pipeline is responsible for sending the instructions in the instruction emission buffer to the structural conflict check module. After two cycles of the emission stage, the instructions in the pipeline will enter the structural conflict check module. The structural conflict check module consumes one cycle for structural conflict check, and then consumes one cycle to read the required source operands from the register file. The source operands enter the data buffer (preset data buffer) through the data path. At this point, the preparatory work required for instruction execution has been completed, and the processing will be carried out in the execution unit in the next cycle. Since the register renaming module and the dispatch module have solved the data dependence problem between instructions, among the instructions in the emission queue, instructions without RAW data dependencies only need to be selected in the queue order, while instructions with RAW dependencies need to be selected in the order of their positions in the RAW chain. Among them, as Figure 6 shown, the emission module is also applicable to sending memory access instructions in the LSQ queue to the corresponding execution unit.
[0073] Specifically, in one embodiment, the number of independent early forwarding units is the same as the number of emission queues, and the independent early forwarding units correspond to the emission queues one by one. Each independent early forwarding unit is provided with a preset data buffer, and communication links are provided between the independent early forwarding units.
[0074] The independent early forwarding unit is further used for:
[0075] receiving the instruction execution result returned by any instruction unit; writing the instruction execution result as a source operand into the preset register file, and at the same time, according to the read data signal currently sent by the structural conflict check unit, transmitting the instruction execution result as an early forwarded operand to the preset data buffer of the corresponding target independent early forwarding unit through the communication link.
[0076] The target independent early forwarding unit is used for:
[0077] judging whether the early forwarded operand matches the currently to-be-emitted instruction. If it matches, using the early forwarded operand as the source operand of the currently to-be-emitted instruction. If it does not match, obtaining the source operand of the currently to-be-emitted instruction from the preset register file, and replacing the early forwarded operand in the preset data buffer with the source operand.
[0078] Specifically, the independent forwarding unit before independence receives the instruction execution result returned by any instruction unit. The independent forwarding unit before independence writes the received instruction execution result as the source operand into the preset register file to update the data in the register file. At the same time, when the structure conflict check unit sends a read data signal, the independent forwarding unit before independence will, according to this signal, transfer the instruction execution result as the forwarded operand to the preset data buffer of the corresponding target independent forwarding unit through the communication link. For example, if the execution result of instruction A is the source operand of instruction B, after instruction A is executed, its result will be sent to the independent forwarding unit corresponding to instruction A. After receiving the read data signal, the independent forwarding unit corresponding to instruction A will transfer the result to the preset data buffer of the target independent forwarding unit corresponding to instruction B.
[0079] Among them, after receiving the forwarded operand, the target independent forwarding unit will perform a matching judgment to determine whether this forwarded operand matches the source operand required by the current instruction to be issued. If it matches, the forwarded operand will be directly used as the source operand of the current instruction to be issued, which can avoid reading data from the preset register file again and save time. If it does not match, the target independent forwarding unit will obtain the source operand of the current instruction to be issued from the preset register file and replace the forwarded operand in the preset data buffer with the newly obtained source operand to ensure that the instruction can be executed using the correct operand.
[0080] Specifically, after the forwarded operand arrives at the target independent forwarding module, it will stay in the preset data buffer and wait until it is replaced by the data forwarding of the execution result of the next instruction. Before the instruction is sent to the execution unit, the dynamic scheduler will detect whether the forwarded operand (source operand) of the execution unit corresponding to the instruction is ready and whether it matches the instruction. If it is ready and matches, the instruction will start to execute directly; if the forwarded operand has not arrived yet, the instruction will wait here; if the forwarded operand does not match, it means that the source operand required by the current instruction has already been written into the register file, and the register file will be read immediately to obtain the source operand.
[0081] It should be noted that since the independent forwarding modules are linked through a communication link (cross-issue queue forwarding module), in the case where the communication distance between the independent forwarding unit corresponding to instruction A and the target independent forwarding unit corresponding to instruction B is relatively long, there will be a certain delay in the transfer of the forwarded operand. When the forwarded operand arrives at the target independent forwarding unit, it is possible that the target independent forwarding unit has already started instruction C after corresponding instruction B, so it will cause the situation where the forwarded operand does not match the source operand required by the current instruction to be issued. The embodiment of the present application is precisely to avoid the occurrence of this situation, and selects to send multiple instructions to be scheduled of the same instruction type belonging to the same dependency list to the same issue queue to shorten the communication distance between the independent forwarding units corresponding to the instructions belonging to the same dependency list.
[0082] Specifically, in one embodiment, each emission queue is provided with a corresponding buffer unit. The buffer unit is used to screen the instructions to be scheduled in the emission queue as the instructions to be emitted according to the head pointer direction of the emission queue; determine whether there is a dependency for the instruction to be emitted; directly cache the instruction to be emitted into a preset emission buffer when there is no dependency for the instruction to be emitted; determine whether an instruction wake-up signal is received currently when there is a dependency for the instruction to be emitted; cache the instruction to be emitted into the preset emission buffer when an instruction wake-up signal is received currently; determine whether an instruction cancellation signal is received for the instruction to be sent cached in the preset emission buffer within a preset waiting period; send the instruction to be emitted to the structural conflict check unit when the instruction to be sent cached in the preset emission buffer does not receive an instruction cancellation signal within the preset waiting period; and return to the process of determining whether an instruction wake-up signal is received currently when the instruction to be sent cached in the preset emission buffer receives an instruction cancellation signal within the preset waiting period.
[0083] Among them, when any instruction belonging to the same dependency linked list has been sent to the execution unit, the emission queue corresponding to this instruction sends an instruction wake-up signal to the emission queue of the next instruction in the dependency linked list.
[0084] It should be noted that by judging and processing the instruction dependency relationship, the instruction is cached into the preset emission buffer only when the dependency condition is met, avoiding unnecessary resource occupation and instruction waiting. By using the instruction wake-up signal mechanism, it is ensured that the instructions in the dependency linked list are executed sequentially in order, making the connection between instructions smoother and avoiding delays caused by chaotic instruction execution order.
[0085] Specifically, as Figure 7As shown, it is a schematic diagram of the operation process of the buffer unit provided by the embodiment of the present application. When an instruction moves to the head of the issue queue, it is judged whether it has a RAW dependency. If there is no RAW dependency, it is directly selected into the preset issue buffer; if there is a RAW dependency, it is checked whether there is an instruction wake-up signal from other issue queues at present. If the wake-up signal of this instruction has been received, it is selected into the preset issue buffer. If there is no wake-up signal of this instruction, it means that the instructions before the RAW chain where this instruction is located have not been executed or selected yet. Then this instruction enters the loop waiting area. The issue queue includes this loop waiting area until the corresponding wake-up signal arrives, and then it is selected. After the instruction enters the instruction buffer, it waits for one cycle. If no instruction cancellation signal is received, it is sent to the structural conflict check module and the Busy field is set to 1. If an instruction cancellation signal is received, it means that this instruction is cancelled or blocked by the upper level from execution. Therefore, it returns to the process of judging whether an instruction wake-up signal is received at present until it is sent to the structural conflict check module. The setting of the instruction cancellation signal enables the system to handle various emergencies. The instruction can return to the waiting state and wait for the right opportunity to execute again, improving the fault tolerance of the system.
[0086] Among them, when any instruction to be sent has no instruction wake-up signal for a long time, that is, when the waiting duration of the instruction to be sent in the loop waiting area reaches the upper limit value, the execution of this instruction can be directly cancelled because continuing to wait may seriously affect the processor performance and instruction execution efficiency. After cancelling the instruction execution, the system can release the relevant resources occupied by this instruction, such as the position in the issue queue, register resources, etc., so as to provide an execution opportunity for other instructions. At the same time, the system can also record the relevant information about the cancellation of this instruction execution, such as the identification of the instruction, the reason for cancellation, etc., which is convenient for subsequent debugging and analysis.
[0087] Correspondingly, in an embodiment, when the waiting duration of the instruction to be sent in the loop waiting area reaches the upper limit value, the exception handling mechanism of the system can also be triggered. The exception handling mechanism can analyze the situation of this instruction in detail and find out the reasons for the long-term inability to meet the dependency relationship, such as hardware failure, software logic error, etc. According to the analysis results, the exception handling program can take corresponding measures, such as performing hardware self-check, fixing software vulnerabilities, etc. At the same time, the exception handling mechanism can also send an alarm message to the user to remind them to pay attention to this problem so as to take further measures in time.
[0088] The instruction scheduling system provided by the embodiments of the present application includes: a dispatch module, a dependency check module, and a launch module; the dispatch module is configured to, for each acquired instruction to be scheduled and the renamed register index corresponding to the instruction to be scheduled, write the combined field of the instruction to be scheduled and the renamed register index as a dispatch table entry into the corresponding dispatch queue; the dependency check module is configured to perform a dependency check on the dispatch table entries in the dispatch queue according to a preset dependency check period, construct a dependency linked list according to the dependency check result, and send the instructions to be scheduled to the corresponding launch queue in the order of instruction execution represented by the dependency linked list; the launch module is configured to send the instructions to be scheduled in the launch queue to the corresponding execution unit and receive the instruction execution result returned by the execution unit. The method provided by the above solution improves the RAW dependency processing efficiency in instruction scheduling of a superscalar out-of-order processor, shortens the waiting duration of instructions, avoids wasting pipeline instruction storage resources, and reduces the instruction execution delay by ensuring that the sending order of instructions with RAW dependencies matches the execution order. Moreover, the RAW dependency problem is actively solved in advance to reduce the waiting time of instructions in the launch stage; the independent forwarding module combines with the RAW chain to enable the instruction execution result to be forwarded to the data buffer before the execution unit in time before the next instruction enters the execution unit, reducing the bubbles in the pipeline.
[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that the system according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0090] The embodiments of the present application provide an instruction scheduling method for performing instruction scheduling for a superscalar out-of-order processor. The execution subject of the embodiments of the present application is an electronic device, such as a server, a desktop computer, a laptop computer, a tablet computer, and other electronic devices that can be used to configure a superscalar out-of-order processor and perform instruction scheduling for the superscalar out-of-order processor.
[0091] As Figure 8 shown, it is a schematic flowchart of the instruction scheduling method provided by the embodiments of the present application, and the method includes:
[0092] Step 801, acquire any instruction to be scheduled and the renamed register index corresponding to the instruction to be scheduled;
[0093] Step 802, write the combined field of the instruction to be scheduled and the renamed register index as a dispatch table entry into the corresponding dispatch queue;
[0094] Step 803, perform a dependency check on the dispatch table entries in the dispatch queue according to a preset dependency check period;
[0095] Step 804: construct a dependency linked list according to the dependency check result;
[0096] Step 805: send the instructions to be scheduled to the corresponding emission queue in the order of instruction execution represented by the dependency linked list;
[0097] Step 806: send the instructions to be scheduled in the emission queue to the corresponding execution unit and receive the instruction execution result returned by the execution unit.
[0098] For the description of the features in the corresponding embodiments of the instruction scheduling method, reference can be made to the relevant descriptions of the corresponding embodiments of the instruction scheduling system, which will not be elaborated here one by one.
[0099] An embodiment of the present application also provides an electronic device, as Figure 9 shown, which is a schematic structural diagram of the electronic device provided by the embodiment of the present application, including a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in the above-mentioned embodiment of the instruction scheduling method.
[0100] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in the above-mentioned embodiment of the instruction scheduling method when running.
[0101] In an exemplary embodiment, the above-mentioned computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disk, magnetic disk or optical disc, etc., all kinds of media that can store computer programs.
[0102] An embodiment of the present application also provides a computer program product. The above-mentioned computer program product includes a computer program, and the steps in the above-mentioned embodiment of the instruction scheduling method are implemented when the computer program is executed by a processor.
[0103] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and the steps in the above-mentioned embodiment of the instruction scheduling method are implemented when the computer program is executed by a processor.
[0104] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0105] The above has introduced in detail an instruction scheduling system, method, and electronic device provided by this application. Specific examples have been used herein to illustrate the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. An instruction scheduling system, characterized in that, It includes: a dispatch module, a dependency check module, and a launch module; The dispatch module is used to, for each obtained pending scheduling instruction and the renamed register index corresponding to the pending scheduling instruction, write the combined field of the pending scheduling instruction and the renamed register index as a dispatch entry into the corresponding dispatch queue; The dependency check module is used to perform dependency checks on the dispatch entries in the dispatch queue according to a preset dependency check period, construct a dependency linked list based on the dependency check results, and send the pending scheduling instructions to the corresponding launch queue in the order of instruction execution represented by the dependency linked list; The launch module is used to send the pending scheduling instructions in the launch queue to the corresponding execution unit and receive the instruction execution results returned by the execution unit.
2. The instruction scheduling system according to claim 1, wherein The dispatch module is used for: Obtain the opcode of the pending scheduling instruction; Determine the instruction type of the pending scheduling instruction according to the opcode of the pending scheduling instruction; Write the combined field of the pending scheduling instruction and the renamed register index as a dispatch entry into the corresponding dispatch queue according to the instruction type of the pending scheduling instruction; wherein, the dispatch queue corresponds one-to-one with the instruction type.
3. The instruction scheduling system according to claim 1, characterized in that, The dependency check module is used for: For any one of the dispatch queues, perform dependency checks according to the constituent fields of each dispatch entry in the dispatch queue to determine the instruction execution order, source register, and destination register of each pending scheduling instruction in the dispatch queue; Construct a dependency linked list according to the instruction execution order, source register, and destination register of each pending scheduling instruction in each dispatch queue; wherein, the dependency check results include the instruction execution order, source register, and destination register of each pending scheduling instruction in each dispatch queue; the dependency linked list uses the pending scheduling instruction as a node, each of the nodes is connected in the order of instruction execution, and each of the nodes is connected through the source register and the destination register, the source register is used as the input end of the node, and the destination register is used as the output end of the node.
4. The instruction scheduling system according to claim 1, wherein The dependency check module is further used for: For any one of the pending scheduling instructions, determine the linked list number and the dependency order number in the dependency linked list of the pending scheduling instruction according to the dependency linked list of the pending scheduling instruction; Add the linked list number and the dependency order number in the dependency linked list of the pending scheduling instruction to the dispatch entry corresponding to the pending scheduling instruction; wherein, the linked list number and the dependency order number are used to represent the order of instruction execution.
5. The instruction scheduling system according to claim 1, characterized in that, The dependency check module is used for: Determine the launch priority of each pending scheduling instruction according to the instruction waiting duration represented by the dispatch entry; wherein, the instruction waiting duration and the launch priority are positively correlated; Send each pending scheduling instruction to the corresponding launch queue in the order of instruction execution represented by the dependency linked list, the launch priority of each pending scheduling instruction, and the instruction type of each pending scheduling instruction; wherein, each instruction type corresponds to at least one launch queue, and multiple pending scheduling instructions of the same instruction type belonging to the same dependency linked list will be sent to the same launch queue.
6. The instruction scheduling system according to claim 1, wherein The transmitting module includes: A buffer unit, which is used to screen a corresponding number of to-be-scheduled instructions as to-be-transmitted instructions in each of the transmitting queues according to the communication bandwidth of each transmitting queue, cache the to-be-transmitted instructions into a preset transmission buffer, and at the same time send the to-be-transmitted instructions to a structural conflict checking unit; A structural conflict checking unit, which is used to perform structural conflict checking on the received to-be-transmitted instructions. When the structural conflict checking passes, it sends a read data signal to an independent forwarding unit to obtain the source operands of the to-be-transmitted instructions in a preset register file based on the independent forwarding unit, and write the source operands into a preset data buffer; A dynamic scheduler, which is used to determine the target execution unit of any to-be-transmitted instruction in the preset transmission buffer, obtain the current occupancy of the target execution unit, and determine the best issuing timing of the to-be-transmitted instruction according to the current occupancy of the target execution unit, so as to send the to-be-transmitted instruction to the target execution unit at the best issuing timing; An independent forwarding unit, which is used to read the source operands of the to-be-transmitted instructions from the preset data buffer, and send the source operands of the to-be-transmitted instructions to the target execution unit while the dynamic scheduler sends the to-be-transmitted instructions to the target execution unit.
7. The instruction scheduling system according to claim 6, wherein The number of the independent forwarding units is the same as the number of the transmitting queues, and the independent forwarding units correspond to the transmitting queues one by one. Each independent forwarding unit is provided with a preset data buffer, and communication links are provided between the independent forwarding units; The independent forwarding unit is further used for: Receiving the instruction execution result returned by any instruction unit; Writing the instruction execution result as a source operand into the preset register file, and at the same time, according to the current read data signal sent by the structural conflict checking unit, transmitting the instruction execution result as a forwarded operand to the preset data buffer of the corresponding target independent forwarding unit through the communication link; The target independent forwarding unit is used for: Judging whether the forwarded operand matches the current to-be-transmitted instruction. If it matches, using the forwarded operand as the source operand of the current to-be-transmitted instruction. If it does not match, obtaining the source operand of the current to-be-transmitted instruction from the preset register file, and replacing the forwarded operand in the preset data buffer with the source operand.
8. The instruction scheduling system according to claim 6, wherein Each of the transmitting queues is provided with a corresponding buffer unit, and the buffer unit is used for: Screening to-be-scheduled instructions as to-be-transmitted instructions in the transmitting queue according to the head pointer direction of the transmitting queue; Judging whether the to-be-transmitted instruction has dependencies; When the to-be-transmitted instruction has no dependencies, directly caching the to-be-transmitted instruction into a preset transmission buffer; When the to-be-transmitted instruction has dependencies, judging whether an instruction wake-up signal is received currently; When an instruction wake-up signal is received currently, caching the to-be-transmitted instruction into a preset transmission buffer; Judging whether an instruction cancellation signal is received for the to-be-transmitted instructions cached in the preset transmission buffer within a preset waiting period; When the instruction to be sent cached in the preset transmission buffer does not receive an instruction cancellation signal within the preset waiting period, the instruction to be transmitted is sent to the structural conflict check unit; When the instruction to be sent cached in the preset transmission buffer receives an instruction cancellation signal within the preset waiting period, return to the process of determining whether an instruction wake-up signal is received currently; Among them, when any instruction belonging to the same dependency linked list has been sent to the execution unit, the transmission queue corresponding to the instruction sends an instruction wake-up signal to the transmission queue of the next instruction in the dependency linked list.
9. An instruction scheduling method, characterized in that, It includes: Obtain any instruction to be scheduled and the rename register index corresponding to the instruction to be scheduled; Use the combined field of the instruction to be scheduled and the rename register index as a dispatch table entry and write it into the corresponding dispatch queue; Perform dependency check on the dispatch table entries in the dispatch queue according to the preset dependency check period; Construct a dependency linked list according to the dependency check result; Send the instructions to be scheduled in sequence to the corresponding transmission queue according to the execution sequence of the instructions represented by the dependency linked list; Send the instructions to be scheduled in the transmission queue to the corresponding execution unit and receive the instruction execution result returned by the execution unit.
10. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the instruction scheduling method as claimed in claim 9 when executing the computer program.
Citation Information
Cited By
Data processing method and device, equipment, storage medium and program product
CN121742905A