Pseudo out-of-order instruction scheduling method based on branch jump
By introducing a pseudo-out-of-order instruction scheduling method based on branch jump into the processor, the branch instructions are conditionally suspended, which solves the high power consumption problem caused by the existing branch instruction scheduling method and achieves more efficient processor performance and stability.
Patent Information
- Application Number
- CN202510663355.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
The unreasonable existing branch instruction scheduling methods lead to large power consumption of the processor, which seriously affects the processor performance.
A pseudo-out-order instruction scheduling method based on branch jump is proposed. By conditional pause of branch instructions, the cost of branch prediction errors caused by out-of-order execution is avoided, and the power consumption of the processor is significantly reduced.
Through pseudo-out of order scheduling of branch instructions, the power consumption of the processor is significantly reduced, the execution unit is avoided and the performance and stability of the processor is improved.
Smart Images

Figure CN120179296A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hardware processors, and particularly relates to a pseudo out-of-order instruction scheduling method based on branch jump. Background Art
[0002] Power consumption optimization in high-performance processors has always been a hot topic in the research of processor technologies and is also a major factor restricting the performance stability of processors. In order to adapt to the increasing number of processor backend execution units and complex data path networks without affecting the computing performance, it is becoming increasingly important to perform power consumption optimization.
[0003] As Figure 1 shown, the general structure of a superscalar processor can execute multiple instructions simultaneously. Among them, the central processing unit (CPU) includes an instruction fetch unit, a level-1 instruction cache, a pre-decode unit, a fetch buffer, a decode unit, a rename register, a dispatch and issue unit, an execute unit (EU), a reorder buffer (ROB), a retire unit, etc. Among them, the execute unit can execute multiple instructions within a single clock cycle, exceeding the ability of a general processor to process only one instruction at a time. The main feature of the superscalar architecture is that it has multiple execute units and can parallelly process different types of operations, such as integer operations, floating-point operations, load operations, store operations, etc.
[0004] Branch instructions are a special type of instructions that can disrupt the entire execution order. Therefore, in the out-of-order unit, the jump of the branch instruction will cause these units to do useless work. The scheduling of branch instructions is an effective method to reduce system power consumption. Currently, many branch instruction scheduling methods adopt a completely out-of-order approach and then make a prediction judgment based on the results. However, this scheduling method will bring two problems: First, when a branch instruction is executed, there are many cases where many instructions after this branch instruction have already been executed. If the branch instruction prediction fails, then the instructions that have already been executed are equivalent to doing useless work; Second, the out-of-order execution of branch instructions will cause the results of branch instructions to be stored in batches, resulting in an increase in resource storage and also an increase in power consumption. Therefore, how to reasonably schedule branch instructions has become a key factor affecting processor performance.
[0005] In summary, the unreasonable existing branch instruction scheduling method leads to a large processor power consumption, seriously affecting the performance of the processor. Summary of the Invention
[0006] In view of this, the present invention proposes a pseudo-out-of-order instruction scheduling method based on branch jump. By conditionally suspending branch instructions, the cost of branch prediction errors caused by out-of-order execution is avoided, and the power consumption of the processor is significantly reduced, so as to solve the technical defect that the existing branch instruction scheduling method is unreasonable and leads to poor processor performance.
[0007] In a first aspect, the present invention provides a pseudo-out-of-order instruction scheduling method based on branch jump, which is applied to the dispatch and issue units in a superscalar processor. The method includes: If the currently received instruction is a branch instruction, dispatch the current branch instruction to the branch instruction issue queue; Based on instruction prediction analysis, determine whether the current branch instruction needs to be speculatively awakened so as to wake up the current branch instruction in advance; If it is determined that the current branch instruction needs to be speculatively awakened, determine whether the oldest branch instruction in the branch instruction issue queue can be awakened within two clock cycles; If it is determined that the oldest branch instruction in the branch instruction issue queue cannot be awakened within two clock cycles, suspend the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, so that other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction cannot be issued to the execution unit.
[0008] Further, in the above pseudo-out-of-order instruction scheduling method based on branch jump, after suspending the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, it further includes: After the oldest branch instruction is awakened, issue the oldest branch instruction to the execution unit, and enable other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to be scheduled and issued to the execution unit after being awakened.
[0009] Further, in the above pseudo-out-of-order instruction scheduling method based on branch jump, after determining whether the oldest branch instruction in the branch instruction issue queue can be awakened within two clock cycles, it further includes: If it is determined that the oldest branch instruction can be awakened within two clock cycles, issue the oldest branch instruction to the execution unit, and enable instructions in other queues in the awakened state to be scheduled and issued to the execution unit.
[0010] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, after determining whether the current branch instruction needs to be speculatively awakened based on instruction prediction analysis and thus awakening the current branch instruction in advance, it further includes: If it is determined that the current branch instruction does not need to be speculatively awakened, the current branch instruction continues to be stored in the branch instruction issue queue and waits to be awakened.
[0011] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, the method further includes: If the currently received instruction is not a branch instruction, dispatch the current non-branch instruction to the non-branch instruction issue queue corresponding to the instruction type; Based on the instruction type and the dependency relationship, determine whether the current non-branch instruction needs to be speculatively awakened so as to awaken the current non-branch instruction in advance; If it is determined that there are no instructions in the non-branch issue instruction queue that need to wait for awakening for more than two clock cycles, launch other non-branch instructions in the non-branch issue instruction queue into the execution unit.
[0012] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, the method further includes: When the oldest branch instruction in the branch instruction issue queue needs to access the operand provided by the memory instruction but has not been awakened after waiting for two clock cycles, or when other dependent instructions that need to provide operands for the oldest branch instruction in the branch instruction issue queue also have not been awakened after waiting for two clock cycles, pause the operation of launching other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are arranged after the oldest branch instruction into the execution unit.
[0013] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, the method further includes: After the branch instruction is executed, when it is determined that the branch prediction result is incorrect, set all the register states awakened by the branch instruction to non-awakened, and rewrite the instructions awakened by the branch instruction into the branch instruction issue queue to achieve context restoration.
[0014] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, the method for determining whether the current branch instruction needs to be speculatively awakened based on instruction prediction analysis includes: If there are instruction-dependent operands in the current branch instruction that can be prepared within the speculative clock period, determine that the current branch instruction needs to be speculatively awakened.
[0015] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, the oldest branch instruction refers to the branch instruction in the branch instruction issue queue that waits for its operand to be prepared earliest.
[0016] Further, in the above-mentioned pseudo-out-of-order instruction scheduling method based on branch jump, when the current oldest branch instruction that needs to access the memory instruction to provide an operand in the branch instruction issue queue has not been awakened after waiting for two clock cycles, the operation of sending the instructions in the branch instruction issue queue and the instructions in the non-branch instruction issue queue that are ranked after the current oldest branch instruction to the execution unit is paused.
[0017] The pseudo-out-of-order instruction scheduling method based on branch jump provided by the present invention has the following beneficial effects: The method for processing branch instructions in the present invention is that when it is possible to speculate and wake up the current branch instruction and the current oldest branch instruction cannot be awakened within two clock cycles, the emission of other instructions in the branch instruction issue queue and other instruction issue queues that are ranked after the current oldest branch instruction is paused, so as to avoid the premature execution of uncertain branches, thereby converting the full out-of-order execution of branch instructions into pseudo out-of-order execution. By conditionally pausing the branch instructions, the cost of branch prediction errors caused by out-of-order execution is avoided.
[0018] The present invention provides a pseudo-out-of-order instruction scheduling method based on branch jump that can reduce power consumption. By dynamically scheduling the instructions after the branch instruction according to the readiness of the branch instruction, it effectively avoids the execution unit from doing useless work, thereby reducing the power consumption of the processor. By analyzing the in-order execution function of the instructions of the current processor, the present invention proposes a special scheduling mode for branch instructions. On this basis, the full out-of-order execution of branch instructions is changed to the pseudo out-of-order execution mode of branch instructions, which can ensure that the subsequent execution components will not do useless work in the case of prediction errors, save power consumption, and improve stability.
[0019] The pseudo-out-of-order instruction scheduling method based on branch jump provided by the present invention reduces unnecessary register read / write access operations and arithmetic unit flip operations, reduces the power consumption of the processor, and also reduces the logical complexity of branch prediction failure recovery, thereby improving the timing. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only for further understanding of the embodiments of the present invention and constitute a part of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 is the general structure of an existing superscalar processor; Figure 2Schematic diagram of the internal structure of a dispatch and issue unit in a superscalar processor provided by an embodiment of the present invention, in which a rename register and an execution unit are also shown; Figure 3 Flowchart of a pseudo out-of-order instruction scheduling method based on branch jump provided by an embodiment of the present invention; Figure 4 Another flowchart of a pseudo out-of-order instruction scheduling method based on branch jump provided by an embodiment of the present invention; Figure 5 Scheduling flowchart when a branch instruction cannot be woken up due to a TLB Miss provided by an embodiment of the present invention; Figure 6 Processing flowchart of the pipeline when branch prediction fails provided by an embodiment of the present invention; Figure 7 Schematic diagram of the structure of a superscalar processor provided by an embodiment of the present invention. Detailed implementation manners
[0021] The following examples are given in conjunction with the accompanying drawings to further illustrate the technical solutions provided by the present invention. It should be understood that the system structure and business scenarios provided in the embodiments of the present invention are mainly for explaining possible implementation manners of the technical solutions of the present invention, and should not be construed as the only limitation of the technical solutions of the present invention. Those of ordinary skill in the art will know that with the evolution of the system structure and the emergence of new business scenarios, the technical solutions provided by the present invention are equally applicable to similar technical problems.
[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. In case of inconsistency, the meaning described in this specification or the meaning obtained according to the content recorded in this specification shall prevail. In addition, the terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0023] An embodiment of the present invention provides a pseudo out-of-order instruction scheduling method based on branch jump, and this method is applied to a dispatch and issue unit in a superscalar processor. As Figure 2 shown, an embodiment of the present invention provides a schematic diagram of the internal structure of a dispatch and issue unit in a superscalar processor. As Figure 2As shown, the Dispatch and Issue unit obtains instructions from the Rename register, and then distributes the instructions to the corresponding Issue Queues (IQs) according to the instruction type. The Issue Queues are queues for storing instructions to be issued to the Execution Units (EUs) for execution. These instructions need to wait for their source operands before being issued. The Issue Queues include: the Branch Issue Queue (Br IQ) for handling branch instructions, the Load / Store Issue Queue (LS IQ) for load / store instructions, the ALU Issue Queue for handling integer operations such as addition, subtraction, shift, and logical processing, and the SPEC Issue Queue for handling special instructions for floating-point operations such as division and multiplication. In a pipelined processor, instructions are executed in multiple stages, such as instruction fetch, decoding, execution, memory access, and write-back. Branch instructions need to determine the control flow of the program during the execution stage, which may change the order of program execution.
[0024] As Figure 3 shown, an embodiment of the present invention provides a pseudo out-of-order instruction scheduling method based on branch jump, including; Step 11: If the currently received instruction is a branch instruction, dispatch the current branch instruction to the Branch Issue Queue.
[0025] Step 12: Based on instruction prediction analysis, determine whether the current branch instruction needs to be speculatively woken up to wake up the current branch instruction in advance. If speculative wake-up is required, execute Step 131; if speculative wake-up is not required, execute Step 132.
[0026] To increase the opportunity for concurrent instruction execution, reduce the stalls caused by waiting for data, and keep the pipeline running at full load as much as possible, thereby improving the processor performance, before a branch instruction is executed, it can be assumed that the branch instructions that meet the conditions obtain the operands, and these branch instructions are woken up in advance, that is, the status of these branch instructions is set to ready.
[0027] Determine whether the current branch instruction can be speculatively woken up, that is, speculate whether the operand with instruction dependence in the branch instruction can be prepared within the speculation cycle, such as within two clock cycles. The dependent operand is the result of the execution of other instructions. For example, for two instructions, the addition instruction add(x1, x2, x3) and the conditional jump instruction bne(x1, x2, label), where the conditional jump instruction is the branch instruction and uses the result of the addition instruction. It is predicted that the result of the addition instruction can be calculated in one clock cycle. Therefore, the branch of the conditional jump instruction can be speculatively woken up. If the current branch instruction can be prepared within the speculation cycle, wake up the current branch instruction in advance. After waking up, the branch instruction can be issued to the execution unit for execution; if it cannot be completed within the speculation cycle, the current branch instruction waits in the branch instruction issue queue to be woken up after it is ready.
[0028] Specifically, in the embodiment of the present invention, the method for determining whether a branch instruction can be speculatively woken up is to determine whether the operand with instruction dependence in the branch instruction can be prepared within the speculation clock cycle, such as within two clock cycles.
[0029] Step 131: If it is determined that the current branch instruction needs to be speculatively woken up, determine whether the oldest branch instruction in the branch instruction issue queue can be woken up within two clock cycles. If the oldest branch instruction cannot be woken up within two clock cycles, execute step 141; if the oldest branch instruction can be woken up within two clock cycles, execute step 142.
[0030] In the embodiment of the present invention, the oldest branch instruction generally refers to the branch instruction that waits for the execution result of other instructions in the branch instruction issue queue, has a dependence relationship with other instructions, and the dependent instruction has a long execution time. That is, the branch instruction in the branch instruction issue queue that waits for its operand to be prepared earliest. The oldest branch instruction is ready after waiting in the issue queue for two clock cycles, that is, the oldest branch instruction is woken up. The oldest branch instruction is not ready after waiting in the issue queue for two clock cycles, that is, the oldest branch instruction is not woken up.
[0031] Step 132: If it is determined that the current branch instruction does not need to be speculatively woken up, the current branch instruction waits in the branch instruction issue queue to be woken up.
[0032] Step 141: If it is determined that the oldest branch instruction cannot be woken up within two clock cycles, suspend the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, so that other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction cannot be issued to the execution unit.
[0033] Suspend the instructions sent to the execution unit, including not only the branch instructions in the branch instruction issue queue, but also the instructions in the non-branch instruction issue queue, that is, the instructions in other instruction issue queues, such as the load / store instruction issue queue, the arithmetic instruction issue queue, or the special instruction issue queue.
[0034] Step 142: If it is determined that the oldest branch instruction can be woken up within two clock cycles, send the oldest branch instruction to the execution unit. All instructions in the wake-up state can be normally scheduled and sent to the execution unit.
[0035] The oldest branch instruction is the branch instruction that has been waiting for its operands to be ready earliest among all the branch instructions in the current branch instruction issue queue. This oldest branch instruction has been waiting in the current branch instruction issue queue for two clock cycles and is still not ready, and is still waiting for the calculation results of its dependent instructions, indicating that the dependent instructions are still in an uncertain state during their execution phase. This will also affect the prediction of the branch result, and it is not certain whether it is correct. Sending its subsequent instructions to the execution unit for early execution may result in the execution of invalid instructions. Therefore, suspend sending other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, so as to avoid executing potentially invalid instructions and reduce the possibility of performing useless work. All instructions after the oldest branch instruction are subsequent instructions related to the branch result of the oldest branch instruction, and these instructions have been woken up by early speculation.
[0036] If the oldest branch instruction is ready and woken up within two clock cycles, it indicates that it has obtained its dependent operands. The scheduling and execution of the subsequent instructions of the oldest branch instruction are determined, and executing the subsequent instructions of the oldest branch instruction will not cause invalid execution.
[0037] After suspending sending all instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, it is found that the oldest branch instruction is ready and in the wake-up state. At this time, it indicates that the oldest branch instruction has obtained its dependent operands. The scheduling and execution of the subsequent instructions of the oldest branch instruction are determined (it will not jump to special situations such as exceptions), and executing the subsequent instructions of the oldest branch instruction will not cause invalid execution. Therefore, after suspending sending all instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, after the oldest branch instruction is woken up, send the oldest branch instruction to the execution unit, and make all instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction be able to be scheduled and sent to the execution unit after being woken up, so that all instructions after the oldest branch instruction can be executed in the execution unit.
[0038] In the prior art, due to the complete out-of-order execution of instructions, there are ineffective operation phenomena in multiple arithmetic units. Moreover, with the increase in execution units and the growth of the number of instruction channels of various types, the number of devices in the processor backend design increases rapidly, and this power consumption waste situation will become more obvious. The embodiments of the present invention utilize the characteristics of branch instructions to conditionally pause the branch instructions, avoid the cost of branch prediction errors caused by out-of-order execution, significantly reduce the power consumption of the processor, solve the technical defect that the existing branch instruction scheduling method is unreasonable and leads to poor processor performance, effectively schedule the operations of the entire execution unit, and produce a significant effect on the overall power consumption optimization.
[0039] Figure 4 This is a flowchart of another pseudo out-of-order instruction scheduling method based on branch jump provided by the embodiments of the present invention. As Figure 4 shown, from the start stage to the dispatch stage of the instruction, it is judged whether this instruction is a branch instruction. If it is, it is sent to the branch instruction issue queue, otherwise the current instruction is sent to other instruction issue queues corresponding to the instruction type, such as the load / store instruction issue queue, the arithmetic instruction issue queue, or the special instruction issue queue.
[0040] If the currently received instruction is not a branch instruction, the current non-branch instruction is dispatched to the non-branch instruction issue queue corresponding to the instruction type. Then, based on the instruction type and the dependency relationship, it is judged whether the non-branch instruction needs to be speculatively woken up to wake up the non-branch instruction in advance. If it is determined that there is no instruction in the non-branch instruction issue queue that needs to wait for wake-up for more than two clock cycles, that is, the instructions in the non-branch instruction issue queue can be woken up within two clock cycles, then other non-branch instructions in the non-branch instruction issue queue are issued to the execution unit.
[0041] Processing method for non-branch instructions: Based on the ready state of the instruction itself, it is directly issued to the corresponding execution unit. If it is not ready, then based on the instruction type and the dependency relationship, predictive analysis is performed on the instruction to judge whether the instruction needs to be speculatively woken up. The specific method is that if the execution cycle of the instruction is determined, or the non-branch instruction is a memory access instruction, it is determined that the instruction needs to be woken up in advance.
[0042] If it is determined that the instructions in the non-branch instruction issue queue can be woken up within two clock cycles, other non-branch instructions in the non-branch instruction issue queue are all issued to the execution unit for execution. If it is determined that the instruction cannot be speculatively woken up, the instruction continues to wait in the instruction issue queue to be woken up.
[0043] The processing method for the branch instructions in the branch instruction issue queue is as follows: Determine whether the current branch instruction needs speculative wake-up. If it does not need speculative wake-up, the branch instruction continues to be stored in the instruction issue queue waiting to be woken up. Speculative wake-up means that there is an operand that needs to be used, which is the execution result of other instructions being executed. Determine whether the execution time of other instructions exceeds two clock cycles (ordinary integer operations are one clock cycle). If the execution time does not exceed two clock cycles, it is considered that speculative wake-up is possible. If speculative wake-up is required, determine whether the oldest branch instruction in the current instruction issue queue can be woken up within two beats, that is, within two clock cycles. If the current oldest branch instruction is not ready after two beats, suspend the issue of all instructions in the branch instruction issue queue and other instruction issue queues that are after the oldest branch instruction until the oldest branch instruction is ready and woken up. After that, issue the oldest branch instruction to the execution unit, and at the same time release the issue of other instructions in the branch instruction issue queue and other instruction issue queues that are after the oldest branch instruction, so that other instructions in the branch instruction issue queue and other instruction issue queues that are after the oldest branch instruction can be scheduled to be issued to the execution unit after being woken up.
[0044] If the current oldest branch instruction can be woken up within two beats, issue the oldest branch instruction to the execution unit, and it does not affect the issue of all instructions after the current branch instruction. That is, instructions in the branch instruction issue queue and other instruction issue queues that are in the wake-up state can be scheduled to be issued to the execution unit.
[0045] The method for processing branch instructions in the embodiments of the present invention is that when it is possible to speculate and wake up the current branch instruction but the current oldest branch instruction cannot be woken up within two beats, suspend the issue of all instructions in the branch instruction issue queue and other instruction issue queues that are after the current oldest branch instruction, and avoid executing uncertain branches in advance, so as to convert the fully out-of-order execution of branch instructions into pseudo out-of-order execution. By conditionally suspending branch instructions, the cost of branch prediction errors caused by out-of-order execution is reduced.
[0046] The scheduling method provided by the embodiments of the present invention can be applied to multiple scenarios, such as Figure 5 the scenario of TLB Miss (TLB miss) shown and Figure 6 the scenario when branch prediction fails shown.
[0047] Figure 5The scenario shown involves memory access. When the processor attempts to access memory, it first checks the Translation Lookaside Buffer (TLB) to determine whether the virtual address has been translated into a physical address. If the corresponding mapping is found in the TLB, i.e., a TLB hit, the physical address can be directly used for memory access. If no match is found, i.e., a TLB Miss occurs, the page table entry needs to be read from memory into the TLB to complete the address translation.
[0048] When a branch instruction enters the wake-up state in advance after being dispatched to the branch instruction issue queue and passing the speculative wake-up judgment, if the instruction before the current oldest branch instruction in the branch instruction issue queue, that is, the memory access instruction that feeds the operands for the current oldest branch instruction, experiences a TLB Miss during memory access, it will cause the operands of the current oldest branch instruction to not be ready within more than two clock cycles. Because when there is a TLB Miss, the page table entry needs to be read from memory into the TLB, and reading memory is a relatively time-consuming process. Therefore, to avoid useless work being done on the pipeline when there is a TLB Miss, when the current oldest branch instruction in the branch instruction issue queue that requires the memory access instruction to provide operands has not been ready after waiting for two clock cycles, the embodiments of the present invention pause sending the instructions in the branch instruction issue queue and other instruction issue queues that are after the current oldest branch instruction to the execution unit, avoiding the execution of instructions in the uncertain branch prediction on the pipeline due to TLB misses, thereby generating useless work on the pipeline, reducing unnecessary register read / write access operations and arithmetic unit flip operations, and reducing the power consumption of the processor. After pausing to send the instructions in the branch instruction issue queue and other instruction issue queues that are after the current oldest branch instruction to the execution unit, when it is found that a TLB Miss occurs during the memory access of the instruction before the current oldest branch instruction, the register states of all the registers that wake up the instruction before the current oldest branch instruction, i.e., the memory access instruction, are set to non-wake-up, and the instructions woken up by this memory access instruction are rewritten into the branch instruction issue queue to achieve context restoration, and then they re-enter the branch instruction issue queue waiting for their operands to be ready.
[0049] Figure 6 The scenario shown is the processing flow when branch prediction fails. As Figure 6As shown, after the execution of a branch instruction is completed, if it is found that the branch prediction result is incorrect, that is, the prediction is not to jump but actually needs to jump, or the prediction is to jump but actually does not jump, it is necessary to set the states of all registers awakened by this branch instruction to non-awakened, and rewrite the instructions awakened by this branch instruction into the branch instruction issue queue to achieve the restoration of the scene when the branch prediction fails. If the instruction before the oldest branch instruction in the branch instruction issue queue currently is a branch instruction, that is, there are multiple branch instructions awakened simultaneously, but since the final branch result can be calculated within three clock cycles, the validity of the entire instruction can also be quickly controlled. However, if the oldest branch instruction waits for two clock cycles without being awakened, it indicates that the branch results of the instructions before the oldest branch instruction are still undetermined, and these instructions may be invalid instructions. Moreover, the instructions after the oldest branch instruction have been awakened in advance, such as being issued to the execution unit for execution, which may be a waste of effort. In this case, the embodiment of the present invention pauses the sending of the instructions after the oldest branch instruction in the branch instruction issue queue and other instruction issue queues to the execution unit, avoiding the execution of the instructions after this branch instruction, that is, the execution of invalid instructions, thereby reducing the power consumption of the processor, reducing the logical complexity of the branch prediction failure recovery, and thus improving the timing.
[0050] As Figure 7 As shown, the present invention also provides a superscalar processor, which includes a dispatch unit 701 and an issue unit 702.
[0051] The dispatch unit 701 is configured to, when the current instruction to be dispatched is a branch instruction, dispatch the current branch instruction to the branch instruction issue queue; based on instruction prediction analysis, determine whether the current branch instruction needs to be speculatively awakened so as to awaken the current branch instruction in advance; if it is determined that the current branch instruction needs to be speculatively awakened, determine whether the oldest branch instruction in the branch instruction issue queue can be awakened within two clock cycles.
[0052] The issue unit 702 is configured to, when the dispatch unit 701 determines that the oldest branch instruction cannot be awakened within two clock cycles, pause the operation of issuing the other instructions after the oldest branch instruction in the branch instruction issue queue and the non-branch instruction issue queue to the execution unit, so that all the instructions after the oldest branch instruction in the branch instruction issue queue and the non-branch instruction issue queue cannot be issued to the execution unit.
[0053] Further, the issue unit 702 is configured to, after the oldest branch instruction is awakened, issue the oldest branch instruction to the execution unit, and enable all the instructions after the oldest branch instruction in the branch instruction issue queue and the non-branch instruction issue queue to be scheduled and issued to the execution unit after being awakened.
[0054] Further, when the dispatch unit 701 determines that the oldest branch instruction can be woken up within two clock cycles, the emission unit 702 is configured to emit the oldest branch instruction to the execution unit, and the instructions in the wake-up state can be scheduled and emitted to the execution unit simultaneously.
[0055] Further, when it is determined that there is no need to speculatively wake up the current branch instruction, the dispatch unit 701 is configured to continue to store the current branch instruction in the branch instruction emission queue and wait to be woken up.
[0056] Further, when the currently received instruction is not a branch instruction, the dispatch unit 701 is configured to dispatch the current non-branch instruction to the non-branch instruction emission queue corresponding to the instruction type; based on the instruction type and the dependency relationship, determine whether the non-branch instruction needs to be speculatively woken up so as to wake up the non-branch instruction in advance.
[0057] Further, when it is determined that there is no instruction in the non-branch instruction emission queue that needs to wait for more than two clock cycles to be woken up, the emission unit 702 is configured to emit the non-branch instruction to the execution unit.
[0058] Further, when the oldest branch instruction in the branch instruction emission queue needs to access the operand provided by the memory instruction but has not been woken up after waiting for two clock cycles, or when other dependent instructions that need to provide operands to the oldest branch instruction in the branch instruction emission queue also have not been woken up after waiting for two clock cycles, the dispatch unit 701 is configured to pause the operation of emitting other instructions in the branch instruction emission queue and the non-branch instruction emission queue that are ranked after the oldest branch instruction to the execution unit.
[0059] Further, the superscalar processor further includes a recovery unit, which is configured to, after the branch instruction is executed, when it is determined that the branch prediction result is incorrect, set the states of all registers woken up by the branch instruction to non-wake-up, and rewrite the instructions woken up by the branch instruction into the branch instruction emission queue to achieve context recovery.
[0060] Further, when it is determined that the operands on which the instructions in the current branch instruction depend can be prepared within the speculative clock period, the dispatch unit 701 is configured to determine that the current branch instruction needs to be speculatively woken up.
[0061] Further, when the oldest branch instruction in the branch instruction emission queue that needs to access the operand provided by the memory instruction has not been woken up after waiting for two clock cycles, the emission unit 702 is configured to pause the operation of sending the instructions in the branch instruction emission queue and other instruction emission queues that are ranked after the oldest branch instruction to the execution unit.
[0062] It should be noted that the embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all embodiments. The components of the embodiments of the present invention generally described and illustrated in the drawings can be arranged and designed in various different configurations. Therefore, the above detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents the selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0063] The terms "first", "second", "third", etc. or terms such as module A, module B, module C, etc. in the description and claims are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that, where permitted, the specific order or sequence can be interchanged so that the embodiments of the present invention described herein can be implemented in an order other than that illustrated or described herein.
[0064] In the above description, the reference numerals representing steps do not necessarily mean that the steps will be executed in this order. It may also include intermediate steps or be replaced by other steps. Where permitted, the order of the front and rear steps can be interchanged or executed simultaneously.
[0065] The term "comprising" used in the description and claims should not be construed as limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the presence of the stated features, wholes, steps or components, but does not exclude the presence or addition of one or more other features, wholes, steps or components and their groups. Therefore, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.
[0066] The phrase "an embodiment" or "embodiments" mentioned in this specification means that the specific features, structures or characteristics described in connection with the embodiment are included in at least one embodiment of the present invention. Therefore, the phrase "in an embodiment" or "in embodiments" that appears throughout this specification does not necessarily refer to the same embodiment, but may refer to the same embodiment. In addition, in various embodiments of the present invention, if there is no special explanation and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0067] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, all of which fall within the protection scope of the present invention.
Claims
1. A method for pseudo-out-of-order instruction scheduling based on branch jump, characterized in that Applied to the dispatch and issue unit in a superscalar processor, the method includes: If the currently received instruction is a branch instruction, dispatch the current branch instruction to the branch instruction issue queue; Based on instruction prediction analysis, determine whether the current branch instruction needs to be speculatively woken up to wake up the current branch instruction in advance; If it is determined that the current branch instruction needs to be speculatively woken up, determine whether the oldest branch instruction in the branch instruction issue queue can be woken up within two clock cycles; If it is determined that the oldest branch instruction in the branch instruction issue queue cannot be woken up within two clock cycles, suspend the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, so that other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction cannot be issued to the execution unit.
2. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that After suspending the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit, it further includes: After the oldest branch instruction is woken up, issue the oldest branch instruction to the execution unit, and enable other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to be scheduled and issued to the execution unit after being woken up.
3. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that After determining whether the oldest branch instruction in the branch instruction issue queue can be woken up within two clock cycles, it further includes: If it is determined that the oldest branch instruction can be woken up within two clock cycles, issue the oldest branch instruction to the execution unit, and instructions in other queues in the wake-up state can be scheduled and issued to the execution unit.
4. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that After determining based on instruction prediction analysis whether the current branch instruction needs to be speculatively woken up to wake up the current branch instruction in advance, it further includes: If it is determined that the current branch instruction does not need to be speculatively woken up, the current branch instruction continues to be stored in the branch instruction issue queue waiting to be woken up.
5. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that The method further includes: If the currently received instruction is not a branch instruction, dispatch the current non-branch instruction to the non-branch instruction issue queue corresponding to the instruction type; Based on the instruction type and dependency relationship, determine whether the current non-branch instruction needs to be speculatively woken up to wake up the current non-branch instruction in advance; If it is determined that there are no instructions in the non-branch issue instruction queue that need to wait for more than two clock cycles to be woken up, issue other non-branch instructions in the non-branch issue instruction queue to the execution unit.
6. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that The method further includes: When the oldest branch instruction in the branch instruction issue queue needs to access the operand provided by the memory instruction but has not been woken up after waiting for two clock cycles, or when other dependent instructions that need to provide operands to the oldest branch instruction in the branch instruction issue queue also have not been woken up after waiting for two clock cycles, suspend the operation of issuing other instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the oldest branch instruction to the execution unit.
7. The method for pseudo-out-of-order instruction scheduling based on branch jump according to claim 1, characterized in that The method further includes: After the branch instruction is executed, when it is determined that the branch prediction result is incorrect, set the states of all registers awakened by the branch instruction to non-awakened, and rewrite the instructions awakened by the branch instruction into the branch instruction issue queue to achieve context restoration.
8. The method for pseudo-out-of-order instruction scheduling based on branch jump according to any one of claims 1 to 6, characterized in that The method for determining whether the current branch instruction needs to be speculatively awakened based on instruction prediction analysis includes: If there are operand dependencies in the current branch instruction that can be prepared within the speculative clock period, determine that the current branch instruction needs to be speculatively awakened.
9. The method for pseudo-out-of-order instruction scheduling based on branch jump according to any one of claims 1 to 6, characterized in that The oldest branch instruction refers to the branch instruction in the branch instruction issue queue that waits for its operands to be prepared earliest.
10. The method for pseudo-out-of-order instruction scheduling based on branch jump according to any one of claims 1 to 6, characterized in that When the current oldest branch instruction that needs to access the memory instruction to provide an operand in the branch instruction issue queue has not been awakened after waiting for two clock cycles, pause the operation of sending the instructions in the branch instruction issue queue and the non-branch instruction issue queue that are after the current oldest branch instruction to the execution unit.
Citation Information
Patent Citations
Information processing method and device and storage medium
CN111290786A
High-performance embedded processor based on RISC-V architecture
CN116661870A
Instruction dispatching method, dispatching device and related equipment
CN118093024A
Instruction transmitting method and device for out-of-order processor
CN118760475A
Branch lookahead prefetch for microprocessors
US20060149933A1
Cited By
Instruction allocation method and device, processor, electronic equipment and storage medium
CN121501345A