Macro operation fusion
By introducing instruction decoding buffers and fusion predictors into the integrated circuit of the processor, detecting and processing macro operation sequences, the problem that macro operation fusion in the prior art is difficult to handle control flow instructions, achieving higher processor performance and less pipeline refresh.
Patent Information
- Application Number
- CN201980080209.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-10
- Filing Date
- 2019-12-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-12-10
AI Technical Summary
When existing processors perform macro operations fusion, it is difficult to effectively process control flow instructions, resulting in performance degradation and frequent pipeline refreshes.
The prefix of the macro operation sequence is detected by introducing an instruction decoding buffer and a fusion predictor circuit in the integrated circuit, and performing macro operations delays or advances based on predicting whether there is a fusion sequence.
Effectively remove control flow instructions, improve processor performance, reduce pipeline refresh, and avoid performance degradation caused by branch prediction misses.
Smart Images

Figure CN113196244B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to macro-operation fusion. Background Art
[0002] Processors sometimes perform macro-operation fusion, in which several Instruction Set Architecture (ISA) instructions are fused during the decoding phase and processed as one internal operation. Macro-operation fusion is a powerful technique for reducing the effective number of instructions. Summary of the Invention
[0003] An implementation of macro-operation fusion is disclosed herein.
[0004] In a first aspect, the subject matter described in this specification may be embodied in an integrated circuit for executing instructions, the integrated circuit including one or more execution resource circuits configured to execute micro-operations to support an instruction set including macro-operations; an instruction decoding buffer configured to store macro-operations fetched from a memory; an instruction decoder circuit configured to: detect a sequence of macro-operations stored in the instruction decoding buffer, the sequence of macro-operations including a control flow macro-operation followed by one or more additional macro-operations; determine micro-operations equivalent to the detected sequence of macro-operations, and forward the micro-operations to at least one of the one or more execution resource circuits for execution.
[0005] In a second aspect, the subject matter described in this specification may be embodied in a method including: fetching macro-operations from a memory and storing the macro-operations in an instruction decoding buffer; and detecting a sequence of macro-operations stored in the instruction decoding buffer, the sequence of macro-operations including a control flow macro-operation followed by one or more additional macro-operations; determining micro-operations equivalent to the detected sequence of macro-operations; forwarding the micro-operations to at least one execution resource circuit for execution.
[0006] In a third aspect, the subject matter described in this specification may be embodied in an integrated circuit for executing instructions, the integrated circuit including one or more execution resource circuits configured to execute micro-operations to support an instruction set including macro-operations; an instruction decoding buffer configured to store macro-operations fetched from a memory; a fusion predictor circuit configured to: detect a prefix of a sequence of macro-operations in the instruction decoding buffer, determine a prediction as to whether the sequence of macro-operations will complete and fuse in a next fetch from the memory, and based on the prediction, delay execution of the prefix until after the next fetch, enabling fusion of the sequence of macro-operations; an instruction decoder circuit configured to: detect a sequence of macro-operations stored in the instruction decoding buffer, determine micro-operations equivalent to the detected sequence of macro-operations, and forward the micro-operations to at least one of the one or more execution resource circuits for execution.
[0007] These and other aspects of the present disclosure are disclosed in the following detailed description, the appended claims, and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The present disclosure will be best understood from the following detailed description when read in conjunction with the accompanying drawings. It should be emphasized that, according to common practice, the various features of the drawings are not drawn to scale. Instead, the dimensions of the various features are arbitrarily enlarged or reduced for clarity.
[0009] Figure 1 is a block diagram of an example of a system for executing instructions from an instruction set through macro-operation fusion.
[0010] Figure 2 is a block diagram of an example of a system for executing instructions from an instruction set through macro-operation fusion with fusion prediction.
[0011] Figure 3 is a block diagram of an example of a system for fusion prediction.
[0012] Figure 4 is a flowchart of an example of a process for executing instructions from an instruction set through macro-operation fusion.
[0013] Figure 5 is a flowchart of an example of a process for predicting beneficial macro-operation fusion. DETAILED DESCRIPTION
[0014] Systems and methods for macro-operation fusion are disclosed. An integrated circuit (e.g., a processor or a microcontroller) can decode and execute macro-operation instructions of an instruction set (e.g., the RISC V instruction set). A sequence of multiple macro-operations decoded by the integrated circuit can be fused (i.e., combined) into a single equivalent micro-operation to be executed by the integrated circuit. In some embodiments, a control flow instruction can be fused with a subsequent data-independent instruction to form an instruction that does not require a control flow event in the pipeline. For example, a branch macro-instruction can be replaced with a non-branch micro-operation. By effectively removing control flow instructions through macro-operation fusion, performance can be improved. For example, a performance degradation associated with a branch prediction miss can be avoided.
[0015] In some conventional processors, a conditional branch is predicted, and if the prediction is correct as expected, a pipeline flush is typically initiated. If the prediction made is incorrect, the pipeline will be flushed again to restart on the sequential path. If a conditional branch is predicted as not taken but is actually taken, the pipeline will also be flushed. Pipeline flushing is avoided only when a predicted conditional branch is not taken and the branch is actually not taken. Table 1 below shows the number of pipeline flushes that a conventional processor can perform using branch prediction.
[0016] Table 1:
[0017]
[0018] In some cases, branches may be difficult to predict, which can not only cause many pipeline flushes but also contaminate the branch predictor, thus degrading the performance of other predictable branches.
[0019] For example, an unconditional jump with a short forward offset can be fused with one or more subsequent instructions. The unconditional jump and the skipped instructions can be fused into a single non-jump micro-operation that has no impact on the machine except for advancing the program counter by the jump offset. Benefits may include removing the pipeline flush typically required to execute a jump using a no-operation (NOP) instruction that only advances the program counter without a pipeline flush. In some embodiments, more than one instruction can be skipped. In some embodiments, one or more target instructions can also be fused into the micro-operation.
[0020] For example, conditional branches on one or more instructions can be fused. In some embodiments, the conditional branch is fused with the following instruction such that the combination is executed as a single non-branch instruction. For example, if the condition is false, the internal micro-operation can disable the write to the target or can be defined to always write to the target, which can simplify the operation of an out-of-order superscalar machine with register renaming. The fused micro-operation can be executed as a non-branch instruction, thus avoiding a pipeline flush and also avoiding contaminating the branch predictor state. The sequence of macro-operations to be fused can include multiple instructions following a control flow instruction.
[0021] For example, a branch instruction and a single jump instruction can be fused. Compared with a conditional branch, an unconditional jump instruction can direct the instruction to a farther place, but sometimes a conditional branch to a long-distance target is desired. This can be done by an instruction sequence that can be fused internally by the processor. For example, a branch instruction and a function call sequence can be fused. A function call instruction is not conditional, so it can be paired with a separate branch to make it conditional. For example, a branch instruction and a long jump sequence can be fused. The scope of an unconditional jump instruction is also limited. For an arbitrary branch at a distance, using a 3-instruction sequence and a scratch register, they can be fused into a single micro-operation. For example, a branch instruction and a long function call sequence can be fused.
[0022] In some implementations, a dynamic fusion predictor can be used to cross-promote macro-operation fusion for instructions that straddle the fetch boundary in an instruction decode buffer. When fetching instructions into the instruction decode buffer, the following situation may exist: there is a prefix of a potentially fusible sequence in the fetch buffer, but the processor will have to wait to fetch additional instructions from memory before knowing for sure whether a fusible sequence exists. In some cases, it may be beneficial to send the existing buffered prefix instructions to execution, while in other cases, it may be beneficial to wait for the remaining instructions in the fusible sequence to be fetched and then fuse them with the buffered instructions. Generally, rushing to execute the prefix or waiting for the trailing instructions may offer performance or power consumption advantages. A fixed policy may lead to suboptimal performance.
[0023] For example, a dynamic "beneficial fusion" predictor can be utilized to inform the processor whether to defer execution of the current one or more instructions in the fetch buffer and wait until additional instructions have been fetched. In some implementations, the fusion predictor is queried and updated only when one or more of the buffered instructions in a potential fusion sequence may have been sent to execution (i.e., execution resources are available); otherwise, the fusion predictor is neither queried nor updated.
[0024] For example, one of several forms can be used to index and / or tag fusion predictor entries, such as indexing by program counter; indexing by a hash of the current program counter and program counter history; tagged, where each entry is tagged with a program counter; or untagged, where each entry is used without regard to the program counter. For example, the program counter used to index the fusion predictor can be the program counter for fetching the last set of instructions, or the program counter of the potential fusion prefix, or the program counter of the next set to be fetched. For example, an entry in the fusion predictor may contain a K-bit counter (K >= 1) to provide hysteresis. The system can execute instruction sequences correctly regardless of the predictions made by the beneficial fusion predictor, so an error prediction recovery mechanism can be omitted from the system.
[0025] The beneficial fusion predictor can be updated based on a performance model that examines the instructions fetched after a potential fusion sequence to determine whether waiting for these additional instructions will be beneficial. The performance model may include many potential components, such as: 1) Can the newly fetched instructions be fused with the buffered instructions? 2) Can fusion prevent parallel issue of instructions following the fusible sequence in the new fetch group? 3) Are there instructions in the new fetch group that depend on instructions in the buffered fusion prefix, creating stalls that would be eliminated by eagerly executing the prefix instructions?
[0026] As used herein, the term "circuit" refers to an arrangement of electronic components (e.g., transistors, resistors, capacitors, and / or inductors) configured to implement one or more functions. For example, a circuit may include one or more transistors interconnected to form logic gates that jointly implement a logical function.
[0027] Figure 1 FIG. 4 is a block diagram of an example of a system 100 for executing instructions from an instruction set with macro-operation fusion. System 100 includes a memory 102 that stores instructions and an integrated circuit 110 configured to execute the instructions. For example, the integrated circuit may be a processor or a microcontroller. Integrated circuit 110 includes an instruction fetch circuit 112; a program counter register 114; an instruction decode buffer 120 configured to store macro-operations 122 that have been fetched from memory 102; an instruction decoder circuit 130 configured to decode the macro-operations from instruction decode buffer 120 to generate corresponding micro-operations 132, which are passed to one or more execution resource circuits (140, 142, 144, and 146) for execution. For example, integrated circuit 110 may be configured to implement Figure 4 process 400. The correspondence between macro-operations 122 and micro-operations is not always one-to-one. Instruction decoder circuit 130 is configured to fuse certain sequences of macro-operations 122 detected in instruction decode buffer 120 and determine a single equivalent micro-operation 132 to execute using one or more execution resource circuits (140, 142, 144, and 146).
[0028] Instruction fetch circuit 112 is configured to fetch macro-operations from memory 102 and store them in instruction decode buffer 120 while processing macro-operations 122 through the pipeline architecture of integrated circuit 110. Instruction fetch circuit 112 may use a direct memory access (DMA) channel to fetch blocks of macro-operations.
[0029] Program counter register 114 may be configured to store a pointer to the next macro-operation in memory. The program counter value stored in program counter register 114 may be updated based on the execution progress of integrated circuit 110. For example, when an instruction is executed, the program counter may be updated to point to the next instruction to be executed. For example, the program counter may be updated to one of multiple possible values based on the test result of a condition by a control flow instruction. For example, the program counter may be updated to a target address.
[0030] Integrated circuit 110 includes an instruction decode buffer 120 configured to store macro-operations fetched from memory 102. For example, the instruction decode buffer 120 may have a depth (e.g., 4, 8, 12, 16, or 24 instructions) that aids the pipelining and / or superscalar architecture of the integrated circuit 110. The macro-operations may be members of an instruction set supported by the integrated circuit 110 (e.g., RISC-V instruction set, x86 instruction set, ARM instruction set, or MIPS instruction set).
[0031] Integrated circuit 110 includes one or more execution resource circuits (140, 142, 144, and 146) configured to execute micro-operations to support an instruction set including macro-operations. For example, the instruction set may be the RISC-V instruction set. For example, one or more execution resource circuits (140, 142, 144, and 146) may include adders, shift registers, multipliers, and / or floating-point units. One or more execution resource circuits (140, 142, 144, and 146) may update the state of the integrated circuit 110 based on the results of executing the micro-operations, including internal registers and / or flags or status bits ( Figure 1 not explicitly shown in). The results of executing the micro-operations may also be written to the memory 102 (e.g., during a subsequent stage of pipelined execution).
[0032] Integrated circuit 110 includes an instruction decoder circuit 130 configured to decode the macro-operations 122 in the instruction decode buffer 120. The instruction decode buffer 120 may convert the macro-operations into corresponding micro-operations 132 that are internally executed by the integrated circuit using one or more execution resource circuits (140, 142, 144, and 146). The instruction decoder circuit 130 is configured to implement macro-operation fusion, where multiple macro-operations are converted into a single micro-operation for execution.
[0033] For example, the instruction decoder circuit 130 may be configured to detect a sequence of macro-operations stored in the instruction decode buffer 120. For example, detecting a sequence of macro-operations may include detecting a sequence of operation codes as part of the respective macro-operations. The sequence of macro-operations may include control-flow macro-operations (e.g., branch instructions or call instructions) and one or more additional macro-operations. The instruction decoder circuit 130 may determine a micro-operation equivalent to the detected sequence of macro-operations. The instruction decoder circuit 130 may forward the micro-operation to at least one of the one or more execution resource circuits (140, 142, 144, and 146) for execution. In some implementations, the control-flow macro-operation is a branch instruction, while the micro-operation is not a branch instruction.
[0034] For example, a sequence of macro operations can include an unconditional jump and one or more macro operations to be skipped, and the micro operation can be a NOP that advances the program counter to the unconditional jump target. For example, an unconditional jump with a short forward offset can be fused with one or more subsequent instructions. The unconditional jump and the skipped instructions can be fused into a single non-jump micro operation, which has no impact on the machine other than advancing the program counter by the jump offset. Benefits may include removing the pipeline flush that is typically required to execute a no-operation (NOP) instruction to perform a jump, where the no-operation (NOP) instruction only advances the program counter without a pipeline flush. For example, a sequence of macro operations:
[0035] j target
[0036] add x3,x3,4
[0037] target: <nextinstruction>
[0038] Can be replaced with a fused micro-operation:
[0039] Nop_pc+8#Advance program counter over skipped instruction
[0040] In some embodiments, more than one instruction can be skipped. In certain implementations, one or more target instructions can also be fused into the micro-operations:
[0041] <nextinstruction>_PC + 12
[0042] In some implementations, the sequence of macro operations includes an unconditional jump, one or more macro operations that will be skipped, and a macro operation at the target of the unconditional jump; and the micro operation executes the function of the macro operation at the target of the unconditional jump and advances the program counter to point to the next macro operation after the target of the unconditional jump.
[0043] For example, conditional branches on one or more instructions can be fused. In some implementations, the sequence of macro operations includes a conditional branch and one or more macro operations after the conditional branch and before the target of the conditional branch, and the micro operation advances the program counter to the target of the conditional branch. In some embodiments, the conditional branch is fused with the following instruction such that the combination is executed as a single non-branch instruction. For example, the sequence of macro operations:
[0044] bne x1,x0,target
[0045] addi x3,x3,1
[0046] target:sub x4,x2,x3
[0047] can be replaced with fused micro operations:
[0048] ifeqz_addix3,x3,x1,1 # If x1 == 0, x3 = x3 + 1, else x3 = x3; PC += 8 For example, if the condition is false, the internal micro operation can disable the write to the destination, or it can be defined to always write to the destination, which can simplify the operation of an out-of-order superscalar machine with register renaming. The fused micro operation can be executed as a non-branch instruction, so pipeline flushing can be avoided, and the branch predictor state can also be avoided from being contaminated. For example, the sequence of macro operations:
[0049] bne x2,x3,target
[0050] sub x5,x7,x8
[0051] target: <unrelatedinstruction>
[0052] Can be replaced by fused micro-operations:
[0053] ifeq_sub x5,x7,x8,x2,x3 # If x2 == x3, x5 = x7 - x8, else x5 = x5;
[0054] PC += 8
[0055] In some implementations, the sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and a macro-operation at the target of the conditional branch; while the micro-operations advance the program counter to point to the next macro-operation after the conditional branch target.
[0056] The sequence of macro-operations to be fused may include multiple instructions following a control flow instruction.
[0057] For example, the sequence of macro-operations:
[0058]
[0059] Can be replaced by fused micro-operations:
[0060] ifeqz_sllori x1,x2,2,1 # If x2 == 0 then x1 = (x1 << 2) | 1 else x1 = x1; PC += 12
[0061] For example, a branch instruction and a single jump instruction can be fused. Compared with a conditional branch, an unconditional jump instruction can direct the instruction to a farther place, but sometimes a conditional branch to a distant target is desired. This can be done through a sequence of instructions that can be fused internally by the processor. In some implementations, the sequence of macro-operations includes a conditional branch followed by an unconditional jump. For example, the sequence of macro-operations:
[0062]
[0063]
[0064] Can be replaced by fused micro-operations:
[0065] jne x8,x9,target # If x8!= x9, then PC = target else PC += 8
[0066] For example, a branch instruction and a function call sequence can be fused. A function call instruction is not conditional, so it can be paired with a separate branch to make it conditional. In some implementations, the sequence of macro operations includes a conditional branch followed by a jump and link. For example, the sequence of macro operations:
[0067]
[0068] can be replaced with fused micro-operations:
[0069] jalez x1,x8,target#If x8 == 0,then x1 = PC+6,PC = subroutine
[0070] #else x1 = x1,PC = PC+6
[0071] For example, a branch instruction and a long jump sequence can be fused. The scope of an unconditional jump instruction is also limited. For arbitrary branching, a sequence of 3 instructions and a scratch register can be used, which can be fused into a single micro-operation. In some implementations, the sequence of macro operations includes a conditional branch followed by a pair of macro operations that implement a long jump and link. For example, the sequence of macro operations:
[0072]
[0073] can be replaced with fused micro-operations:
[0074] jnez_far x6,x8,target_hi,target#If x8 != 0,then
[0075] #x6 = target_hi,PC = target
[0076] #else x6 = x6,PC = PC+10
[0077] For example, a branch instruction and a long function call sequence can be fused. In some implementations, the sequence of macro operations includes a conditional branch followed by a pair of macro operations that implement a long unconditional jump. For example, the sequence of macro operations:
[0078]
[0079]
[0080] can be replaced with fused micro-operations:
[0081] jalgez_far x1,x8,target#If x8 >= 0,then
[0082] #x1 = PC + 12, PC = target
[0083] #else x1 = x1, PC = PC + 12
[0084] The instruction decoder circuit 130 may be configured to detect and fuse sequences of multiple different macro-operations including control flow instructions, such as the sequences of macro-operations described above.
[0085] In some implementations ( Figure 1 not shown), the memory 102 may be included in the integrated circuit 110.
[0086] Figure 2 is a block diagram of an example of a system 200 for using fused speculation to execute instructions from an instruction set with macro-operation fusion. System 200 is similar to Figure 1 system 100, where a fused predictor circuit 210 is added, and the fused predictor circuit 210 is configured to facilitate the detection and beneficial fusion of candidate sequences of macro-operations. For example, system 200 may be used to implement Figure 4 process 400. For example, system 200 may be used to implement Figure 4 process 500. For example, the fused predictor circuit 210 may include Figure 3 the fused predictor circuit 310.
[0087] System 200 includes a fused predictor circuit 210 configured to detect a prefix of a sequence of macro-operations in an instruction decode buffer. For example, when the instruction decoder circuit 130 is configured to detect a macro-instruction sequence consisting of instructions 1 to N (e.g., N = 2, 3, 4, or 5) when it appears in the instruction decode buffer 120, the fused predictor circuit 210 may be configured to detect a prefix including one or more macro-operation instructions 1 to m when it appears in the instruction decode buffer 120, where 1 <= m < N.
[0088] The fused predictor circuit 210 is configured to determine whether to predict the completion and fusion of a sequence of macro-operations upon the next fetch of a macro-operation from the memory. For example, a prediction counter table maintained by the fused predictor circuit 210 may be used to determine the prediction. The prediction counter may be used as an estimate of the likelihood that the prefix will be part of a sequence of completed and fused macro-operations. For example, the prediction counter may be a K-bit counter with K > 1 (e.g., K = 2) to provide some hysteresis. In some implementations, the prediction counter table is indexed by the program counter stored in the program counter register 114. In some implementations, the prediction counter table is tagged with program counter values.
[0089] Maintaining the prediction counter table may include updating the prediction counter after detecting the corresponding prefix and fetching the next instruction set from the memory. For example, the fused predictor circuit 210 may be configured to update the prediction counter table based on whether a sequence of macro-operations is completed by the next fetch of a macro-operation from the memory. For example, the fused predictor circuit 210 may be configured to update the prediction counter table based on whether there is an instruction in the next fetch that depends on the instruction in the prefix. For example, the fused predictor circuit 210 may be configured to update the prediction counter table based on whether fusion will prevent the parallel issue of instructions following a fusible sequence in the next fetch group.
[0090] The fused predictor circuit 210 is configured to, based on the prediction, delay the execution of the prefix until after the next fetch to enable the fusion of the sequence of macro-operations, or start executing the prefix before the next fetch and abandon any possible fusion of the sequence containing the prefix.
[0091] In some implementations ( Figure 2 not shown), the fused predictor circuit 210 is implemented as part of the instruction decoder circuit 130.
[0092] Figure 3 is a block diagram of an example of a system 300 for fused prediction. The system 300 includes an instruction decode buffer 120 and a fused predictor circuit 310. The fused predictor circuit 310 may be configured to examine the macro-operation instructions in the instruction decode buffer 120 to determine a prediction 332 as to whether a sequence of macro-operations including a prefix will be completed and fused by the next fetch of a macro-operation from the memory. The fused predictor circuit 310 includes a prefix detector circuit 320, a prediction determination circuit 330, a prediction counter table 340, and a prediction update circuit 350. The fused predictor circuit 310 may also be configured to examine the macro-operation instructions in the instruction decode buffer 120 to maintain the prediction counter table 340. For example, the system 300 may be used as part of a larger system (e.g., Figure 2 system 200) to implement Figure 5 process 500.
[0093] The fused predictor circuit 310 includes a prefix detector circuit 320 configured to detect a prefix of a sequence of macro-operations in the instruction decode buffer 120. For example, where the instruction decoder (e.g., instruction decoder circuit 130) is configured to detect a sequence of macro-operation instructions consisting of instructions 1 to N (e.g., N = 2, 3, 4, or 5) when a macro-instruction appears in the instruction decode buffer 120, the prefix detector circuit 320 may be configured to detect a prefix including one or more macro-operation instructions 1 to m, where 1 <= m < N, when it appears in the instruction decode buffer 120. For example, the prefix detector circuit 320 may include a network of logic gates configured to set a flag when an m-bit sequence of operation codes corresponding to the prefix is read in the last m macro-operations stored in the instruction buffer.
[0094] The fused predictor circuit 310 includes a prediction determination circuit 330 configured to determine whether to predict 332 a sequence of macro-operations that are completed and fused upon the next fetch from memory. For example, the prediction 332 may include a binary value indicating whether fusion with the detected prefix is expected to occur after the next fetch of a macro-operation. For example, the prediction 332 may include an identifier of the prefix that has been detected. The prediction 332 may be determined by looking up a corresponding prediction counter in the prediction counter table 340 and based on the value of the prediction counter. The prediction counter may be used as an estimate of the likelihood that the prefix will be part of the sequence of macro-operations that are completed and fused. For example, the prediction counter stored in the prediction counter table 340 may be a K-bit counter, where K > 1 (e.g., K = 2) to provide some hysteresis. For example, if the corresponding prediction counter has a current value >= 2^K (e.g., the last bit of the counter is 1), the prediction 332 may be determined to be true, otherwise false. For example, the prediction determination circuit 330 may determine the binary part of the prediction as the most significant bit of the corresponding K-bit prediction counter in the prediction counter table 340.
[0095] In some implementations, the prediction counter table 340 is indexed by the program counter. In some implementations, the prediction counter table 340 is indexed by a hash of the program counter and the program counter history. In some implementations, the prediction counter table 340 is tagged with the program counter value. For example, the program counter used to index the prediction counter table 340 can be the program counter for fetching the last set of instructions or potential fused prefixes, or the program counter for the next set to be fetched. In some implementations, the prediction counter table 340 is untagged when entries are used without considering the program counter. In some implementations, when looking for multiple sequences of macro-operations and / or prefixes for potential fusion, the prediction counter table 340 can be tagged or indexed by the identifier of the detected prefix (e.g., the juxtaposition of one or more opcodes for the prefix or an index value associated with the prefix).
[0096] The fusion predictor circuit 310 includes a prediction update circuit 350, which can be configured to maintain the prediction counter table 340. For example, the prediction update circuit 350 can be configured to update the prediction counter table based on whether the next macro-operation fetched from memory completes the sequence of macro-operations. For example, the prediction update circuit 350 can be configured to update the prediction counter table based on whether there are instructions in the next fetch that depend on the instructions in the prefix. For example, the prediction update circuit 350 can be configured to update the prediction counter table based on whether fusion would prevent the parallel issue of instructions following a fusible sequence in the next fetch group. In some implementations, the prediction counter table 340 is queried and updated only when one or more buffered macro-operations of the prefix of a potential fusion sequence may have been sent to execution (i.e., execution resources are available), otherwise the prediction counter table 340 is neither queried nor updated.
[0097] The fusion predictor circuit 310 can delay the execution of the prefix until after the next fetch based on this prediction, in order to enable the fusion of the sequence of macro-operations. For example, delaying execution can include holding one or more macro-operations of the prefix for multiple clock cycles in the decode stage of the pipeline.
[0098] For example, system 300 can be part of a larger system, such as an integrated circuit for executing instructions (e.g., a processor or a microcontroller). Instruction decode buffer 120 can be configured to store macro-operations fetched from memory. The integrated circuit can also include one or more execution resource circuits configured to execute micro-operations to support an instruction set that includes macro-operations (e.g., the RISC V instruction set, the x86 instruction set, the ARM instruction set, or the MIPS instruction set). The integrated circuit can also include an instruction decoder circuit configured to detect a sequence of macro-operations stored in the instruction decode buffer, determine micro-operations equivalent to the detected sequence of macro-operations, and forward the micro-operations to at least one of the one or more execution resource circuits for execution.
[0099] Figure 4 is a flowchart of an example of process 400 for executing instructions from an instruction set with macro-operation fusion. Process 400 includes fetching 410 macro-operations from memory; and detecting 420 a sequence of macro-operations stored in the instruction decode buffer, the sequence of macro-operations including a control-flow macro-operation followed by one or more additional macro-operations; determining 430 micro-operations equivalent to the detected sequence of macro-operations; and forwarding 440 the micro-operations to at least one execution resource circuit for execution. For example, process 400 can be implemented using Figure 1 system 100 of. For example, process 400 can be implemented using Figure 2 system 200 of.
[0100] Process 400 includes fetching 410 macro-operations from memory and storing the macro-operations in an instruction decode buffer (e.g., instruction decode buffer 120). For example, a direct memory access (DMA) channel can be used to fetch (410) a block of macro-operations. The instruction decode buffer can be configured to store macro-operations fetched from memory while processing the macro-operations through the pipeline architecture of the integrated circuit (e.g., a processor or a microcontroller). For example, the instruction decode buffer can have a depth (e.g., 4, 8, 12, 16, or 24 instructions) that facilitates the pipeline and / or superscalar architecture of the integrated circuit. The macro-operations can be members of an instruction set supported by the integrated circuit (e.g., the RISC V instruction set, the x86 instruction set, the ARM instruction set, or the MIPS instruction set).
[0101] Process 400 includes detecting 420 a sequence of macro-operations stored in an instruction decode buffer, the sequence of macro-operations including a control-flow macro-operation followed by one or more additional macro-operations. For example, detecting 420 the sequence of macro-operations can include detecting an opcode sequence that is part of respective macro-operations. The sequence of macro-operations can include a control-flow macro-operation (e.g., a branch instruction or a procedure call instruction) and one or more additional macro-operations. In some implementations, timely detection 420 of the sequence of macro-operations is enabled to facilitate macro-operation fusion by using a fused predictor (e.g., Figure 3 the fused predictor circuit 310 of Figure 5 to first detect a prefix and a delayed prefix of the sequence until the remainder of the sequence of macro-operations is fetched 410 from memory. For example, a process 500 of
[0102] can be implemented to facilitate detection and fusion of a sequence of macro-operations.
[0103] Process 400 includes determining 430 micro-operations equivalent to the detected sequence of macro-operations. For example, a control-flow instruction can effectively be removed from a program where the micro-operations determined 430 do not include control-flow aspects. In some implementations, the control-flow macro-operation is a branch instruction and the micro-operations are not branch instructions. Removing a branch or other control-flow instruction can improve the performance of an integrated circuit executing a program that includes macro-operations. For example, performance can be improved by avoiding pipeline flushes associated with control-flow instructions and / or by avoiding contaminating the branch predictor state.
[0104] j target
[0105] add x3,x3,4
[0106] target: <nextinstruction>
[0107] The micro-operation can be determined 430 as:
[0108] Nop_pc + 8 # Advance program counter over skipped instruction
[0109] In some implementations, more than one instruction can be skipped. In some implementations, one or more target instructions can also be fused into the micro-operation. For example, the micro-operation can be determined 430 as:
[0110] <nextinstruction>_PC + 12
[0111] In some implementations, the sequence of macro-operations includes an unconditional jump, one or more macro-operations that will be skipped, and a macro-operation at the target of the unconditional jump; and the micro-operations execute the function of the macro-operation at the target of the unconditional jump and advance the program counter to point to the next macro-operation after the target of the unconditional jump.
[0112] For example, conditional branches on one or more instructions can be fused. In some implementations, the sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and the micro-operations advance the program counter to the target of the conditional branch. In some implementations, the conditional branch is fused with the following instruction such that the combination is executed as a single non-branch instruction. For example, for the sequence of macro-operations:
[0113] bne x1,x0,target
[0114] addi x3,x3,1
[0115] target:sub x4,x2,x3
[0116] The micro-operations can be determined 430 as:
[0117] ifeqz_addi x3,x3,x1,1 # If x1 == 0, x3 = x3 + 1, else x3 = x3; PC += 8 For example, if the condition is false, the internal micro-operations can disable the write to the destination, or it can be defined to always write to the destination, which can simplify the operation of an out-of-order superscalar machine with register renaming. The fused micro-operations can be executed as a non-branch instruction, so pipeline flushes can be avoided, and the branch predictor state can also be avoided from being contaminated. For example, for the sequence of macro-operations:
[0118] bne x2,x3,target
[0119] sub x5,x7,x8
[0120] target:<unrelated instruction>
[0121] The micro-operations can be determined 430 as:
[0122] ifeq_sub x5,x7,x8,x2,x3 # If x2 == x3, x5 = x7 - x8, else x5 = x5;
[0123] PC += 8
[0124] In some implementations, the sequence of macro operations includes a conditional branch and one or more macro operations after the conditional branch and before the target of the conditional branch, and a macro operation at the target of the conditional branch; and the micro-operations advance the program counter to point to the next macro operation after the conditional branch target.
[0125] The sequence of macro operations to be fused can include multiple instructions following a control flow instruction. For example, for the sequence of macro operations:
[0126] bne x2,x0,target
[0127] sllix1,x1,2
[0128] ori x1,x1,1
[0129] target:<unrelated instruction>
[0130] The micro-operations can be determined430 as:
[0131] ifeqz_sllori x1,x2,2,1#If x2==0then x1=(x1<<2)|1else x1=x1;PC+=12
[0132] For example, a branch instruction and a single jump instruction can be fused. Compared with a conditional branch, an unconditional jump instruction can direct the instruction to a farther place, but sometimes a conditional branch to a long-distance target is desired. This can be done through a sequence of instructions that can be fused internally by the processor. In some implementations, the sequence of macro operations includes a conditional branch followed by an unconditional jump. For example, for the sequence of macro operations:
[0133]
[0134] The micro-operations can be determined430 as:
[0135] jne x8,x9,target#If x8!=x9,then PC=target else PC+=8
[0136] For example, a branch instruction and a function call sequence can be fused. A function call instruction is not conditional, so it can be paired with a separate branch to make it conditional. In some implementations, the sequence of macro operations includes a conditional branch followed by a jump and link. For example, for the sequence of macro operations:
[0137]
[0138] The micro-operations can be determined430 as:
[0139] Branch x1,x8,target # If x8 == 0, then x1 = PC + 6, PC = subroutine
[0140] # else x1 = x1, PC = PC + 6
[0141] For example, a branch instruction and a long jump sequence can be fused. The scope of an unconditional jump instruction is also very limited. For arbitrary branching, three instruction sequences and a scratch register can be used, and they can be fused into a single micro - operation. In some implementations, the sequence of macro - operations includes a conditional branch followed by a pair of macro - operations that implement a long jump and link. For example, for the sequence of macro - operations:
[0142]
[0143] The micro - operation can be determined430 as:
[0144] jnez_far x6,x8,target_hi,target # If x8!= 0, then
[0145] # x6 = target_hi, PC = target
[0146] # else x6 = x6, PC = PC + 10
[0147] For example, a branch instruction and a long function call sequence can be fused. In some implementations, the sequence of macro - operations includes a conditional branch followed by a pair of macro - operations that implement a long unconditional jump. For example, for the sequence of macro - operations:
[0148]
[0149] The micro - operation can be determined430 as:
[0150] jalgez_far x1,x8,target # If x8 >= 0, then
[0151] # x1 = PC + 12, PC = target
[0152] # else x1 = x1, PC = PC + 12
[0153] Process 400 includes forwarding440 the micro - operation to at least one execution resource circuit for execution. At least one execution resource circuit (e.g., Figure 1 140, 142, 144, and / or 146) can be configured to perform micro-operations to support an instruction set including macro-operations. For example, the instruction set can be the RISC V instruction set. For example, at least one execution resource circuit can include an adder, a shift register, a multiplier, and / or a floating-point unit. At least one execution resource circuit can update the state of an integrated circuit (e.g., a processor or a microcontroller) including internal registers and / or flags or status bits that implements process 400 based on the result of executing the micro-operations. The execution result of the micro-operations can also be written to a memory (e.g., during a subsequent stage of pipelined execution).
[0154] Figure 5 is a flowchart of an example of process 500 for predicting beneficial macro-operation fusion. Process 500 includes detecting 510 a prefix of a sequence of macro-operations; determining 520 a prediction as to whether the sequence of macro-operations will complete and fuse in the next fetch of macro-operations from memory; when no fusion is predicted, starting 530 execution of the prefix before fetching 532 the next batch of one or more macro-operations; when fusion is predicted, delaying 540 execution of the prefix until after fetching 542 the next batch of one or more macro-operations; if a sequence of complete macro-operations is detected 545, fusing 548 the sequence of macro-operations including the prefix; and updating 550 a prediction counter table. For example, process 500 can be implemented using Figure 2 the fusion predictor circuit 210. For example, process 500 can be implemented using Figure 3 the fusion predictor circuit 310. Process 500 can be used to facilitate the fusion of sequences of various different types of macro-operations, including sequences that may lack control flow instructions.
[0155] Process 500 includes detecting 510 a prefix of a sequence of macro-operations in an instruction decoding buffer (e.g., instruction decoding buffer 120). For example, in an instruction decoder configured to detect a sequence of macro-operation instructions including instructions 1 to N (e.g., N = 2, 3, 4, or 5) when it appears in the instruction decoding buffer, when it appears in the instruction decoding buffer, detecting 510 a prefix that includes one or more macro-operation instructions 1 to m, where 1 <= m < N. For example, detecting 510 the prefix can include detecting an opcode sequence as part of each macro-operation that is the prefix.
[0156] Process 500 includes determining 520 whether a prediction that a sequence of macro operations will complete and fuse on the next fetch from memory. For example, a prediction counter table maintained by a fusion predictor circuit can be used to determine 520 the prediction. The prediction counter can be used as an estimate of the likelihood that a prefix will be part of a sequence of macro operations that will complete and fuse. For example, the prediction counter can be a K-bit counter with K > 1 (e.g., K = 2) to provide some hysteresis. For example, if the corresponding prediction counter has a current value ≥ 2^K (e.g., the last bit of the counter is 1), the prediction determination 520 can be determined to be yes or true, otherwise determination 520 is determined to be no or false. In some implementations, the prediction counter table is indexed by a program counter. In some implementations, the prediction counter table is indexed by a hash of the program counter and program counter history. In some implementations, the prediction counter table is tagged with program counter values. For example, the program counter used to index the prediction counter table can be the program counter for fetching the last set of instructions or a potential fusion prefix, or the program counter for the next set of instructions to be fetched. In some implementations, the prediction counter table is untagged when entries are used without regard to the program counter.
[0157] Process 500 includes, if (at operation 525) it is predicted that no fusion will occur, starting execution of the prefix 530 before fetching 532 the next batch of one or more macro operations. For example, starting 530 execution of the prefix can include forwarding micro-operation versions of the macro operations of the prefix to one or more execution resources for execution. For example, a direct memory access (DMA) channel can be used to fetch 532 a block of macro operations.
[0158] Process 500 includes, if (at operation 525) it is predicted that fusion will occur, delaying 540 execution of the prefix based on that prediction until after the next fetch to enable fusion of the sequence of macro operations. For example, delaying 540 execution can include holding one or more macro operations of the prefix for a number of clock cycles in the decode stage of the pipeline.
[0159] After fetching 542 the next batch of one or more macro operations, if (at operation 545) a complete sequence of macro operations is detected, the complete sequence of macro operations including the prefix is fused 548 to form a single micro-operation for execution. For example, the process 400 of Figure 4 can be used to fuse 548 the sequence of macro operations. If (at operation 545) a complete sequence of macro operations is not detected, execution proceeds normally starting from the delayed 540 instructions of the prefix.
[0160] Process 500 includes maintaining a prediction counter table for determining 520 predictions. For example, process 500 includes updating 550 the prediction counter table after detecting 510 the next batch of one or more macro-operations for prefix fetch (532 or 542). For example, the prediction counter table 550 can be updated based on whether a sequence of macro-operations is completed by the next batch of macro-operations from memory. For example, the prediction counter table 550 can be updated based on whether there is an instruction in the next fetch that depends on an instruction in the prefix. For example, the prediction counter table 550 can be updated based on whether fusing will prevent the parallel issue of instructions following a fusible sequence in the next fetch group.
[0161] Although the present disclosure has been described in connection with certain embodiments, it should be understood that the present disclosure is not limited to the disclosed embodiments. On the contrary, it is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims, which scope should be accorded the broadest interpretation so as to encompass all such modifications and equivalent structures as are permitted under the law.< / nextinstruction> < / nextinstruction> < / unrelatedinstruction> < / nextinstruction> < / nextinstruction>
Claims
1. An integrated circuit for executing instructions, comprising: One or more execution resource circuits configured to execute micro-operations to support an instruction set including macro-operations, An instruction decode buffer configured to store macro-operations fetched from memory, and An instruction decoder circuit configured to: Detect a sequence of macro-operations stored in the instruction decode buffer, the sequence of macro-operations including a control flow macro-operation followed by one or more additional macro-operations; Determine micro-operations equivalent to the detected sequence of macro-operations; And Forward the micro-operations to at least one of the one or more execution resource circuits for execution, and A fusion predictor circuit configured to: Detect a prefix of the sequence of macro-operations in the instruction decode buffer; Determine a prediction as to whether the sequence of macro-operations will be completed and fused in the next fetch of macro-operations from memory; And Based on the prediction, delay execution of the prefix until after the next fetch, enabling fusion of the sequence of macro-operations.
2. The integrated circuit according to claim 1, wherein, The control flow macro-operation is a branch instruction and the micro-operations are not branch instructions.
3. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes an unconditional jump and one or more macro-operations to be skipped, and the micro-operations are NOPs that advance the program counter to the target of the unconditional jump.
4. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes an unconditional jump, one or more macro-operations to be skipped, and a macro-operation at the target of the unconditional jump; And the micro-operations execute the function of the macro-operation at the target of the unconditional jump and advance the program counter to point to the next macro-operation after the target of the unconditional jump.
5. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and the micro-operations advance the program counter to the target of the conditional branch.
6. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and a macro-operation at the target of the conditional branch; And the micro-operations advance the program counter to point to the next macro-operation after the target of the conditional branch.
7. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch followed by an unconditional jump.
8. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch followed by a jump and link.
9. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch followed by a pair of macro-operations that implement a long jump and link.
10. The integrated circuit according to claim 1, wherein, The sequence of macro-operations includes a conditional branch followed by a pair of macro-operations that implement a long unconditional jump.
11. The integrated circuit according to claim 1, wherein, The fusion predictor circuit is configured to maintain a prediction counter table, wherein the prediction counter table is used to determine the prediction.
12. The integrated circuit according to claim 11, wherein, The prediction counter table is indexed by the program counter.
13. The integrated circuit according to claim 11, wherein, The prediction counter table is marked with program counter values.
14. The integrated circuit according to any one of claims 11 to 13, wherein, The fusion predictor circuit is configured to update the prediction counter table based on whether the sequence of macro-operations is completed by the next fetch of macro-operations from memory.
15. The integrated circuit according to any one of claims 11 to 13, wherein, The fusion predictor circuit is configured to update the prediction counter table based on whether there is an instruction in the next fetch that depends on the instruction in the prefix.
16. The integrated circuit according to any one of claims 1 to 13, wherein, The instruction set is the RISC-V instruction set.
17. A method, comprising: Fetch a macro-operation from memory and store the macro-operation in the instruction decoding buffer; Detect a sequence of macro-operations stored in the instruction decoding buffer, the sequence of macro-operations including a control flow macro-operation followed by one or more additional macro-operations; Determine micro-operations equivalent to the detected sequence of macro-operations; Forward the micro-operations to at least one execution resource circuit for execution; Detect a prefix of the sequence of macro-operations in the instruction decoding buffer; Determine a prediction as to whether the sequence of macro-operations will be completed and fused in the next fetch of macro-operations from memory; And Based on the prediction, delay the execution of the prefix until after the next fetch, enabling fusion of the sequence of macro-operations.
18. The method according to claim 17, wherein, The control flow macro-operation is a branch instruction, and the micro-operations are not branch instructions.
19. The method according to claim 17, wherein, The sequence of macro-operations includes an unconditional jump and one or more macro-operations to be skipped, and the micro-operations are NOPs that advance the program counter to the target of the unconditional jump.
20. The method according to claim 17, wherein, The sequence of macro-operations includes an unconditional jump, one or more macro-operations to be skipped, and a macro-operation at the target of the unconditional jump; And the micro-operations execute the function of the macro-operation at the target of the unconditional jump and advance the program counter to point to the next macro-operation after the target of the unconditional jump.
21. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and the micro-operations advance the program counter to the target of the conditional branch.
22. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch and one or more macro-operations after the conditional branch and before the target of the conditional branch, and a macro-operation at the target of the conditional branch; And the micro-operations advance the program counter to point to the next macro-operation after the target of the conditional branch.
23. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch followed by an unconditional jump.
24. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch followed by a jump and link.
25. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch followed by a pair of macro-operations that implement a long jump and link.
26. The method according to claim 17, wherein, The sequence of macro-operations includes a conditional branch followed by a pair of macro-operations that implement a long unconditional jump.
27. The method according to claim 17, comprising: Maintain a prediction counter table, where the prediction counter table is used to determine the prediction.
28. The method according to claim 27, wherein, The prediction counter table is indexed by the program counter.
29. The method according to claim 27, wherein, The prediction counter table is marked with program counter values.
30. The method according to any one of claims 27 to 29, comprising: Update the prediction counter table based on whether the sequence of macro-operations is completed by the next fetch of macro-operations from memory.
31. The method according to any one of claims 27 to 29, comprising: Update the prediction counter table based on whether there is an instruction in the next fetch that depends on the instruction in the prefix.
32. An integrated circuit for executing instructions, comprising: One or more execution resource circuits configured to execute micro-operations to support an instruction set including macro-operations; An instruction decoding buffer configured to store macro-operations fetched from a memory; A fusion predictor circuit configured to: Detect a prefix of a sequence of macro-operations in the instruction decoding buffer, Determine a prediction as to whether the sequence of macro-operations will be completed and fused in a next fetch of macro-operations from the memory, and Based on the prediction, delay execution of the prefix until after the next fetch to enable fusion of the sequence of macro-operations; And An instruction decoder circuit configured to: Detect the sequence of macro-operations stored in the instruction decoding buffer, Determine micro-operations equivalent to the detected sequence of macro-operations, and Forward the micro-operations to at least one of the one or more execution resource circuits for execution.
33. The integrated circuit according to claim 32, wherein, The fusion predictor circuit is configured to: maintain a prediction counter table, wherein the prediction counter table is used to determine the prediction.
34. The integrated circuit according to claim 33, wherein, The prediction counter table is indexed by a program counter.
35. The integrated circuit according to claim 33, wherein, The prediction counter table is marked with program counter values.
36. The integrated circuit according to any one of claims 33 to 35, wherein, The fusion predictor circuit is configured to update the prediction counter table based on whether the sequence of macro-operations is completed by the next fetch of macro-operations from the memory.
37. The integrated circuit according to any one of claims 33 to 35, wherein, The fusion predictor circuit is configured to update the prediction counter table based on whether there is an instruction in the next fetch that depends on an instruction in the prefix.
Citation Information
Patent Citations
Instruction and logic to perform fused single cycle increment-compare-jump
CN107077321A
Mechanism in a microprocessor for executing native instructions directly from memory and method
CN1632744A
Microprocessor that fuses MOV / ALU instructions
US20110264896A1