Performing flush recovery using parallel walks of sliced reorder buffers (SROBS)
Through the parallel walking slice reordering buffer (SROBs) mechanism, the problem of low RMT recovery efficiency when pipeline crashes is solved, and a faster and more efficient processor recovery process is achieved.
Patent Information
- Application Number
- JP2023504717
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-08-06
- Filing Date
- 2021-05-03
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2041-05-03
AI Technical Summary
The prior art rename map table (RMT) recovery process is less efficient when processing pipeline crashes, resulting in processing delays and performance degradation.
The parallel-walking slice reordering buffer (SROBs) mechanism is adopted, and the traditional reordering buffer (ROB) is divided into multiple SROB slices. Each SROB slice only tracks uncommitted instructions for the corresponding logical register number, achieving rapid recovery of RMT.
By processing SROB slices in parallel, the efficiency of RMT recovery is significantly improved, processing delay is reduced, and processor performance is improved.
Smart Images

Figure 0007676532000001 
Figure 0007676532000002 
Figure 0007676532000003
Abstract
Description
[Technical field]
[0001] The techniques of this disclosure relate to speculative instruction execution in processor-based devices, and in particular to recovering from pipeline flushes that may occur during the execution of speculative instructions. [Background technology]
[0002] One common feature of typical modern processor devices is branch prediction, which uses mechanisms to predict the direction of branches resulting from control flow instructions (such as conditional branch instructions) and to allow speculative execution of instructions along the predicted execution path. Because some branch predictions will inevitably be incorrect, such branch prediction mechanisms also include hardware to recover from the effects of speculative instruction execution resulting from a mispredicted branch. To achieve this recovery, the branch prediction mechanism must "flush" speculatively executed instructions younger than the mispredicted branch instruction from the execution pipeline, undo all updates performed by the speculatively executed instructions in different microarchitectural structures, and restore the original state of these structures that existed before the mispredicted branch instruction. In addition to branch mispredictions, pipeline flushes may also be performed in response to other control hazards. For example, a pipeline flush may need to be performed after an attempt to execute a load or store instruction such that the calculated address of a memory location is invalid or inaccessible.
[0003] One microarchitectural structure affected by flush recovery is the rename map table (RMT). The RMT, provided by processor units that support register renaming, stores the most recent logical-to-physical register mappings used by the processor unit to establish true data dependencies. Because the mappings stored by the RMT may be modified by speculatively executed instructions, recovering from a flush requires that each RMT mapping be restored to its previous mapping state that existed when the instruction that triggered the flush (the "target instruction") underwent a register renaming. Restoring the RMT is a time-sensitive operation, since instructions on the "correct" execution path fetched after a flush cannot proceed through the processor unit's execution pipeline until the RMT has been restored to its previous mapping state. If the RMT's previous mapping state is not restored quickly enough, the processor unit may be forced to stall the execution pipeline.
[0004] Existing techniques for RMT recovery can generally be categorized according to when the recovery process is initiated. In one approach, known as "lazy recovery", the recovery process is not initiated until the target instruction is the oldest uncommitted instruction in the reorder buffer (ROB), a queue that tracks the state of in-flight instructions in program order after register renaming. Once the target instruction is the oldest uncommitted instruction in the ROB, the RMT can be recovered by simply copying the contents of the committed mapping table (CMT) to the RMT. However, while lazy recovery is simple to implement, it can result in significant degradation of processor performance in situations where a large number of older uncommitted instructions are present at the time the flush is initiated.
[0005] Another approach, known as "immediate recovery", involves initiating the recovery process as soon as the flush is initiated. Some immediate recovery mechanisms may utilize an RMT snapshot. An RMT snapshot is created for each branch instruction and may be used to restore the RMT if the corresponding branch instruction is determined to have been mispredicted. The RMT snapshot may be used alone or in conjunction with "walking" the ROB (i.e., sequentially accessing entries in the ROB between the entry for the target instruction and the time the RMT snapshot was taken, undoing the changes made to the RMT by each corresponding instruction). Other immediate recovery techniques may involve using the contents of the CMT as a starting point to walk the ROB from the oldest uncommitted instruction towards the target instruction, undoing the changes made by each corresponding instruction. Yet another immediate recovery technique may involve using the contents of the RMT as a starting point to walk the ROB from the youngest uncommitted instruction towards the target instruction, undoing the changes. In general, under each of these approaches, the performance of the instant recovery mechanism may depend on the number of snapshots required and / or the number of instructions that need to be walked to restore the previous mapping state of the RMT. Summary of the Invention [Problem to be solved by the invention]
[0006] Therefore, a mechanism to more efficiently restore the RMT after a pipeline flush is desirable. [Means for solving the problem]
[0007] Exemplary embodiments disclosed herein include performing flush recovery using a parallel walk of a sliced reorder buffer (SROB). In this regard, in one exemplary embodiment, a processor device includes a register mapping circuit that provides a rename mapping table (RMT). The RMT includes a number of RMT entries, each of which represents a mapping from a logical register number (LRN) to a physical register number (PRN). The register mapping circuit also provides an SROB that includes a number of SROB slices. Each SROB slice corresponds to a respective LRN and includes a number of SROB slice entries. Each SROB slice is functionally similar to a conventional reorder buffer (ROB), except that the SROB slice tracks only uncommitted instructions that write to the LRN corresponding to the SROB slice and maintains those instructions in program order only among themselves. In an exemplary operation, when the register mapping circuit detects an uncommitted instruction that writes to an LRN in the execution pipeline of the processor device, the register mapping circuit allocates an SROB slice entry for the uncommitted instruction in the SROB slice that corresponds to the LRN. Thereafter, when the register mapping circuitry receives a pipeline flush indication from a target instruction in the execution pipeline, the register mapping circuitry restores the plurality of RMT entries in the RMT to their prior mapping states based on a parallel walk of SROB slices of the SROB. Because the walk of the SROB slices is performed in parallel and each SROB slice is likely to contain fewer instructions than a traditional ROB, flush recovery may be accomplished more efficiently than traditional approaches.
[0008] In some embodiments, the number of SROB slice entries in each SROB slice may be the same size as the number of ROB entries in the ROB provided by the register mapping circuitry, or may be smaller than the number of ROB entries in the ROB. In the latter case, if the register mapping circuitry needs to allocate an SROB slice entry for an uncommitted instruction, but there is no SROB slice entry available in the appropriate SROB slice, the register mapping circuitry of some embodiments may initiate a stall of the execution pipeline until an SROB slice entry is available in the SROB slice. In some embodiments, instead of initiating a stall of the execution pipeline, the register mapping circuitry may allocate the oldest SROB slice entry for the uncommitted instruction. Thereafter, if the register mapping circuitry determines that the overwritten contents of the oldest SROB slice entry are needed for flush recovery, the register mapping circuitry may perform a walk of the ROB. Similarly, some embodiments may provide a partially serial ROB (PSROB) in which the oldest SROB slice entry may be evicted before it is allocated for an uncommitted instruction. In such an embodiment, if the register mapping circuitry determines that the overwritten contents of the oldest evicted SROB slice entry are needed for flash recovery, the register mapping circuitry may perform a walk of the PSROB.
[0009] In another exemplary embodiment, a register mapping circuit in a processor device is provided. The register mapping circuit includes an RMT including a plurality of RMT entries each representing a mapping of an LRN of a plurality of LRNs to a PRN of a plurality of PRNs. The register mapping circuit further includes an SROB subdivided into a plurality of SROB slices, each SROB slice corresponding to a respective LRN of the plurality of LRNs and including a plurality of SROB slice entries. The register mapping circuit is configured to detect an uncommitted instruction in an execution pipeline of the processor device, the uncommitted instruction including a write instruction to a destination LRN of the plurality of LRNs. The register mapping circuit is further configured to assign an SROB slice entry of a plurality of SROB slice entries of an SROB slice corresponding to a destination LRN of the plurality of SROB slices of the SROB to the uncommitted instruction. The register mapping circuit is also configured to receive a pipeline flush indication from a target instruction in the execution pipeline. The register mapping circuitry is further configured, in response to receiving the indication of a pipeline flush, to restore the RMT entries to corresponding prior mapping states based on a parallel walk of the SROB slices corresponding to LRNs of the RMT entries.
[0010] In another exemplary embodiment, a method for performing flush recovery using a parallel walk of SROBs is provided. The method includes detecting, by a register mapping circuit of the processor device, an uncommitted instruction in an execution pipeline of the processor device, the uncommitted instruction including a write instruction to a destination LRN of a plurality of LRNs. The method further includes assigning to the uncommitted instruction an SROB slice entry of a plurality of SROB slice entries of a SROB slice of a plurality of SROB slices of an SROB of the processor device corresponding to the destination LRN, each SROB slice of the plurality of SROB slices corresponding to a respective LRN of the plurality of LRNs. The method also includes receiving an indication of a pipeline flush from a target instruction in the execution pipeline. The method further includes, in response to receiving the indication of a pipeline flush, restoring a plurality of RMT entries of an RMT to corresponding a plurality of previous mapping states based on a parallel walk of the plurality of SROB slices corresponding to the LRNs of the plurality of RMT entries.
[0011] In another exemplary embodiment, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium has stored thereon computer-executable instructions that, when executed by a processor device, cause the processor device to detect an uncommitted instruction in an execution pipeline of the processor device. The uncommitted instruction includes a write instruction to a destination LRN of a plurality of LRNs. The computer-executable instructions further cause the processor device to assign an SROB entry of a plurality of SROB slice entries of a SROB slice of a plurality of SROB slices of an SROB of the processor device corresponding to a destination LRN to the uncommitted instruction, where each SROB slice of the plurality of SROB slices corresponds to a respective LRN of the plurality of LRNs. The computer-executable instructions also cause the processor device to receive a pipeline flush indication from a target instruction in the execution pipeline. The computer-executable instructions further cause the processor device to restore a plurality of RMT entries of an RMT to corresponding a plurality of prior mapping states based on a parallel walk of the plurality of SROB slices corresponding to the LRNs of the plurality of RMT entries in response to receiving the pipeline flush indication.
[0012] Those skilled in the art will appreciate the scope of the present disclosure and realize additional embodiments thereof after reading the following detailed description of the preferred embodiments in conjunction with the accompanying drawing figures. [Brief description of the drawings]
[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate several embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0014] [Figure 1] FIG. 2 is a block diagram illustrating an example processor-based apparatus configured to perform flush recovery using parallel walks of a sliced reorder buffer (SROB).
[0015] [Diagram 2] FIG. 2 illustrates example contents of the reorder buffer and SROB of FIG. 1, in accordance with some embodiments.
[0016] [Diagram 3] 2 is a flowchart illustrating an example operation of the register mapping circuit of FIG. 1 to perform flash recovery using a parallel walk of the SROB of FIG. 1, in accordance with some embodiments.
[0017] [Figure 4] 2 is a flowchart illustrating an example operation of the register mapping circuit of FIG. 1 to allocate a new SROB slice entry, according to some embodiments.
[0018] [Diagram 5] 10 is a flowchart illustrating an example operation of the register mapping circuit of FIG. 1 to allocate a new SROB slice entry in embodiments that overwrite an older SROB slice entry, according to some embodiments.
[0019] [Figure 6] 2 is a flowchart illustrating an example operation of the register mapping circuit of FIG. 1 to allocate a new SROB slice entry in an embodiment that ejects old entries to the partial serial reorder buffer (PSROB) of FIG. 1 , according to some embodiments.
[0020] [Figure 7] 2 is a block diagram of an example processor-based device, such as the processor-based device of FIG. 1, configured to perform flash recovery using a parallel walk of SROBs, in accordance with some embodiments. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] Exemplary embodiments disclosed herein include performing flush recovery using a parallel walk of a sliced reorder buffer (SROB). In this regard, in one exemplary embodiment, the processor device includes a register mapping circuit that provides a rename mapping table (RMT). The RMT includes a plurality of RMT entries, each of which represents a mapping from a logical register number (LRN) to a physical register number (PRN). The register mapping circuit also provides an SROB that includes a plurality of SROB slices. Each SROB slice corresponds to a respective LRN and includes a plurality of SROB slice entries. Each SROB slice is functionally similar to a conventional reorder buffer (ROB), except that it tracks only uncommitted instructions that write to the LRN corresponding to that SROB slice and maintains those instructions in program order only among themselves. In an exemplary operation, upon detecting an uncommitted instruction that writes to an LRN in the execution pipeline of the processor device, the register mapping circuit allocates an SROB slice entry for the uncommitted instruction in the SROB slice that corresponds to the LRN. Later, when the register mapping circuit receives a pipeline flush indication from a target instruction in the execution pipeline, the register mapping circuit restores multiple RMT entries in the RMT to their previous mapping state based on a parallel walk of the SROB slices of the SROB. Because the walk of the SROB slices is performed in parallel and each SROB slice is likely to contain fewer instructions than a traditional ROB, flush recovery may be achieved more efficiently than traditional approaches.
[0022] In some embodiments, the number of SROB slice entries in each SROB slice may be the same size as the number of ROB entries in the ROB provided by the register mapping circuitry, or may be smaller than the number of ROB entries in the ROB. In the latter case, if the register mapping circuitry needs to allocate an SROB slice entry for an uncommitted instruction, but there is no SROB slice entry available in the appropriate SROB slice, the register mapping circuitry in some embodiments may initiate a stall of the execution pipeline until an SROB slice entry is available in the SROB slice. In some embodiments, instead of initiating a stall of the execution pipeline, the register mapping circuitry may allocate the oldest SROB slice entry for the uncommitted instruction. Thereafter, if the register mapping circuitry determines that the overwritten contents of the oldest SROB slice entry are needed for flush recovery, the register mapping circuitry may perform a walk of the ROB. Similarly, some embodiments may provide a partial serial ROB (PSROB) in which the oldest SROB slice entry may be evicted before it is allocated for an uncommitted instruction. In such an embodiment, if the register mapping circuitry determines that the overwritten contents of the oldest evicted SROB slice entry are needed for flash recovery, the register mapping circuitry may perform a walk of the PSROB.
[0023] In this regard, FIG. 1 illustrates an exemplary processor-based device 100 that provides a processor unit 102 for processing executable instructions. In some embodiments, the processor unit 102 may be one of multiple processor units of the processor-based device 100. The processor unit 102 of FIG. 1 includes an execution pipeline 104 that includes circuitry configured to perform execution of an instruction stream 106 that includes computer-executable instructions. In the example of FIG. 1, the execution pipeline 104 includes fetch circuitry 108 configured to fetch the instruction stream 106 of executable instructions from an instruction memory 110. The instruction memory 110 may be provided within or as part of a system memory (not shown) of the processor-based device 100, as a non-limiting example. To reduce the latency of the fetch circuitry 108, an instruction cache 112 may be provided within the processor unit 102 to cache instructions fetched from the instruction memory 110. The fetch circuitry 108 of the example of FIG. 1 may fetch instructions to one or more instruction pipelines I, II, III, IV ... 0 ~I N Prior to execution of the fetched instructions in the execution circuitry 114, the instruction pipeline I 0 ~I N are provided across different processing circuits (or "stages") of the execution pipeline 104, which process the fetched instructions in parallel.
[0024] The execution pipeline 104 of FIG. 1 further decodes instructions fetched by the fetch circuitry 108 into decoded instructions, determines the instruction type and action required, and further determines which instruction pipeline the decoded instructions are to be executed in. 0 ~I N The decoded instruction is then passed to the instruction pipeline I. 0 ~I Nand then provided to renaming circuitry 118. The renaming circuitry 118 determines whether register names in any of the decoded instructions should be renamed to avoid register dependencies that may hinder parallel or out-of-order processing of instructions.
[0025] The renaming circuit 118 invokes a renaming map table (RMT) 120 provided by a register mapping circuit 122 to map the logical source register operands and / or logical destination register operands of the decoded instruction to a plurality of physical registers 124(0) through 124(P) (respectively, a plurality of physical register numbers (PRNs) PRN0, PRN1, PRN2, PRN3, PRN4, PRN5, PRN6, PRN7, PRN8, PRN9, PRN10, PRN11, PRN12, PRN13, PRN14, PRN15, PRN16, PRN17, PRN18, PRN19, PRN20, PRN21, PRN22, PRN23, PRN24, PRN25, PRN26, PRN27, PRN28, PRN39, PRN40, PRN41, PRN42, PRN43, PRN44, PRN45, PRN46, PRN47, PRN48, PRN49, PRN50, PRN51, PRN52, PRN53, PRN64, PRN65, PRN70, PRN81, PRN92, PRN103, PRN114, PRN125, PRN135, PRN146, PRN154, PRN165, PRN175, PRN186, PRN2155, PRN2256, PRN23, PRN24, PRN257, PRN258, PRN39, PRN41, PRN51, PRN659, PRN70, PRN82, PRN93, PRN194, PRN195, PRN196, PRN196, PRN246, PRN259, PRN260, PRN275, PRN39 1 , …PRN P Each physical register 124(0)-124(P) in PRF 126 is configured to store data for a source register operand and / or a destination register operand of a decoded instruction. RMT 120 includes a number of RMT entries 128(0)-128(L), each of which is associated with a number of logical register numbers (LRNs). 0 ~LRN LRMT entries 128(0)-128(L) correspond to respective ones of physical registers 124(0)-124(P) in PRF 126. RMT entries 128(0)-128(L) are configured to store information in the form of address pointers to physical registers of a plurality of physical registers 124(0)-124(P) in PRF 126. In some embodiments, RMT entries 128(0)-128(L) are also associated with respective program order identifiers 130(0)-130(L), each of which provides an indication of the program order location of an instruction that caused the logical register to physical register mapping represented by RMT entry 128(0)-128(L) to be created. In the event that a flush occurs, program order identifiers 130(0)-130(L) may be used to determine which of RMT entries 128(0)-128(L) have been updated by speculatively executed instructions that are older than the target instruction and therefore should be restored to their previous mapping state.
[0026] The execution pipeline 104 of FIG. 1 includes logical register numbers LRNs, which are shown as source register operands of the decoded instructions. 0 ~LRN LThe execution pipeline 104 also includes a register access circuit ("RACC circuit") 132 configured to access one of the physical registers 124(0)-124(P) in the PRF 126 designated by a mapping entry of the RMT entries 128(0)-128(L) corresponding to one of the physical registers 124(0)-124(P) in the PRF 126. The RACC circuit 132 retrieves values in the PRF 126 generated by instructions previously executed in the execution circuit 114. The execution pipeline 104 further provides a scheduler circuit ("SCHED circuit") 134 configured to store decoded instructions in a reservation entry (not shown) until all source register operands for the decoded instruction are available. The execution pipeline 104 is further provided with a write circuit 136 for writing (i.e., committing) generated values from the executed instructions back to a memory, such as the PRF 126, a data cache memory system (not shown), or a main memory (not shown). It should be understood that in some embodiments, the elements of the execution pipeline 104 may be provided in a different configuration or order than that shown in FIG. 1. For example, according to some embodiments, the register access circuitry 132 may be after the scheduler circuitry 134 in the execution pipeline 104, rather than before the scheduler circuitry 134 as shown in FIG.
[0027] 1 also provides branch prediction circuitry 138. Branch prediction circuitry 138 is configured to speculatively predict the outcome of a condition of a fetched conditional flow control instruction (not shown), such as a conditional branch instruction, that controls which path in the instruction control flow paths of instruction stream 106 is fetched into instruction pipelines I0-IN for execution. In an accurate speculative prediction method, the condition of the fetched conditional flow control instruction does not need to be resolved in execution by execution circuitry 114 before execution pipeline 104 can continue processing the speculatively fetched instructions.
[0028] However, if, when the conditional flow control instruction is executed in execution circuitry 114, it is determined that the condition of the conditional flow control instruction is mispredicted, then the speculatively fetched instructions following the mispredicted conditional flow instruction (i.e., the target instruction) in execution pipeline 104 are flushed because the direction of program flow is not as predicted and does not include processing of the speculatively fetched instructions. When a flush occurs (e.g., as a result of a branch misprediction), register mapping circuitry 122 maps the instruction pipeline I instruction of execution pipeline 104 after the target instruction. 0 ~I N 128(L) in the RMT 120).
[0029] To facilitate restoration of a previous mapping state of the RMT 120, the register mapping circuitry 122 provides a reorder buffer (ROB) 140 that includes a number of ROB entries 142(0)-142(R) that are assigned to "in-flight" instructions (i.e., "uncommitted instructions") that are being processed by the execution pipeline 104 but not yet committed. The ROB entries 142(0)-142(R) are assigned to the uncommitted instructions sequentially in program order. Logical Register Number (LRN) by the RMT 120 0 ~LRN LInformation regarding the instruction's modification of the logical-to-physical register mappings (i.e., "register mapping information") is stored in association with each ROB entry 142(0)-142(R) assigned to that instruction. The register mapping information stored by RMT 120 for uncommitted instructions may be used in accordance with conventional techniques to accomplish recovery in response to a flush. Register mapping circuitry 122 of FIG. 1 also includes a committed map table (CMT) 144 that provides a number of mapping entries 146(0)-146(L) in which the logical-to-physical register mappings resulting from committed instructions are stored. CMT 144 is updated only when instructions are committed, and as a result, is not modified in response to a flush.
[0030] As previously discussed, conventional techniques for restoring the RMT 120 to a previous mapping state after a flush (including delayed and immediate recovery techniques) can be inefficient in situations where a large number of older uncommitted instructions are involved at the time of the flush. Accordingly, exemplary embodiments disclosed herein provide a sliced ROB (SROB) 148. The SROB 148 is subdivided into multiple SROB slices ("slices") 150(0)-150(L), each of which contains multiple SROB slice entries, such as SROB slice entries 152(0)-152(X), 154(0)-154(X), etc. Each SROB slice 150(0)-150(L) functions in a manner similar to the ROB 140, except that each SROB slice 150(0)-150(L) contains multiple LRNs. 0 ~LRN L 1, and tracks only uncommitted instructions that write to the destination LRN corresponding to that SROB slice 150(0)-150(L). For example, SROB slice 150(0) in the example of FIG. 0 , and therefore the destination LRN LRN 0SROB slice entries 152(0)-152(X) and 154(0)-154(X) store the same data for each uncommitted instruction as ROB entries 142(0)-142(R) and are allocated sequentially in program order with respect to other SROB slice entries in the same SROB slice 150(0)-150(L).
[0031] In an exemplary operation, when the register mapping circuitry 122 detects an uncommitted instruction in the execution pipeline 104 that writes to a destination LRN, the register mapping circuitry 122 allocates one of the SROB slice entries 152(0)-152(X), 154(0)-154(X) for the uncommitted instruction in the SROB slice 150(0)-150(L) corresponding to the destination LRN. Thereafter, when the register mapping circuitry 122 receives a pipeline flush indication from a target instruction in the execution pipeline 104, the register mapping circuitry 122 restores the multiple RMT entries 128(0)-128(L) of the RMT 120 to their previous mapping state based on a parallel walk of the SROB slices 150(0)-150(L) of the SROB 148. For example, the register mapping circuitry 122 may perform a walk of each SROB slice 150(0)-150(L) in parallel. This is by accessing SROB slice entries 152(0)-152(X), 154(0)-154(X) corresponding to uncommitted instructions younger than the target instruction that caused the flush, and using the data stored in each SROB slice entry 152(0)-152(X), 154(0)-154(X) to undo the changes made by said uncommitted instructions to the RMT entries 128(0)-128(L) corresponding to the LRN for each SROB slice 150(0)-150(L). Because the walk of SROB slices 150(0)-150(L) is performed in parallel, and each SROB slice 150(0)-150(L) is likely to contain fewer instructions than the ROB 140, flush recovery may be accomplished more efficiently than with conventional approaches.
[0032] In some embodiments, the number of SROB slice entries 152(0)-152(X), 154(0)-154(X) may be the same as the number of ROB entries 142(0)-142(R) (i.e., X=R). Such an embodiment may be such that SROB slices 150(0)-150(L) are allocated to ROB 140 such that all uncommitted write instructions in ROB 140 are allocated to multiple LRNs. 0 ~LRN L SROB 148. The SROB 148 may be large enough to handle situations where multiple LRNs target the same destination LRN, thus providing improved performance. However, the improved performance comes at the cost of increased processor resources required to implement the SROB 148.
[0033] Another embodiment may be such that the respective numbers of SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) may be less than the number of ROB entries 142(0) to 142(R) (i.e., X < R). In such an embodiment, the register mapping circuit 122 provides special handling for situations where any of the SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) are not available for assignment to new uncommitted instructions. According to some embodiments, if any of the SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) are not available for assignment, the register mapping circuit 122 may initiate a stall of the execution pipeline 104 until one of the SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) within the appropriate SROB slices 150(0) to 150(L) becomes available for assignment. Once the stall is resolved (i.e., when one of the SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) becomes available), the register mapping circuit 122 then assigns one of the SROB slice entries 152(0) to 152(X) and 154(0) to 154(X) within the appropriate SROB slices 150(0) to 150(L).
[0034] In some embodiments, in response to determining that none of SROB slice entries 152(0)-152(X), 154(0)-154(X) are available for allocation, register mapping circuitry 122 may overwrite the oldest SROB slice entry 152(0)-152(X), 154(0)-154(X) for new uncommitted instructions. If register mapping circuitry 122 later determines that the overwritten contents of the oldest SROB slice entry 152(0)-152(X), 154(0)-154(X) are needed for flash recovery, register mapping circuitry 122 may perform a walk of ROB 140 in a conventional manner to restore RMT entries 128(0)-128(L) of RMT 120 to their previous mapping states. Because walking ROB 140 can incur the same performance penalty as conventional mechanisms for flush recovery, register mapping circuitry 122 in some embodiments may provide a partial serial ROB (PSROB) 156 that includes multiple PSROB entries 158(0)-158(P). PSROB 156 in such an embodiment functions in a manner similar to ROB 140, but only allocates PSROB entries 158(0)-158(P) to store SROB slice entries 152(0)-152(X), 154(0)-154(X) that are evicted from SROB slices 150(0)-150(L) when there are no free SROB slice entries 152(0)-152(X), 154(0)-154(X) to allocate for new uncommitted instructions. If the register mapping circuit 122 later determines that the overwritten contents of the oldest SROB slice entries 152(0)-152(X), 154(0)-154(X) are needed for flash recovery, the register mapping circuit 122 can perform a walk of the PSROB 156 to restore the RMT entries 128(0)-128(L) of the RMT 120 to their previous mapping states.
[0035] FIG. 2 is provided to illustrate example contents of ROB 140 and SROB 148 of FIG. 1, according to some embodiments. In FIG. 2, register mapping circuit 122, ROB 140, and SROB 148 of FIG. 1 are shown. ROB 140 includes ROB entries 142(0)-142(6), each of which stores an uncommitted instruction I 0 ~I 6 Corresponds to command I 0 ~I 6 is the destination LRN 0 and L.R.N. 1 2, which in the example of FIG. 2 correspond to SROB slices 150(0) and 150(1) of FIG. 1, respectively. In an exemplary operation, the register mapping circuit 122 writes to the uncommitted instruction I 0 ~I 6 (e.g., execution pipeline 104 of FIG. 1 ) and each uncommitted instruction I 0 ~I 6 Allocate ROB entries 142(0) through 142(6) for command I. 0 , I 2 , I 4 , and I 6 , the register mapping circuit 122 also allocates SROB slice entries 152(0), 154(0), 152(1), and 154(1), respectively.
[0036] Instruction I 0 ~I 6 During execution of target instruction I 3 does not execute as expected (for example, target instruction I 3 is determined to be a mispredicted branch instruction, or the target instruction I 3 (Because the calculated address of the memory location is an invalid or inaccessible load or store instruction). A pipeline flush is triggered and the register mapping circuit 122 maps the target instruction I 3 Receives pipeline flush instruction 200 from the target instruction I 3All instructions younger than (i.e., instruction I 4 ~I 6 ) is flushed from the execution pipeline 104 and the target instruction I 3 This requires that the RMT 120 of FIG. 1 be restored to a previous mapping state that corresponds to the state of the RMT 120 when execution of was attempted.
[0037] Instead of walking ROB 140 in the conventional manner to restore RMT 120, register mapping circuitry 122 performs a parallel walk of SROB slices 150(0) and 150(1) to retrieve the flushed instruction I 4 ~I 6 Identify the SROB slice entry that corresponds to one of the flushed instructions I 4 ~I 6 In the example of FIG. 2, as a result of the parallel walk of SROB slices 150(0) and 150(1), the register mapping circuitry 122 undoes the mapping changes made to RMT 120 by target instruction I 3 Instruction I, which is younger than and is an uncommitted instruction that modified a mapping in RMT 120 4 and I 6 The register mapping circuit 122 then uses the data in the SROB slice entries 152(1) and 154(1) to map the LRN 0 and L.R.N. 1 (e.g., RMT entries 128(0) and 128(1) in FIG. 1) corresponding to the previous mapping state.
[0038] In some examples, the register mapping circuit 122 may use the program order identifiers 130(0)-130(L) of the RMT entries 128(0)-128(L) to determine which LRN 0 ~LRN LFor example, the register mapping circuit 122 determines whether the mapping state of one of the RMT entries 128(0)-128(L) needs to be restored to a previous mapping state based on the program order identifiers 130(0)-130(L). 3 The register mapping circuit 122 may then determine that the LRN of the particular RMT entry was modified by an uncommitted instruction older than 0. The register mapping circuit 122 may then optimize the restoration of the RMT 120 by not performing a walk of the SROB slice 150(0)-150(L) that corresponds to the LRN of that particular RMT entry.
[0039] 3 provides a flowchart 300 illustrating an example operation of the register mapping circuitry 122 of FIG. 1 to perform flush recovery using a parallel walk of the SROB 148 of FIG. 1, according to some embodiments. Elements of FIG. 1 and FIG. 2 are referenced in describing FIG. 3 for clarity. In FIG. 3, the operation begins when the register mapping circuitry 122 of the processor unit 102 maps instruction I of FIG. 2 to instruction 1 in the execution pipeline 104 of the processor unit 102. 0 It starts by detecting uncommitted instructions such as 0 is multiple LRN LRN 0 ~LRN L Of the destination LRN LRN 0 The register mapping circuit 122 selects the destination LRN LRN from among the multiple SROB slices 150(0) to 150(L) of the SROB 148 of the processor unit 102 (block 302). 0 Among the multiple SROB slice entries 152(0) to 152(X) in the SROB slice 150(0) corresponding to the uncommitted instruction I 0 Here, each of the multiple SROB slices 150(0) to 150(L) is assigned to multiple LRNs (block 304). 0 ~LRN LIn some embodiments, the register mapping circuitry 122 also maps a number of ROB entries 142(0)-142(R) of the ROB 140 to a number of corresponding uncommitted instructions (instructions I1-I2 in FIG. 2 ) in the execution pipeline 104 of the processor unit 102. 0 ~I 6 , etc.), where multiple uncommitted instructions I 0 ~I 6 is an uncommitted instruction I 0 (Block 306).
[0040] The register mapping circuit 122 then maps the target instruction I 3 The register mapping circuit 122 receives a pipeline flush indication 200 from the RMT 120 (block 308). In response, the register mapping circuit 122 restores the RMT entries 128(0)-128(L) of the RMT 120 to their corresponding prior mapping states based on a parallel walk of the SROB slices 150(0)-150(L) corresponding to the LRNs of the RMT entries 128(0)-128(L) (block 310). In some embodiments, the register mapping circuit 122 may restore the RMT entries 128(0)-128(L) to their corresponding prior mapping states based further on the program order identifiers 130(0)-130(L) of the RMT entries 128(0)-128(L) (block 312).
[0041] FIG. 4 provides a flowchart 400 to illustrate an exemplary operation of the register mapping circuit 122 of FIG. 1 for allocating new SROB slice entries 152(0)-152(X), 154(0)-154(X) according to some embodiments. For clarity, reference is made to elements of FIG. 1 in describing FIG. 4. In FIG. 4, it is assumed that the number of SROB slice entries 152(0)-152(X), 154(0)-154(X) in the SROB 148 of FIG. 1 is less than the number of ROB entries 142(0)-142(R) in the ROB 140 of FIG. 1. In FIG. 4, the operation begins with the register mapping circuit 122 determining whether an SROB slice entry is available for allocation among the multiple SROB slice entries 152(0)-152(X) in an SROB slice (e.g., SROB slice 150(0)) (block 402). If so, the register mapping circuit 122 allocates an SROB slice entry, such as SROB slice entry 152(0) (block 404).
[0042] However, if the register mapping circuitry 122 determines in decision block 402 that no SROB slice entries are available for allocation, the register mapping circuitry 122 initiates a stall of the execution pipeline 104 of the processor unit 102 (block 406). Once the stall is resolved (i.e., by one of the SROB slice entries 152(0)-152(X) becoming available for allocation), the register mapping circuitry 122 allocates an SROB slice entry, such as SROB slice entry 152(0) (block 408).
[0043] FIG. 5 provides a flowchart 500 illustrating an example operation of the register mapping circuitry 122 of FIG. 1 to allocate new SROB slice entries 152(0)-152(X), 154(0)-154(X) in an embodiment that evicts old entries to the ROB 140 of FIG. 1. Elements of FIG. 1 are referenced in describing FIG. 5 for clarity. In FIG. 5, the number of SROB slice entries 152(0)-152(X), 154(0)-154(X) of the SROB 148 of FIG. 1 is assumed to be less than the number of ROB entries 142(0)-142(R) of the ROB 140 of FIG. 1. The operation of FIG. 5 begins with the register mapping circuitry 122 determining whether an SROB slice entry of the plurality of SROB slice entries 152(0)-152(X) of an SROB slice (e.g., SROB slice 150(0)) is available for allocation (block 502). If so, the register mapping circuit 122 allocates an SROB slice entry, such as SROB slice entry 152(0) (block 504).
[0044] However, if the register mapping circuit 122 determines in decision block 502 that there are no SROB slice entries available for allocation, the register mapping circuit 122 allocates the oldest SROB slice entry, such as SROB slice entry 152(0) (block 506). The register mapping circuit 122 then determines whether the overwritten contents of the oldest SROB slice entry 152(0) are needed for flash recovery in the course of performing flash recovery (block 508). For example, the register mapping circuit 122 determines that the oldest of the remaining SROB slice entries 152(0)-152(X) of SROB slice 150(0) corresponds to an instruction that follows the target instruction in program order. If it is determined that the overwritten contents of the oldest SROB slice entry 152(0) are needed for flash recovery, the register mapping circuit restores the plurality of RMT entries 128(0)-128(L) to the corresponding plurality of prior mapping states, further based on the walk of ROB 140 (block 510). Otherwise, processing continues as described in the embodiments disclosed herein (block 512).
[0045] FIG. 6 illustrates a flowchart 600 to explain an exemplary operation of the register mapping circuit 122 of FIG. 1 for allocating new SROB slice entries 152(0)-152(X), 154(0)-154(X) in an embodiment that evicts old entries to the PSROB 156 of FIG. 1. For clarity, the description of FIG. 6 references elements of FIG. 1. FIG. 6 assumes that the number of SROB slice entries 152(0)-152(X), 154(0)-154(X) in the SROB 148 of FIG. 1 is less than the number of ROB entries 142(0)-142(R) in the ROB 140 of FIG. 1. 6, the operation begins with the register mapping circuit 122 determining whether an SROB slice entry, such as SROB slice entry 152(0)-152(X), of an SROB slice (e.g., SROB slice 150(0)) is available for allocation (block 602). If so, the register mapping circuit 122 allocates an SROB slice entry, such as SROB slice entry 152(0) (block 604).
[0046] However, if the register mapping circuit 122 determines in decision block 402 that there is no SROB slice entry available for allocation, the register mapping circuit 122 evicts the oldest SROB slice entry, such as SROB slice entry 152(0), among the SROB slice entries 152(0)-152(X) of the SROB slice 150(0) to the PSROB 156 (block 606). The register mapping circuit 122 then allocates the oldest SROB slice entry 152(0) (block 608). The register mapping circuit 122 later, in the course of performing flash recovery, determines whether the overwritten contents of the evicted oldest SROB slice entry 152(0) are needed for flash recovery (block 610). If so, the register mapping circuit restores the RMT entries 128(0)-128(L) to their corresponding previous mapping states based further on the walk of the PSROB 156 (block 612). If not, processing continues as described in the embodiments disclosed herein (block 614).
[0047] FIG. 7 is a block diagram of an exemplary processor-based device 700, such as the processor-based device 100 of FIG. 1, that provides exception stack management using stack panic fault exceptions. The processor-based device 700 may be a circuit included in a printed circuit board (PCB), an electronic board card, such as a server, a personal computer, a desktop computer, a laptop computer, a personal digital assistant (PDA), a computing pad, a mobile device, or any other device, and may represent, for example, a server or a user's computer. In this example, the processor-based device 700 includes a processor 702. The processor 702 may represent one or more general-purpose processing circuits, such as a microprocessor, a central processing unit, or the like, and may correspond to the processor device 102 of FIG. 1. The processor 702 is configured to execute processing logic with instructions to perform the operations and steps described herein. In this example, the processor 702 includes an instruction cache 704 for temporary fast access memory storage of instructions, and an instruction processing circuit 710. Instructions fetched or prefetched from memory, such as from a system memory 708 via a system bus 706, are stored in the instruction cache 704. The instruction processing circuitry 710 is configured to process instructions fetched into the instruction cache 704 and process the instructions for execution.
[0048] The processor 702 and the system memory 708 are coupled to a system bus 706, which may interconnect peripheral devices included in the processor-based device 700. As is well known, the processor 702 communicates with these other devices by exchanging address, control, and data information over the system bus 706. For example, the processor 702 may communicate bus transaction requests to a memory controller 712 in the system memory 708, as an example of a peripheral device. Although not shown in FIG. 7, multiple system buses 706 may be provided, with each system bus constituting a different fabric. In this example, the memory controller 712 is configured to provide memory access requests to a memory array 714 in the system memory 708. The memory array 714 is comprised of an array of storage bit cells for storing data. The system memory 708 may be, by way of non-limiting examples, a read-only memory (ROM), a flash memory, a dynamic random access memory (DRAM), such as a synchronous DRAM (SDRAM), and a static memory (e.g., a flash memory, a static random access memory (SRAM), etc.).
[0049] Other devices may also be connected to the system bus 706. As shown in FIG. 7, these devices may include, by way of example, a system memory 708, one or more input devices 716, one or more output devices 718, a modem 724, and one or more display controllers 720. The input devices 716 may include any type of input device, including but not limited to input keys, switches, voice processors, and the like, and the output devices 718 may include any type of output device, including but not limited to audio, video, other visual indicators, and the like. The modem 724 may be any device configured to allow the exchange of data to and from a network 726. The network 726 may be any type of network, including but not limited to a wired or wireless network, a private or public network, a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a BLUETOOTH™ network, and the Internet. The modem 724 may be configured to support any type of communication protocol desired. The processor 702 may access a display controller 720 via the system bus 706 and be configured to control information sent to one or more displays 722. The display 722 may include any type of display, including, but not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, etc.
[0050] The processor-based device 700 of FIG. 7 may include a set of instructions 728, which may be encoded in a reach-based explicit consumer naming model, executed by the processor 702 for any application desired in accordance with the instructions. The instructions 728 may be stored in the system memory 708, the processor 702, and / or the instruction cache 704, as examples of non-transitory computer-readable media 730. The instructions 728 may also reside, completely or at least partially, within the system memory 708 and / or the processor 702 during its execution. The instructions 728 may further be transmitted or received over the network 726 via the modem 724, such that the network 726 includes the computer-readable medium 730.
[0051] Although computer readable medium 730 is shown to be a single medium in the exemplary embodiment, the term "computer readable medium" should be interpreted to include a single medium or multiple media (e.g., a centralized or distributed database and / or associated caches and servers) that store the one or more sets of instructions 728. The term "computer readable medium" should also be interpreted to include a medium that can store, encode, or carry a set of instructions for execution by a processing device that causes the processing device to perform any one or more of the methodologies of the embodiments disclosed herein. The term "computer readable medium" thus includes, but is not limited to, solid-state memory, optical media, and magnetic media.
[0052] The embodiments disclosed herein include various steps that may be formed by hardware components or embodied in machine-executable instructions that may be used to cause a general-purpose or special-purpose processor programmed with the instructions to perform the steps, or may be performed by a combination of hardware and software processes.
[0053] The embodiments disclosed herein may be provided as a computer program product, or software process, which may include a machine-readable medium (or computer-readable media) having instructions stored thereon, which may be used to program a computer system (or other electronic device) to perform a process according to the embodiments disclosed herein. A machine-readable medium includes any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer). For example, machine-readable media include: machine-readable storage media (e.g., ROM, random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.), and the like.
[0054] Unless otherwise indicated, and as is evident from the preceding discussion, discussions utilizing terms such as "processing," "computing," "determining," "displaying," and the like throughout this document are understood to refer to operations and processes of a computer system or similar electronic computing device that manipulate or transform data represented as physical (electronic) quantities in the computer system's registers and memory into other data similarly represented as physical quantities in the computer system's registers or memory or other such information storage, transmission, or display device.
[0055] The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. A variety of systems may be used with programs in accordance with the teachings of the present application, or it may prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the above description. Further, the embodiments described herein are not described with reference to any particular programming language. It will be understood that a variety of programming languages can be used to implement the teachings of the embodiments described herein.
[0056] Those skilled in the art will further appreciate that the various exemplary logic blocks, modules, circuits, and algorithms described in connection with the embodiments disclosed herein may be implemented as electronic hardware, instructions stored in a memory or another computer-readable medium and executed by a processor or other processing device, or a combination of both. The components of the distributed antenna system described herein may be used in any circuit, hardware component, integrated circuit (IC), or IC chip, as examples. The memories disclosed herein may be of any type and size and may be configured to store any type of information desired. To clearly illustrate this interchangeability, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. How such functionality is implemented depends on the particular application, design choices, and / or design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the embodiments of the present application.
[0057] The various example logic blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Additionally, the controller may be a processor. The processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration).
[0058] The embodiments disclosed herein may be embodied in hardware and instructions stored in the hardware, which may reside in, for example, a RAM, a flash memory, a ROM, an Electrically Programmable ROM (EPROM), an Electrically Erasable Programmable ROM (EEPROM), a register, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable medium known in the art. An exemplary storage medium is coupled to the processor, such that the processor can read information from the storage medium and write information to the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a remote station. In the alternative, the processor and the storage medium may reside as discrete components in a remote station, a base station, or a server.
[0059] It should also be noted that the operational steps described in any of the exemplary embodiments herein are described to provide examples and discussion. The described operations may be performed in many different sequences other than the sequence illustrated. Furthermore, an operation described in a single operational step may actually be performed in several different steps. Furthermore, one or more operational steps discussed in the exemplary embodiments may be combined. Those skilled in the art will also appreciate that information and signals may be represented using any of a variety of technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips, as may be referenced throughout the above description, may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0060] Unless expressly stated otherwise, it is not intended that any method described herein be construed as requiring that its steps be performed in a particular order. Thus, unless a method claim actually recites the order in which its steps are to be followed, or unless it is otherwise expressly stated in the claim or description that the steps are limited to a particular order, no particular order is intended to be inferred.
[0061] It will be apparent to those skilled in the art that various modifications and changes can be made without departing from the spirit or scope of the present invention. Since modifications, combinations, subcombinations and variations of the disclosed embodiments incorporating the spirit and substance of the present invention may occur to those skilled in the art, the present invention should be construed as including all within the scope of the appended claims and their equivalents.
Claims
1. 1. A register mapping circuit in a processor device, the register mapping circuit comprising: a rename map table (RMT) including a plurality of RMT entries each representing a mapping of a logical register number (LRN) among a plurality of LRNs to a physical register number (PRN) among a plurality of PRNs; a reorder buffer (ROB) containing multiple ROB entries; a sliced reorder buffer (SROB) subdivided into a plurality of SROB slices, each SROB slice including a plurality of SROB slice entries, each SROB slice configured to track only uncommitted instructions that write to a respective LRN among the plurality of LRNs; The register mapping circuit: detecting an uncommitted instruction in an execution pipeline of the processor device, the uncommitted instruction including a write instruction to a destination LRN among the plurality of LRNs; assigning an SROB slice entry of the plurality of SROB slice entries of an SROB slice of the plurality of SROB slices of the SROB corresponding to the destination LRN to the uncommitted instruction; receiving a pipeline flush indication from a target instruction in the execution pipeline; configured to, in response to receiving the indication of the pipeline flush, restore the RMT entries to corresponding previous mapping states based on sequential accesses performed in parallel to the SROB slices corresponding to the LRNs of the RMT entries; Register mapping circuitry.
2. the plurality of RMT entries including a respective plurality of program order identifiers; the register mapping circuitry is configured to restore the RMT entries to the corresponding prior mapping states further based on the program order identifiers of the RMT entries.
2. The register mapping circuit of claim 1.
3. The register mapping circuitry is further configured to assign the plurality of ROB entries to a corresponding plurality of uncommitted instructions in the execution pipeline of the processor unit, the plurality of uncommitted instructions including the uncommitted instruction.
2. The register mapping circuit of claim 1.
4. 4. The register mapping circuit of claim 3, wherein a number of the plurality of SROB slice entries of each SROB slice of the plurality of SROB slices of the SROB is equal to a number of the plurality of ROB entries.
5. Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; The register mapping circuitry is further configured to allocate the SROB slice entries: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: initiate a stall of the execution pipeline of the processor unit; Allocating the SROB slice entry in response to resolving the stall. By being configured as follows:
4. The register mapping circuit of claim 3.
6. Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; The register mapping circuitry is further configured to allocate the SROB slice entries: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: Allocate the oldest SROB slice entry By being configured as follows:
4. The register mapping circuit of claim 3.
7. the register mapping circuitry being configured to restore the RMT entries to the corresponding prior mapping states; While performing the walk of the SROB slices, determining that the overwritten contents of the oldest SROB slice entry are required for flash recovery; and configured to, in response to determining that the overwritten contents of the oldest SROB slice entry are necessary for flash recovery, restore the RMT entries to the corresponding previous mapping states based on the ROB walk.
7. The register mapping circuit of claim 6.
8. The register mapping circuit further includes a partial serial ROB (PSROB); Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; The register mapping circuitry is further configured to allocate the SROB slice entries: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: evict the oldest SROB slice entry of the plurality of SROB slice entries of the SROB slice to the PSROB; Allocate the oldest SROB slice entry By being configured as follows:
4. The register mapping circuit of claim 3.
9. the register mapping circuitry being configured to restore the RMT entries to the corresponding prior mapping states; While performing the walk of the SROB slices, determining that the overwritten contents of the oldest evicted SROB slice entry are required for flash recovery; and configured to, in response to determining that the overwritten contents of the oldest evicted SROB slice entry are required for flash recovery, restore the RMT entries to the corresponding previous mapping states based on the walk of the PSROB.
9. The register mapping circuit of claim 8.
10. 1. A method for performing flush recovery using a parallel walk of a sliced reorder buffer (SROB), the method comprising: detecting an uncommitted instruction in an execution pipeline of the processor device by a register mapping circuit of the processor device, the uncommitted instruction including a write instruction to a destination logical register number (LRN) of a plurality of LRNs, the register mapping circuit having a reorder buffer (ROB); assigning, to the uncommitted instruction, an SROB slice entry among a plurality of SROB slice entries of a SROB slice corresponding to the destination LRN among a plurality of SROB slices of an SROB of the processor device, wherein each SROB slice of the plurality of SROB slices is configured to track only uncommitted instructions that write to a respective LRN among the plurality of LRNs; receiving a pipeline flush indication from a target instruction in the execution pipeline; and in response to receiving the indication of the pipeline flush, restoring a plurality of rename map table (RMT) entries of an RMT to a corresponding plurality of prior mapping states based on sequential accesses performed in parallel to the plurality of SROB slices corresponding to the LRNs of the plurality of RMT entries. method.
11. The method of claim 10, further comprising allocating a plurality of reorder buffer (ROB) entries of the ROB to a corresponding plurality of uncommitted instructions in the execution pipeline of the processor unit, the plurality of uncommitted instructions including the uncommitted instruction. The method of claim 10.
12. Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; Allocating the SROB slice entry comprises: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: initiate a stall of the execution pipeline of the processor unit; Allocating the SROB slice entry in response to resolving the stall. Including, The method of claim 11.
13. Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; Allocating the SROB slice entry comprises: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: Allocate the oldest SROB slice entry Including, The method of claim 11.
14. Among the plurality of SROB slices of the SROB, a number of the plurality of SROB slice entries of each SROB slice is less than a number of the plurality of ROB entries; Allocating the SROB slice entry comprises: determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice; In response to determining that no SROB slice entries are available for allocation among the plurality of SROB slice entries of the SROB slice: Evicting an oldest SROB slice entry of the plurality of SROB slice entries of the SROB slice to a partial serial ROB (PSROB); Allocate the oldest SROB slice entry Including, The method of claim 11.
15. Restoring the plurality of RMT entries to the corresponding plurality of prior mapping states includes: While performing the walk of the SROB slices, determining that the overwritten contents of the oldest evicted SROB slice entry are required for flash recovery; in response to determining that the overwritten contents of the oldest evicted SROB slice entry are necessary for flash recovery, further comprising: restoring the RMT entries to the corresponding prior mapping states based on the walk of the PSROB. The method of claim 14.
Citation Information
Patent Citations
Recovery device
JP1995056760A
Data-less history buffer with banked restore ports in a register mapper
US20190188140A1