Efficient hybrid free list for fast register name allocation

CN122804219APending Publication Date: 2026-09-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480088657.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

并且,随着PRN的总数增加,找出可用PRN的复杂性也增加,这进一步推动了面积和功率要求

Benefits of technology

[0011]可以实现本说明书中描述的主题的特定实施例,以便实现以下优点中的一个或多个。混合空闲列表实现了FIFO空闲列表和位向量空闲列表两者的益处,同时减轻了它们相应的缺点。例如,混合空闲列表可以在冲刷恢复期间高效地清除经分配PRN。附加地,混合空闲列表包含要求较少存储空间的经简化PRN信息。因此,混合空闲列表实现了原本在等效FIFO空闲列表中将无法实现的面积和功率减少。附加地,混合空闲列表是完全可扩展的,因为随着PRN的数量增加,冲刷恢复不会像其在严格位向量空闲列表中那样变得不成比例地更具挑战性。本说明书中描述的空闲列表还可以高效地分配PRN的组以满足超标量处理器的带宽。例如,空闲列表可以持续地填充输入暂存缓冲区和输出暂存缓冲区,使得在每一个周期上,N个可用且完整的PRN已经排好队。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804219A_ABST
    Figure CN122804219A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatuses for maintaining a hybrid free list for register renaming. One of the methods includes maintaining a hybrid free list representing physical registers available for register renaming, where the hybrid free list includes a plurality of entries having an order, each entry including a pair of bits for each of a plurality of physical registers, where the pair of bits of each entry includes a valid bit and a commit bit respectively representing a valid state and a commit state of the physical register. An in-order allocation of free physical register names is made from the plurality of entries of the free list using a head pointer and a tail pointer of the free list.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Modern out-of-order (OOO) processing utilizes register renaming to remove spurious data dependencies, allowing more instructions to be executed in parallel, which typically involves out-of-order execution of those instructions. In this specification, register renaming refers to assigning a physical register name (PRN) to a logical register name (LRN) within an instruction being executed. Using register renaming, subsequent instructions that write to a logical register can be executed before or in parallel with other instructions that use or read from the same logical register name, because the register renaming process assigns a different underlying physical register to the logical register name.

[0002] The "free list" is a microarchitectural component of the processor that maintains a pool of available physical resources (e.g., physical registers). In this specification, available resources (e.g., physical registers) will be described as being in an idle state or the corresponding PRN being a free PRN. Therefore, the free list can provide free PRNs for allocating physical registers for register renaming. When a PRN is no longer needed, it can be returned to the free list for reuse.

[0003] The size of the free list and the complexity of the control logic are major considerations in processor design, affecting the silicon area, power, and performance of machines using OOO processors.

[0004] Traditionally, there are two main approaches to implementing free lists in OOO processors: using a First-In-First-Out (FIFO) queue and using a bit vector. In a FIFO free list, PRNs are allocated one after another from the beginning of the free list, and when no longer needed, they are returned to the end. Therefore, in a pure FIFO free list, PRNs are allocated based on age, where older PRNs that have not been allocated for a longer period are prioritized over younger PRNs that have been recently allocated. However, since entire PRNs, which can be 10 bits or longer, are stored in the FIFO, returning a PRN to the FIFO free list requires writing a multi-bit PRN entry. Therefore, frequently writing multi-bit PRN entries to the FIFO free list requires large area and power requirements, which poses serious problems for performance, cost, and scalability.

[0005] In contrast, bit-vector free lists use a single bit to represent each PRN and therefore have smaller storage requirements than FIFO free lists. However, bit-vector free lists require very complex control logic for routine tasks. For example, flushing recovery (e.g., after a mispredicted branch) in a bit-vector free list requires significantly more complex logic than in a FIFO free list because the bit-vector free list does not contain age information for each PRN. Furthermore, as the total number of PRNs increases, the complexity of finding available PRNs also increases, further driving up area and power requirements. For instance, when multiple available entries are provided from a bit-vector free list, finding the first available PRN or the first pair of available PRNs becomes more challenging as request frequency and vector length increase, representing a significant scalability bottleneck. Summary of the Invention

[0006] This specification describes techniques related to hybrid free lists for fast and efficient register renaming, which offer the benefits of both FIFO and bit-vector free lists, as well as many other additional advantages. Hybrid free list maintenance requires only a few bits of stored PRN entries. As an example, each PRN entry may have 1) valid bits indicating whether the corresponding PRN is free or has been allocated, and 2) a commit bit indicating whether the PRN is in an architected state, meaning that the data in the corresponding physical register is part of the machine's permanent state at a given point in time, and its corresponding instruction is no longer subject to flushing or re-execution.

[0007] In this specification, a PRN entry is a structure that stores data from which a PRN can be generated based on the position of the PRN entry in the free list. In this specification, describing the free list as a means of generating a PRN means using the position of the PRN entry to generate the PRN, and does not imply that the free list stores the complete PRN itself.

[0008] The hybrid free list maintains a head pointer, which indicates the position in the free list where to search for the next free PRN. In some implementations, the head pointer advances through segments of the free list containing multiple PRN entries rather than individual PRN entries. The hybrid free list also maintains a tail pointer, which indicates the entry corresponding to the youngest free PRN that can be reassigned. PRNs before the head pointer and after the tail pointer can be assigned sequentially, skipping PRNs that have already been assigned but not yet returned.

[0009] The tail pointer is generally used to protect the ordered arrangement of allocated PRN entries. Therefore, when allocating PRN entries, the head pointer is not allowed to lead the tail pointer. By maintaining the head and tail pointers, the processor can essentially perform cyclic allocation of PRNs by sequentially cycling through the free list, skipping PRNs that have already been allocated.

[0010] The head pointer can be used for efficient flush recovery. When recovering a PRN during flush recovery, the head pointer is simply redirected to the flush point, and the valid bits of the PRN between the flush point and the old head pointer are reset, except for PRNs in a structured state (as indicated by their commit bits).

[0011] Specific embodiments of the subject matter described in this specification can be implemented to achieve one or more of the following advantages. The hybrid free list combines the benefits of both FIFO free lists and bit-vector free lists while mitigating their respective disadvantages. For example, the hybrid free list can efficiently clear allocated PRNs during flush recovery. Additionally, the hybrid free list contains simplified PRN information that requires less storage space. Therefore, the hybrid free list achieves area and power reductions that would be impossible in an equivalent FIFO free list. Furthermore, the hybrid free list is fully scalable because flush recovery does not become disproportionately more challenging as the number of PRNs increases compared to a strict bit-vector free list. The free list described in this specification can also efficiently allocate groups of PRNs to meet the bandwidth requirements of a superscalar processor. For example, the free list can continuously fill input and output buffers such that N available and complete PRNs are queued in each cycle.

[0012] The hybrid free list described in this specification can also be used for register renaming of multiple different types of physical registers. This allows the processor to maintain a single free list instead of multiple free lists, saving power and reducing the required silicon area. For example, the free list can be used for register renaming of both integer registers and flag registers. The free list can also be used for register renaming of integer registers and floating-point registers, even if the sizes of the corresponding physical register files are different. Furthermore, even when the free list is used for register renaming of multiple different types of physical registers, the basic scalable flush recovery mechanism remains largely unchanged and can be used to return the names of all types of physical registers to the free list after a flush.

[0013] This specification also describes how to use a free list to borrow already allocated physical register names. This arrangement improves processor efficiency, reduces the likelihood of the processor running out of register space, and all of this does not add any additional naming storage.

[0014] This specification also describes how a processor can improve floating-point instruction bandwidth by introducing a new microarchitecture structure (i.e., a structured floating-point physical register file), which is an auxiliary physical register file used to maintain the values ​​of floating-point registers whose names are held in a structured mapping table.

[0015] This specification also describes techniques for saving processor power through clock gating using a sequential structure of hybrid free lists. Because the free list inherently maintains information about which physical registers are in use during execution, the processor can utilize this information to perform highly targeted clock gating on unused physical registers, saving power with minimal impact on processor performance.

[0016] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the description, drawings, and claims. Attached Figure Description

[0017] Figure 1A This is an overview of the example lifecycle of a PRN in a mixed free list.

[0018] Figure 1B This is an example overview of a mixed free list.

[0019] Figure 2 This is an example procedure for allocating PRNs from a mixed free list.

[0020] Figure 3 This is an example procedure for deallocating PRNs into a mixed free list.

[0021] Figure 4 An example flush recovery using a mixed free list is shown.

[0022] Figure 5 This is a sample procedure for implementing a mixed free list.

[0023] Figure 6 It is a diagram of a processor system with a shared free list, which uses a single namespace of physical register names for two different physical register files.

[0024] Figure 7 It is a diagram illustrating how the system can maintain a single shared free list for integer registers and flag registers.

[0025] Figure 8 This is a flowchart of an example process for maintaining a shared free list.

[0026] Figure 9 PRN borrowing in a system with two physical register sets is shown.

[0027] Figure 10 This is a flowchart of an example procedure for borrowing physical register names.

[0028] Figure 11 An example technique for associating PRN entries for floating-point registers is shown.

[0029] Figure 12 It is a graph showing how the system can maintain a shared free list between integer registers and floating-point registers.

[0030] Figure 13 This is a flowchart of an example procedure for maintaining virtual PRN entries during flush recovery when managing multiple different register sets.

[0031] Figure 14 This is a flowchart of an example procedure for allocating virtual PRNs to manage free lists of multiple different register sets.

[0032] Figure 15 A sliding window for PRNs in the free list is shown.

[0033] Figure 16 The data structure for handling the floating-point PRN is shown.

[0034] Figure 17 This is a flowchart of an example process for managing a virtual PRN for managing free lists of multiple different register sets.

[0035] Figure 18 A mixed free list with 512 entries is shown, which are divided into banks, each with 16 entries.

[0036] Figure 19 This is a flowchart illustrating an example of clock gating of a physical register file based on a free list.

[0037] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation

[0038] Figure 1A This is an overview of a sample lifecycle 100 of the PRN maintained by a hybrid free list. Lifecycle 100 consists of three phases maintained by three corresponding structures: Hybrid Free List Phase 102, Reorder Buffer (ROB) Phase 130, and Architectural Map (AMT) Phase 140. Throughout the OOO process, the PRN cycles between these three phases.

[0039] The PRN begins in an idle state in free list stage 102. The idle PRN can then be allocated from the mixed free list, renamed, and transitioned to ROB stage 130, which stores the mapping between the PRN and its corresponding LRN in instruction order. When the PRN transitions from free list stage 120 to ROB stage 130 or to AMT stage 140, it is generally no longer idle and therefore cannot be allocated to another logical register until it returns to free list stage 120. However, in certain cases described in detail below, the PRN can be reassigned to another logical register before returning to the free list.

[0040] The ROB can also associate each PRN-LRN mapping with a corresponding position in the free list, so that when a flush occurs, the head pointer of the free list can efficiently and quickly jump back to the flush point.

[0041] When the instruction corresponding to an entry in the ROB has completed execution and reached the head of the ROB, the instruction is committed by retiring the PRN to the AMT, which also stores the mapping between the PRN and its corresponding LRN. The data stored in the physical register corresponding to the PRN is now part of the machine's permanent state and is not subject to flushing or re-execution.

[0042] When subsequent instructions are written to the logical register of the PRN maintained in the AMT, the PRN mapping is no longer required, and the PRN is evicted from the AMT and looped back to the hybrid free list, in which it re-enters hybrid free list stage 102.

[0043] A hybrid free list can maintain PRN entries, each with a pair of values ​​(e.g., a pair of bits) to track the valid and committed status of each PRN. The "valid" bit can be used to track the valid status of a PRN, indicating whether the PRN is idle and can be used for renaming. Therefore, the valid status of a PRN can be referred to as being in an idle state or an allocated state.

[0044] The "Commit" bit is used to track the commit status of the PRN. The commit status indicates whether the instructions using the allocated PRN have been committed. Being committed means that the values ​​in the corresponding physical registers have become the permanent state of the program being executed by the processor and will no longer be overwritten or re-executed. Therefore, the commit status of the PRN can be referred to as either committed or uncommitted. The change from uncommitted to committed generally corresponds to the PRN leaving the ROB and entering the AMT.

[0045] Therefore, a PRN or its corresponding PRN entry may be referred to in this specification as having an idle state or an allocated state, which indicates the value of the valid bit of the PRN entry. Similarly, a PRN or its corresponding PRN entry may be referred to as having a committed state or an uncommitted state, which indicates the value of the committed bit of the PRN entry.

[0046] In some cases, the valid and commit bits have binary values, where "1" indicates a set value and "0" indicates a cleared value. A set value for the valid bit (e.g., "1") can indicate that the PRN has been allocated for renaming, while a set value for the commit bit can indicate that the PRN has been retired to AMT. In some cases, a cleared valid bit value (e.g., "0") and a set commit bit value (e.g., "1") indicate an illegal or improper state, as this would mean that the PRN was committed before it was allocated. This is not expected to happen, but if it does, the processor can enter a fault state to restore the integrity of the free list. Figure 1A The table presents example combinations of valid and submitted bit values. In alternative implementations, values ​​considered set and cleared, or high and low, can be reversed.

[0047] In some implementations, PRN availability can be dynamically constructed rather than explicitly represented. For example, the free state of each PRN can be determined using the head and tail pointers of the free list, rather than maintaining an explicit valid bit for each PRN. In this case, the head pointer indicates the region of the free list used to search for a PRN to allocate. Alternatively, the commit state of a PRN can be constructed by checking the AMT instead of maintaining a commit bit. However, when the processor has hundreds of physical registers, it is often faster to simply maintain an explicit commit bit in the free list rather than searching the AMT.

[0048] Figure 1B Here is an example overview 101 of a hybrid free list 132. Hybrid free list 132 comprises a set of PRN entries representing PRNs. Each PRN entry has a valid bit 104 and a commit bit 106. Additionally, hybrid free list 132 includes a head pointer 108 and a tail pointer 110.

[0049] When a PRN needs to be allocated in OOO processing, the mixed free list 132 will use the head pointer 108 to find the PRN entries that can be allocated. Although Figure 1B A PRN entry at head pointer 108 in the hybrid free list 132 is shown, but the head pointer in the hybrid free list can be implemented as a coarse-grained pointer that steps through multiple PRN entries in segments. For example, Figure 1BThe diagram also illustrates how the set of PRN entries in the free list 132 can be implemented as a memory bank 116, with the head pointer stepping one memory bank at a time. In some implementations, the size of each memory bank is based on the processor's bandwidth. For example, if the processor might need to rename ten registers per cycle, the free list could have a memory bank size of ten or more PRN entries to account for PRNs still in the committed state. Typically, each instruction is written to only one logical register, and therefore, the processor's instruction bandwidth typically corresponds to the required PRNs to be generated per cycle.

[0050] When a PRN retires from the ROB, the tail pointer 110 is updated. The tail pointer 110 is updated when the next PRN entry retires from the ROB by setting its valid and committed bits to the state {0,0} or {1,1}. This means that the PRN retires from the ROB without entering the AMT, or the instruction using the PRN is committed and the PRN enters the AMT. Alternatively or additionally, if the tail pointer steps by the memory bank of the PRN entry, the tail pointer is updated when all entries in the next memory bank have the state {0,0} or {1,1}. Therefore, the tail pointer 110 indicates the youngest PRN entry or the memory bank of the PRN entry that can be reallocated.

[0051] In contrast, the head pointer indicates the storage of the oldest PRN that can be allocated, which corresponds to the PRN that has not been used within the longest number of cycles.

[0052] As mentioned above, the system maintains the tail pointer 110 to keep the PRN entries in the free list in order by preventing the head pointer 108 from ever leading the tail pointer 110. This is important for maintaining or deriving age information about the PRN entries, because if the structured PRNs returned to the free list are reallocated, determining the age of the corresponding PRN entries during flushing may be impossible or impractical. Maintaining the tail pointer 110 also improves power performance by preventing the head pointer and temporary buffer from being continuously updated for numerous non-free PRN entries that cannot be allocated.

[0053] When a PRN entry is read from the hybrid free list 102, a decoding process is performed to generate the corresponding PRN based on the PRN entry's position in the free list. This process saves storage space because it allows the free list to efficiently store two bits of each PRN, instead of all bits of each PRN.

[0054] The process of using pointers to iterate through the PRN memory is highly scalable. In other words, the head pointer 108 and the tail pointer 110 can be used to allocate any number of PRNs from the mixed free list without disproportionately increasing processing requirements.

[0055] In some implementations, the free list continuously fills one or more scratch buffers, ensuring that available PRNs are always ready in each cycle. The input scratch buffer receives PRN entries from memory banks indicated by the head pointer. The output scratch buffer is then filled with the full physical register number corresponding to the PRN entry in the input scratch buffer. At either the input or output scratch buffer, allocated PRN entries can be removed (e.g., as indicated by their valid bits). In some cases, the input scratch buffer includes only the valid bit 104 value of each PRN entry identified from the mixed free list 102. In other words, the input scratch buffer can essentially be a bit vector representing which PRNs are free in memory banks indicated by the head pointer. The output scratch buffer is then filled with full PRNs using the positions of the set bits in the input scratch buffer and the position of the head pointer. For example, the output scratch buffer may include memory banks with 16 different PRNs. In other cases, the buffer can be scaled to include different numbers of PRNs and / or bits. References will follow below. Figure 2 The operation of the temporary buffer will be explained in more detail.

[0056] Figure 2 This is a diagram of an example procedure 200 for allocating a PRN from a mixed free list 202. This example procedure, as well as other procedures described herein, can be implemented using a digital logic circuitry system of a processor (e.g., an out-of-order processor).

[0057] In this example, for simplicity, each PRN entry in free list 202 is represented as having a single valid bit, but as described above, mixed free list entries will typically also have a commit bit. The processor includes input buffer 204 and output buffer 206. In some cases, when a PRN memory bank identified by the head pointer is read, the PRNs that can be allocated within that memory bank (e.g., valid bit 104 is cleared, e.g., "0", or an alternative availability state) are expanded into a series of PRNs in input buffer 204 and output buffer 206. This is done by placing the PRN entry from the memory bank indicated by the head pointer into input buffer 204. For example, as... Figure 2 As shown, the input temporary buffer 204 may include 16 bits, where each bit corresponds to a valid bit value from a different PRN entry in the memory bank indicated by the header pointer 108.

[0058] Next, the physical register names corresponding to the PRN entries in input buffer 204 are populated into output buffer 206. The physical register names can be generated from the position of the PRN entry in the free list. In some implementations, the position can be reconstructed from the position of the head pointer plus the offset of the PRN entry in the input or output buffer or in the memory of the PRN entry. Assuming the PRN length is 10 bits, output buffer 204 can include 10 bits of PRN for each free PRN entry in input buffer 204, thus converting the PRN entry into an actual PRN. For example, based on the 16 bits in input buffer 204, the system can generate up to 16 complete 10-bit PRNs in output buffer 206.

[0059] As part of this process, the "holes" in the input buffer caused by the PRN being unavailable due to its valid bits being essentially squeezed out, so that the output buffer only stores PRN entries with the valid bits cleared.

[0060] The processor can then use the available PRNs in the output temporary buffer 204 for renaming, for example, by assigning PRNs to logical register names and storing each assigned PRN to LRN mapping in ROB 302.

[0061] Figure 3 This is an example process 300 that releases a PRN into the mixed free list 102 after it has been retired from the ROB and / or evicted from AMT 304. As described above, the value of the associated PRN entry valid bit 104 can be set to "1" after the PRN has been allocated.

[0062] During OOO processing, the PRN-to-LRN mapping is retained in ROB 302 until the associated instruction is committed, causing the PRN-to-LRN mapping to be retired to AMT or flushed. When a PRN is retired, the commit bit 106 of the associated PRN entry is set to "1" (operation 306), indicating that the instruction for the PRN has been committed and the PRN has been cycled into AMT 304 (operation 308). As discussed above, AMT 304 maintains the PRN-to-LRN mapping for committed PRNs, and a committed PRN entering AMT 304 for the same logical register will evict any previous PRN mapping for that logical register. Therefore, typically, the size of AMT 304 is based on the number of logical registers in the instruction set, while the size of ROB 302 is usually much larger because there are typically many more physical registers than the number of logical register names in the instruction set. When the PRN is evicted from AMT 304, the values ​​of the associated valid bit 104 and commit bit 106 are cleared, for example, returned to "0", thus indicating that the PRN is now available for allocation (operation 310).

[0063] Figure 4 An example flush recovery using the hybrid free list 102 is shown. In some cases, the PRN needs to return from ROB302 to the hybrid free list 102 or be released. This process is called flush recovery. For example, flush recovery can be performed when instructions are speculatively executed along an incorrect control flow path (e.g., a mispredicted branch).

[0064] During flush recovery, the flush point is specified at a point in the mixed free list 102, and the head pointer 108 is redirected to this new point. In some implementations, the flush point is the point where the writer of a particular PRN and all younger PRNs become invalid from the ROB. Any PRN after the flush point and before the original head pointer 108 (e.g., the head pointer 108 before redirection) has its valid 104 bits cleared, for example, cleared to "0", except for PRNs still in use in AMT 304 (as indicated by their commit bit being set).

[0065] In some cases, to support efficient flush recovery, checkpointing operations (operation 402) can be continuously performed on the PRN pointer during the dispatch process. For example, during dispatch, the current head pointer 108 plus an offset value can be associated with each PRN entering ROB 302. Alternatively or additionally, the PRN pointer can be stored in other structures linked to the ROB or in structures that track in-progress instructions (e.g., reservation stations or dispatch queues, to name a few). In other words, the system can use any suitable structure to collect PRNs that have been assigned but not yet committed in AMT, and the corresponding PRN entries will be returned to the free list during flush.

[0066] When flush recovery begins, the current checkpoint is indicated by the oldest PRN to be flushed from ROB 302. The head pointer of the free list can then be reset to the flush point using a pointer to that ROB entry (operation 404). When the head pointer points to a memory bank of a PRN entry, an offset can be used to determine which PRN entries in the memory bank are flushed. In some implementations, the input and output scratch buffers are also cleared, allowing them to be refilled with PRN data from the new flush point. In some cases, the flush recovery process is limited during allocation time to prevent conflicts with ROB 302.

[0067] By maintaining checkpoints, flush recovery can be performed efficiently without exhaustively searching for affected PRNs during the flush. The flush recovery process in the hybrid free list 102 is also highly scalable because the performance of flush recovery is minimally affected as the free list size increases. In other words, even if the free list grows 10 times, the flush recovery process will still be just as fast. This is very different from the bit vector free list, where flush recovery becomes more computationally intensive as the number of PRNs increases and becomes more complex and time-consuming as the ROB size increases. Using checkpoints in the ROB also avoids sidetracking of in-process PRNs (e.g., via the ROB or another structure) to identify and indicate the set of PRNs affected by the flush (as is the case in the bit vector free list).

[0068] Figure 5 This is an example process 500 for implementing a hybrid free list. The process can be performed by a processor configured according to this specification. For convenience, the process will be described as being performed by the system.

[0069] The system uses the free list header pointer to identify PRNs (510) that have an idle status (e.g., a valid bit with a value of "0"). As described above, a PRN in the mixed free list is associated with both a valid bit and a commit bit indicating which stage the PRN is in. The valid bit indicates whether the PRN is currently in use.

[0070] The hybrid free list also maintains a head pointer and a tail pointer. The head pointer indicates the next PRN entry or set of entries to be allocated, while the tail pointer indicates the next PRN to be retired, currently being written to by the oldest in-process instruction being written to the register. In some implementations, the head pointer points to the memory bank of the PRN entry, and the head pointer steps on each memory bank of the PRN entry.

[0071] The system allocates a PRN and modifies its idle state (520). Then, the PRN is allocated and the valid bit is updated to indicate that the PRN is no longer idle; for example, the valid bit value is "1". As described above, once a PRN is allocated from the mixed free list, it is cycled through the ROB, which maintains a sequenced list of instructions for which PRNs have been allocated and their PRN-to-LRN mappings.

[0072] The system retires the PRN and modifies the commit status (530). At commit time, the PRN is retired, and the PRN-to-LRN mapping is stored in the AMT. When this happens, the commit bit of the PRN entry is modified to indicate that the PRN has been retired (e.g., by setting the value of the commit bit to "1"). As described above, when a PRN is transitioned to the AMT, it can evict older PRNs assigned to the same LRN for older instructions.

[0073] The system evicts the PRN from the AMT and modifies the PRN's idle status (540). When the PRN is evicted from the AMT and returned to the mixed free list, both the valid bit and the commit bit are updated to indicate that the PRN is idle; for example, the values ​​of both the valid bit and the commit bit can be set to "0".

[0074] This specification also describes how free lists can be used for various types of physical registers, each with its own physical register namespace. For example, some processors have different physical register files, each with its own namespace. Instructions can reference different logical register names corresponding to different underlying physical register files. While it would be possible to maintain a separate free list for each corresponding physical register file, in many cases this would result in inefficient and redundant hardware.

[0075] Conversely, the techniques described in this specification allow the processor to use a single shared free list to maintain and allocate physical register names for multiple different register files. When an instruction requires a physical register name for register renaming, the free list can allocate the physical register name using a single shared namespace applicable to both types of physical register files. Furthermore, for instructions that reference multiple different types of registers, the free list can assign the same physical register name to both types of logical registers in the instruction.

[0076] A major benefit of using a shared free list for multiple different types of physical register files is that only one flush recovery mechanism needs to be executed during flushing. Additionally, sharing the free list between different types of registers saves hardware resources without significantly impacting performance.

[0077] Figure 6 This is a diagram of a processor system 600 with a shared free list, which uses a single namespace of physical register names for two different physical register files. System 600 includes a first physical register file 610 with physical registers PR1, PR2 through PRN. System 600 also includes a second register file 620 with physical registers FR1, FR2 through FRN.

[0078] System 600 also includes a shared free list 630 that maintains physical register name entries PRN1, PRN2 to PRNN. This single shared namespace PRN entry can be used for register renaming in both the first physical register file 610 and the second physical register file 620.

[0079] Typically, each physical register file has its own Architectural Mapping Table (AMT). Therefore, physical register file 610 will have its own AMT, and the second physical register file 620 will have its own AMT. Since the two physical register files are sharing the same namespace for register renaming, the system can ensure that a physical register name is not returned to the free list 630 unless it is evicted from either AMT or does not exist in either AMT.

[0080] An example of a separate type of physical register file is a special register file. In some processors, special physical registers can be used to maintain information about the processor's state. These special registers are typically maintained in a separate physical register file from the physical register files used for integers or floating-point numbers.

[0081] An example of a special register is the flag register, which is sometimes referred to as the application status register. In this specification, "flag register" will refer to any special register that can be referenced by logical register names in an instruction set that has a physical register file different from an integer register file or a floating-point register file.

[0082] In some processors, instructions can be written to only one logical flag register, but there can be dozens or hundreds of physical flag registers that can be allocated to that single logical flag register. This is because, during OOO execution, flag registers can be used to maintain the conditional states of many different in-progress instruction branches. Therefore, even if the instruction set has only one logical flag register, register renaming can be used to maintain the state of all in-progress conditional branches.

[0083] Figure 7 This is a diagram illustrating how system 700 can maintain a single shared free list 730 for integer registers and flag registers. System 700 includes a set of integer registers 710 and a set of flag registers 720. The system also includes a hybrid free list 730 that maintains a single physical register namespace for both integer registers 710 and flag registers 720. Hybrid free list 730 includes PRN entries, each including a valid bit and a commit bit. Integer register file 710 has its own integer AMT 740. Flag register file 720 also has its own flag AMT 750.

[0084] As described above, the valid bit indicates whether a physical register name is free and can be used for register renaming, or whether the physical register name has already been assigned. The commit bit indicates whether the physical register name is currently still in the AMT. As described above, these two bits determine how the system assigns physical register names and also how the system handles returning physical register names to the free list during flushing, which essentially involves returning PRNs assigned after the flush point that are not yet in the AMT.

[0085] In this example system, the integer register and the flag register share the free list 730 by using the valid bits explicitly represented in the PRN entries of the free list 730.

[0086] System 700 can use the integer commit bit 706 of the PRN entry to track the commit status of the PRN allocated to the integer register. However, for flag PRNs, System 700 can reconstruct the virtual flag commit bit based on the PRN entry and the status of the flag AMT 750. This eliminates the need for the system to maintain a separate free list for the flag register and keeps the size of each PRN entry to two bits.

[0087] In this example, the Flag AMT 750 has only a single entry because, in this example, the Flag register is only a single logical register in the instruction set. Therefore, the physical-to-logical register name mapping in the Flag AMT 750 requires only a single entry. Furthermore, when there is only a single logical flag register, the Flag AMT 750 only needs to store the physical register name, not the complete mapping from physical register name to logical register name. Therefore, the Flag AMT 750 is shown as storing only a single PRN instead of a mapping. In contrast, the Integer AMT 740 has as many integer entries as the Integer logical register, and the Integer AMT 740 stores the mapping between the physical register name and the logical register name for each entry.

[0088] During allocation, physical register names can be assigned based on their valid status. For integer registers and flag registers, this means checking the valid bits in the free list 730. However, to prevent naming conflicts, the system also performs a bitwise OR between the valid status indicated by the PRN entry and the decoded commit bit from the flag AMT 750. This prevents the reassignment of the same PRN until the PRN has been evicted from AMT 740 and 750 or otherwise does not exist in either.

[0089] The decoding module 770 can be used to generate a decoded commit bit for any PRN entry in the free list 730 by checking whether the corresponding PRN is still in the flag AMT 750. In other words, for a given PRN, if the PRN is in the flag AMT 750, the decoding module 770 can generate a 1 for the decoded commit bit, and otherwise, it can generate a 0 for the decoded commit bit.

[0090] Then, based on the decoded commit bit of the PRN, system 700 can generate the virtual valid bit 772 of the PRN using a bitwise OR between 1) the decoded commit bit and 2) the valid bit of the corresponding PRN entry in the free list 730. In practice, this means that when the integer register and the flag register appear in the same instruction, they can be assigned to the same PRN simultaneously.

[0091] Similarly, system 700 can generate the virtual flag commit bit 774 of the PRN by using a bitwise OR between 1) the decoded commit bit and 2) the commit bit of the corresponding PRN entry in the free list 730.

[0092] As described above, the physical register name can become free again when the PRN is no longer in the integer AMT 740 or the flag AMT 750. For the integer AMT 740, the system can use the commit bit of the PRN entry in the free list 730 to determine whether the physical register name is still maintained in the integer AMT 740.

[0093] On the other hand, since the AMT 750 only maintains one logical flag register, the system can save considerable storage space by virtually determining the flag commit bit based on the decoded commit bit. Therefore, for the flag PRN, the system can determine the commit status by performing a bitwise OR operation between 1) the commit bit of the corresponding PRN entry and 2) the decoded commit bit.

[0094] Table 1 summarizes the status of the PRN entries shared between the integer register and the flag register.

[0095]

[0096] Table 1

[0097] The operations for maintaining the shared free list 730 will now be described. During operation 1, physical register names can be assigned by identifying PRN entries with an idle state in the free list. As described above, PRN entries that can be renamed can be identified using an input buffer and an output buffer, which can also be used to generate a PRN based on the free list position of the corresponding PRN entry.

[0098] Since the integer register file 710 and the flag register file 720 are sharing a namespace, if an instruction references a logical register for an integer or a logical register for a flag register, the system simply assigns a physical register name to any free PRN entry in the free list 730.

[0099] In some implementations, the system can assign the same PRN to both the integer register and the flag register. For example, if a single instruction references both the integer logical register and the flag logical register, the system can find an available PRN from the free list and assign the same PRN to both the integer logical register and the flag logical register. Although they are assigned the same physical register name, the values ​​written to the respective logical registers will be stored in different physical registers because the logical register names refer to different physical register files 710 and 720.

[0100] Then, during instruction execution, the PRN-to-LRN mapping is stored in ROB 760. In some implementations, ROB 760 includes fields 711 and 712 indicating whether the mapping is an integer mapping, a flag mapping, or both. Since the system assigns the same physical register name to instructions that reference both the integer register file 710 and the flag register file 720, ROB 760 only needs two bits for each ROB entry—the flag bit 711 and the integer bit 712—to indicate whether the entry is an integer mapping, a flag mapping, or both.

[0101] Operation 2 represents a flush recovery. Since the integer register file 710 and the flag register file 720 are sharing the free list, only a single flush operation is required on the free list 730. As described above, the head pointer of the free list 730 will quickly jump back to the flush point indicated by the checkpoint in ROB 760, as described above. Any allocated PRN entries that did not have their commit bits set after the flush point and before the head pointer will be returned to the free list, for example, by resetting their valid and commit bits.

[0102] Operation 3 indicates a bypass return to the free list 730. This is a special case to avoid unnecessarily storing the PRN in the AMT. For example, if a physical register mapped to a logical register retires in the same cycle as a younger instruction that wrote to the same logical register, there is no need to add the physical register to the AMT, as it will be overwritten by the younger writer. Therefore, the physical register name can bypass the AMT and return directly to the free list. In doing so, the system can clear the valid bits and integer commit bits, for example, by setting them to 0.

[0103] During operation 4, when a PRN retires from ROB 760, if the PRN is only an integer PRN, it will be retired to integer AMT 740. The system can also set the integer commit bit for the corresponding PRN entry, for example, by setting it to 1. If the PRN is only a flag PRN, it will be retired to flag AMT 750. If the PRN is both an integer PRN and a flag PRN, it will be retired to both AMT 740 and 750.

[0104] The PRN can eventually be decommissioned from its corresponding AMT 740 and 750.

[0105] During operation 5, the integer PRN is evicted from the integer AMT 740. The system can clear the valid bit and integer commit bit of the corresponding PRN entry from the free list 730 to indicate that the PRN is no longer in the integer AMT 740.

[0106] However, unless and until the PRN is also not stored in the flag AMT 750, the PRN will not actually become free and available for allocation. As described above, system 700 can enforce this restriction using a virtual valid bit 772, which is a bitwise OR between the valid bit of the PRN entry and the decoded commit bit from the flag AMT 750.

[0107] When a flag PRN is evicted from flag AMT 750, the system can reset the valid bits of the corresponding PRN entry. When this happens, a different PRN will be stored in flag AMT 750, so for that PRN, the decoded commit bit will change from 1 to 0. Therefore, at this time, unless the integer PRN is still in integer AMT 740, the virtual commit bit of the PRN will be reset to 0.

[0108] Figure 8 This is a flowchart of an example process for maintaining a shared free list. The process can be performed by a processor configured according to this specification. For convenience, the process will be described as being performed by the system.

[0109] The system maintains a shared free list for various types of register sets (810). As described above, the processor may have different physical register files, each with its own physical register names. For multiple physical register files, the system may maintain a single physical register namespace. As an example, the system may maintain a single physical register namespace for both the integer register and the flag register, which are maintained in separate physical register files.

[0110] The system generates a first physical register name (820) for a first type of physical register and a second physical register name (830) for a second type of physical register. In other words, even if physical registers are in different physical register files, the system can generate physical register names based on the same shared free list. Furthermore, as described above, if an instruction references a physical register in two physical register files, the system can generate a shared physical register name for the two physical registers in different physical register files.

[0111] The system retires the first physical register name and the second physical register name to different corresponding AMTs (840). As described above, different register files have different AMTs for maintaining the processor's state for committed instructions.

[0112] If the first physical register name is not in the second AMT, the system designates the first physical register name as free (850). In other words, since different physical register files are sharing a namespace, a physical register name should not become free if it is still maintained in one of the AMTs. Therefore, in order to return the first physical register name to the free list, the system may first perform a check to ensure that the physical register name is not still maintained in the second AMT. As described above, in some implementations, the system may perform this check as a bitwise OR between the decoded commit bit and the valid bit in the PRN entry maintained in the free list.

[0113] This specification also describes how a shared free list can be used to borrow already allocated physical register names. In other words, in some cases, even if a physical register name has already been assigned to a logical register of an instruction, the same physical register name can be reassigned to another logical register under certain circumstances. This arrangement improves processor efficiency, reduces the likelihood of the processor running out of register space, and all of this does not add any additional naming memory.

[0114] Figure 9 PRN borrowing in a system with two physical register sets 910 and 920 is illustrated. In this example, a shared free list 930 is used to assign physical register names for the sequence of instructions 905, 915, and 925. In this example, the same physical register name is used for the logical register referenced by instructions 905 and 925.

[0115] The addition instruction 905 adds the values ​​of logical registers X2 and X3 and uses logical register X1 to store the value. This requires the allocation of a physical register, so the shared free list 930 generates a physical register name, PRN2 in this example, which can be, for example, the binary value 0x10. The processor then uses this physical register name to identify the second physical register in the integer register set 910. Therefore, the result of the addition instruction is written to the second physical integer register 911.

[0116] Next, intermediate store instruction 915 is executed. Since store instructions do not require writing any value to any physical register, they do not need to allocate any physical registers using register renaming. Therefore, if the next instruction only references a logical register of a different type, the physical register name allocated for addition instruction 905 can be reassigned to the logical register of instruction 925.

[0117] Instruction 925 compares the values ​​of logical registers X5 and X6 and stores the result in the flag register. Since instruction 925 is a comparison instruction, its syntax only implicitly references the flag register. Because instruction 925 needs to write to the physical flag register, the shared free list generates a physical register name that identifies one of the flag registers in the flag register file 920.

[0118] However, since the comparison instruction 925 satisfies the criterion of borrowing physical register names, the free list 930 reallocates the same physical register name it generated for instruction 905 by regenerating PRN2. Therefore, the result of the comparison instruction will be stored in the second physical flag register (921).

[0119] Figure 10 This is a flowchart of an example procedure for borrowing physical register names. The procedure can be performed by a processor configured according to this specification. For convenience, the procedure will be described as being performed by the system.

[0120] The system maintains a shared free list (1010) for multiple different register sets. As described above, the different register sets have different types and are referenced by logical register names that reflect those types.

[0121] The system assigns a first physical register name (1020) to the first physical register of the first type. The system can use the free list allocation procedure to generate the next available PRN based on the value of its PRN entry and the position of the head pointer.

[0122] The system receives a second instruction (1030) for a second type of physical register, and the system determines whether the second instruction satisfies one or more borrowing criteria (1030) with respect to the first instruction. In other words, the system can check the second logical register type against one or more previously received instructions to determine whether the instruction satisfies one or more borrowing criteria.

[0123] A first example of a borrowing criterion is when instructions are sequential. Therefore, in some implementations, the system checks the instructions against previously executed instructions to determine if they reference physical registers of different types. If so, the system can determine that both instructions satisfy the borrowing criterion.

[0124] A second example of the borrowing rule is when one or more intermediate instructions are instructions that do not require register renaming. For example, a store instruction between two other instructions is an example of an instruction that does not require register renaming. Therefore, any number of store instructions can appear between two instructions, and the two instructions can still satisfy the borrowing rule.

[0125] A third example of the borrowing criterion is whether instructions are renamed within the same cycle. As described above, the use of input and output buffers allows the processor to perform multiple renamings within the same cycle. Therefore, in some implementations, the system may allow borrowing between instructions that are renamed only within the same cycle.

[0126] A fourth example of the borrowing rule is whether a branch instruction exists between two instructions. In some implementations, to simplify the flush recovery logic, the system may disallow borrowing if the intermediate instruction is a branch instruction. However, in some alternative implementations, if the processor treats the instruction as a non-borrowed case after initiating the flush recovery, borrowing may still be allowed despite the presence of an intermediate branch instruction. In other words, the system may allow borrowing before the flush recovery, but may disallow borrowing for instructions after the flush.

[0127] If one or more borrowing criteria are not met, the system generates a new physical register name for the second instruction (branching to 1050).

[0128] On the other hand, if one or more borrowing criteria are met, the system reallocates the same PRN for the second instruction (branching to 1060). As described above, when different instructions are executed using the same PRN, the results will actually go to different physical register files.

[0129] When two instructions have been assigned to the same PRN, the system can ensure that both instructions are evicted from their respective AMTs before returning the PRNs to the free list. Therefore, when the first PRN is evicted from the AMT by the younger writer to the logical register, the system can check the commit status of the other PRN by checking the explicitly maintained commit bit or by searching its AMT. Only when both PRNs have been evicted from their respective AMTs will the system return the shared PRN to the free list.

[0130] This specification also describes how the shared free list can be further extended to manage physical floating-point registers without a significant increase in hardware cost. By sharing the free list between integer and floating-point registers, the processor can effectively eliminate the need for a separate free list for floating-point registers. Furthermore, the free list can still employ the same allocation and flush recovery process, with only a small number of additional records saved.

[0131] Sharing between integer and floating-point registers is more complex than sharing with flag registers because floating-point registers tend to be fewer than either integer or flag registers. Processor design is the result of carefully choosing the number of each type of register to achieve a specific level of performance and efficiency. Therefore, the relative number of physical registers in each register file varies considerably, but floating-point registers tend to be fewer than integer registers, partly because floating-point registers are larger. As described above, the same number of integer and flag registers can exist efficiently because flag registers store significantly less data than integer registers; for example, a flag register is 4 bits, while an integer register is 64 bits. In contrast, the number of physical floating-point registers can be much smaller, for example, only a quarter or half the number of integer registers, to name just a few common examples.

[0132] To support sharing between integer and floating-point registers, the system can essentially allocate virtual PRNs for floating-point registers. In this specification, "virtual" PRN means that the PRN belongs to a namespace that does not correspond to the size of the physical register file. Instead, there is a mapping between each virtual PRN and each floating-point PRN in a floating-point PRN namespace corresponding to the size of the physical register file. When a free list is used to maintain the virtual PRN namespace, this means that every PRN entry in the free list can be assigned to a floating-point register, but since there are fewer floating-point registers than free list entries, multiple PRN entries with different virtual PRNs in the free list can be mapped to the same floating-point PRN. To avoid reallocating already assigned floating-point PRNs, the system can keep track of which PRN entries are associated with each other by having virtual PRNs mapped to the same floating-point PRN. Then, if a PRN has already been assigned to a floating-point register, no other associated PRN entry can be assigned to another logical floating-point register name until the assigned PRN is freed. However, during this time, the PRNs of other associated PRN entries can still be assigned to other register types, such as integer registers. In the following description, a PRN entry described as a paired or associated PRN entry means a group of multiple PRN entries that have different virtual PRNs mapped to the same floating-point PRN.

[0133] The techniques described below for sharing between integer and floating-point registers can also be combined with the techniques described above for sharing with the flag register and for PRN borrowing. Therefore, the shared free list described in this specification can be used to maintain PRN allocations for at least three different types of physical registers: integer registers, flag registers, and floating-point registers.

[0134] Figure 11An example technique for associating PRN entries for floating-point registers is shown. In this example, the physical floating-point register file has as many physical registers as half the number in the integer physical register file. Therefore, the free list can maintain a virtual PRN namespace by associating two PRN entries with each floating-point PRN. To this end, the system can pair PRN entries in any suitable manner such that each pair of PRN entries identifies the same floating-point PRN. Figure 11 An example of pairing PRN entries in a way that does not require overly complex logic or record keeping is shown.

[0135] like Figure 11 As shown, the free list of N PRN entries is conceptually divided into two halves: the first half 1110 spanning from PRN0 to PRNN / 2-1, and the second half 1120 spanning from PRN N / 2 to PRN N-1. Corresponding PRN entries in each half are paired together such that they each identify the same floating-point PRN in the physical floating-point register file 1130 with N / 2 entries.

[0136] For example, the PRN entries for both PRN0 and PRN N / 2 identify the first floating-point register FP PR0. Similarly, the PRN entries for both PRNN / 2-1 and PRN N-1 identify the last floating-point register FP PR N / 2-1.

[0137] With this arrangement, the system can quickly and efficiently check whether a paired PRN entry has already been assigned to a floating-point register. If so, the system can reassign the same PRN to a different register type, but not to another floating-point register.

[0138] Figure 12 This is a diagram illustrating how the system can maintain a shared free list between integer registers and floating-point registers. As shown, the system includes an integer register file 1210 and a floating-point register file 1220. The system also includes a shared free list 1230. The system further has an integer AMT 1240 for the integer registers and a floating-point AMT 1250 for the floating-point registers. Each AMT stores a mapping between the logical register names and physical register names of its corresponding physical register file.

[0139] Unlike the shared free list described above, the shared free list 1230 includes additional bits for each PRN entry to track whether the PRN entry has been allocated to a floating-point register. Therefore, each PRN entry includes a valid bit, a commit bit, and a floating-point bit (“FP bit”).

[0140] Table 2 summarizes the status of PRN entries that also have floating-point bits.

[0141]

[0142] Table 2

[0143] As shown in Table 2, the FP bit of a PRN entry is used to track whether the PRN entry itself, or an associated PRN entry mapped to the same floating-point PRN, has already been allocated to a floating-point register. If the PRN entry itself has already been allocated to a floating-point register, the PRN cannot be allocated to any other register until it returns to the free list. However, if, for a given PRN entry, its associated PRN entry has already been allocated to a floating-point register, that given PRN entry can still be used to allocate the PRN to a non-floating-point register.

[0144] exist Figure 12 In Operation 1, the free list generates the PRN for the physical register and sets the valid bits of the corresponding PRN entries. If the physical register is a floating-point register, the system sets the FP bit for the PRN entry as well as for all other paired or associated PRN entries. This allows paired or associated PRN entries with the FP bit set to be used only for non-floating-point registers until the PRN returns to the free list, for example, during a bypass return, during a flush, or by being evicted from the FP AMT 1250.

[0145] Then, the PRN-LRN mapping is entered into the ROB 1260. To track which AMT the PRN is retired to, the ROB 1260 can also maintain an integer field 1211 and a floating-point field 1212 to indicate whether the PRN is allocated to a floating-point register or a non-floating-point register. During operation 2, the PRN is retired to its corresponding AMT 1240 or 1250, which can be based on the integer field 1211 and floating-point field 1212 of the ROB 1260. The corresponding PRN commit bit is then set to indicate that the PRN has been retired to its AMT, as indicated by the label “Set PRN Commit = 1”.

[0146] During the flush recovery at operation 3, the system can use a checkpoint in the ROB to reset the head pointer to the flush point, and then flush the affected PRN entry by resetting the valid bits of the affected PRN entry between the old head pointer and the flush point, unless the affected PRN entry is stored in the AMT according to its commit bit. Therefore, if the commit bit of the flushed PRN entry is 0, the system clears both the valid bits and the floating-point bits.

[0147] However, the system can perform some additional maintenance to account for the virtual PRN, which may result in restoring the FP bit in the PRN entry. That is, the system can perform a bitwise OR operation across all associated PRN entries, and if any of the associated PRN entries still has the FP bit set, the system sets the FP bit in all associated PRN entries. On the other hand, if none of the associated PRN entries have the FP bit set, the system can refuse to restore any of the FP bits.

[0148] This additional step takes into account the case where a PRN entry is allocated to a physical floating-point register and its paired or associated PRN entry is flushed. In this case, the system can restore the FP bit of the flushed PRN entry so that it is not reassigned to another floating-point register. Therefore, the system can perform a bitwise OR operation between the FP bits of all associated PRN entries and store the resulting value of the FP bits in all associated PRN entries.

[0149] In this example, the bitwise OR procedure includes first determining the paired PRN entries. When the PRN entries are as follows... Figure 11 As shown, when pairing, the paired PRN entries can be obtained using integer addition of N / 2, where N is the size of the free list (1221). Then, the system decodes the value of the FP bit of the paired PRN entries (e.g., using decoding modules 1222 and 1223), performs a bitwise OR between the FP bits (1224), and writes the result to the two PRN entries in the free list 1230.

[0150] A final step can be performed to account for the opposite case where paired PRN entries have not yet been assigned (e.g., proven by their valid bits). In this case, the system clears the FP bits of both PRN entries. See below for reference. Figure 13 This maintenance of the FP bit during flushing is described in more detail.

[0151] During operation 4, when a floating-point PRN is evicted from the FP AMT 1250, the system clears the FP bits of the floating-point PRN entry and all other associated PRN entries. To do this, the system may first identify the paired PRN entries, which in this example may include performing an integer addition of N / 2 (1231). The system may then clear the corresponding FP bits (1232 and 1233) of the paired PRN entries. Following the usual procedure of returning a PRN to the free list, the system may also clear the valid and committed bits of the PRN entry.

[0152] During operation 5 (which is a bypass return of the FP PRN without reaching the FP AMT 1250), the system can clear the valid and committed bits of the PRN entry. Furthermore, since the FP PRN is ready to be reassigned to another floating-point register, the system can also clear all FP bits of all paired or associated PRN entries.

[0153] Figure 13 This is a flowchart of an example procedure for maintaining virtual PRN entries during flush recovery when managing multiple different register sets. The procedure can be performed by a processor configured according to this specification. For convenience, the procedure will be described as being performed by the system.

[0154] As described above, in order to perform flushing recovery, the system resets the head pointer to the flushing point. This will be based on... Figure 13 The operation processes all PRN entries after the flush point and before the old head pointer. Therefore, the system can perform the example procedure described above on all flushed PRN entries.

[0155] The system determines whether a commit bit is set (1310). If the commit bit is not set, the PRN entry's PRN can be returned to the free list. Therefore, the system clears the valid and floating bits of the PRN entry (branch to 1320). If the commit bit is set, the system jumps to the next step (branch to 1330).

[0156] The system performs a bitwise OR operation between the associated PRN entries and writes the result (1330) to the FP bit of each associated PRN entry. In other words, if any one of the paired or associated PRN entries sets its FP bit, then all paired or associated PRN entries will set their FP bits.

[0157] The system determines whether any of the associated PRN entries has a valid bit set (1340). In this case, if the PRN entry in the paired or associated group of PRN entries has a valid bit set, then that PRN entry has been allocated to an unflushed register and will be left alone (branch to end).

[0158] However, if none of the paired or associated PRN entries have valid bits set, the FP bits also need to be cleared so that the PRN entries can be used in the floating-point register in the future. Therefore, if none of the paired or associated PRN entries have valid bits set, the system will clear all floating-point bits of the PRN entries in the group (branch to 1350).

[0159] Figure 14This is a flowchart of an example procedure for allocating a virtual PRN to a free list that manages multiple different register sets. The procedure can be performed by a processor configured according to this specification. For convenience, the procedure will be described as being performed by the system.

[0160] The system maintains a shared free list for various types of register sets (1410). As described above, the system can allocate virtual PRNs by maintaining a virtual PRN namespace, where multiple PRN entries with different virtual PRNs can be mapped to the same physical register name. Therefore, the system can group multiple PRN entries with the same PRN number mapped to a specific type of physical register together.

[0161] The system assigns a first PRN (1420) to the logical register name of the second type and designates one or more other PRN entries as unavailable for allocation to logical registers of the second type (1430). As described above, the system may maintain a separate bit value in each PRN entry to indicate whether any of the associated registers in the group has been allocated to a register of a particular type. For example, the system may maintain floating-point bits to indicate whether any paired or associated PRN entries have been allocated to floating-point registers. If so, other PRN entries in the group are unavailable for allocation to floating-point registers, but they may still be allocated to other types of non-floating-point registers.

[0162] Although the above description uses floating-point registers as an example of a register type with a virtual PRN for clarity, the same technique can be used for any other suitable register type.

[0163] This specification also describes another way in which a shared free list can support multiple different types of register files. In the floating-point register example described above, since there are fewer floating-point registers than integer registers, the free list ensures that only the correct number of floating-point PRNs are allocated according to the floating-point register file. In this example, for a free list of size N, N / 2 floating-point PRNs are allocated for N / 2 physical floating-point registers.

[0164] However, for some applications that heavily utilize floating-point execution, the renaming process itself can become a performance bottleneck. Therefore, the system described above can be configured or reconfigured (e.g., during execution or manufacturing) to allow multiple associated or paired PRN entries to be concurrently assigned to different logical floating-point registers. In other words, a group of PRN entries with different virtual PRNs mapped to the same floating-point PRN can be assigned to logical register names within the same time period. In contrast, in the example described above, paired or associated PRN entries mapped to the same floating-point PRN as the already assigned PRN can only be assigned to registers of different types, such as integer registers or flag registers.

[0165] To prevent clashes between multiple instructions assigned to the same floating-point PRN, the system can manage renaming and execution in different phases. In the first phase, the system assigns a floating-point PRN to an instruction but does not yet allow its execution. In the second phase, the system issues the instruction for execution only if a PRN entry with different virtual PRNs mapped to the same floating-point PRN is available for renaming (e.g., because it was never assigned or because it has been returned to the free list). Returning the same floating-point PRN to the free list means that the floating-point PRN is evicted from the AMT or returned in another way, such as through a bypass return as described above, and means that the program no longer needs the data in the physical registers. The effect of this is to push the resolution of name collisions from the renaming phase to the execution phase, where real data dependency conflicts must be resolved in any case.

[0166] To support the assignment of multiple paired or associated PRN entries to a floating-point register, the system can use modified logic to determine when to set and clear the FP bit of the PRN entry when allocating and freeing PRN entries in the free list for the floating-point register, and to use modified logic to determine when to allow the instruction to be issued when the instruction is written to the floating-point register.

[0167] To maintain information about which instructions using floating-point registers are ready to be issued, the system can maintain a sliding window of PRN entries in the free list. The sliding window separates which PRNs are ready to be issued for execution from which PRNs must wait for the floating-point PRNs to return to the free list before being issued. The sliding window is a data item maintained separately from the head and tail pointers of the free list itself.

[0168] This specification also describes the techniques and circuitry for handling stray floating-point PRNs that might otherwise never enter the sliding window for execution. There are many situations in which a PRN can remain in the AMT. As described above, a PRN can remain in the AMT until a logical writer evicts it. However, in many programs, there may be instructions that are written to a logical register once and then not written to that register again long after the program has started, or never written to that register again. In this case, the floating-point PRN may remain in the AMT indefinitely, preventing its paired or associated PRN entry from being issued for execution.

[0169] Figure 15 A sliding window for PRNs in the free list is shown. In this example, the free list is maintained in storage units, each with 16 PRN entries. Therefore, the first storage unit 1510 contains PRN entries 0 through 15. The total free list size is 512 entries, so the last entry 1520 contains PRN entries 495 through 511.

[0170] In this example, the number of floating-point registers is half the number of integer registers. Therefore, the sliding window 1545, indicated by the dashed line, has 256 entries, which is half the number of integer registers. Using this arrangement, the system can assign two PRN entries to two different instructions, even when PRN entries are mapped to the same floating-point PRN. In other words, different virtual PRNs for PRN entries are mapped to the same floating-point PRN.

[0171] To implement the sliding window, the system can maintain a floating-point tail pointer (FP tail pointer), which points to the oldest PRN entry that has been allocated to a floating-point register in the main floating-point register file. The FP tail pointer differs from the free list tail pointer described above; for clarity, the free list tail pointer may be referred to as the FL tail pointer to distinguish it from the FP tail pointer that defines the sliding window. Based on the FP tail pointer, the system can consider the next M PRN entries as being within the sliding window, where M is based on the size of the physical floating-point register file. Therefore, the system can use the FP tail pointer to determine which PRN entries are designated as ready for execution by entering the sliding window. Instructions using PRN entries within the sliding window can be executed safely because, since the sliding window is based on the size of the physical floating-point register file, the instructions are guaranteed to have a corresponding physical floating-point register at execution time.

[0172] Therefore, in order to move the sliding window, the next PRN entry or the next group of PRN entries (e.g., a bank or row) must clear its FP bit.

[0173] References above Figures 11 to 14 In the previous example described, the FP bit was set when a paired or associated PRN entry was assigned to a floating-point register. However, in this example, this requirement is modified, and the setting and clearing of each FP bit depends solely on the state of the PRN entry itself, not on the associated or paired PRN entry. Instead, the system uses a sliding window instead of the FP bit to avoid name collisions.

[0174] Therefore, when a PRN entry becomes free and available for renaming (e.g., by returning to the free list after being evicted from the FP AMT, by being part of a flush, or by entering the structured floating-point register file, which is described in more detail below), the system can clear the FP bit of the PRN entry. The system can check the valid bits of the PRN entry to determine if it is free and available for renaming.

[0175] The FP tail pointer can step by PRN entry or by each memory bank of a PRN entry. When the FP tail pointer steps by PRN entry, it can advance as the FP bit of the oldest allocated PRN entry is cleared. For example, the system can continuously check the valid bits of the PRN entry indicated by the FP tail pointer, and if a valid bit indicates that the PRN entry is available for renaming, the system can clear the FP bit of the PRN entry and advance the FP tail pointer. When the FP tail pointer steps by memory bank of a PRN entry, it can advance as all FP bits in the oldest memory bank are cleared.

[0176] When the FP tail pointer is updated, thereby moving the sliding window, the system can issue instructions to the PRN entries that have entered the sliding window sequentially or concurrently.

[0177] Figure 16 The data structure used to handle backlogged floating-point PRNs is shown. Backlogged PRNs in the AMT can delay the execution of floating-point instructions because the FP bit of the PRN entry will never be cleared. Furthermore, when the FP bit is never cleared, the FP tail pointer cannot advance, and new floating-point instructions will not enter the sliding window.

[0178] Therefore, to address this issue, the system can use a new auxiliary physical register file that stores the structured floating-point values ​​of the physical registers with the PRN still in the AMT. The Structured Floating-Point Physical Register File (APRF) 1620 is a separate physical register file from the Main Floating-Point Physical Register File (MPRF) 1610.

[0179] The APRF 1620 requires only as many physical registers as the logical registers used by the processor, while the MPRF 1610 typically has more physical registers. For example, if the processor has 32 logical registers, the APRF only needs 32 physical registers. Meanwhile, the MPRF typically has hundreds of physical registers or more, such as 256 or 512 registers.

[0180] The APRF 1620 allows the processor to clear the FP bit of a floating-point PRN entry while still retaining the possibility that the value of the pending PRN will be read in the future. The processor cannot simply remove pending PRNs from the AMT because it cannot know whether the logic register will be read at some future time.

[0181] APRF 1620 essentially provides another way to clear the FP bit of a PRN entry. Therefore, the ways to clear the FP bit of a PRN entry when using APRF include: PRN being evicted from the FP AMT, being returned to the free list via bypass or flushing, and PRN entering the APRF.

[0182] Each entry in APRF 1620 has a corresponding PRN in virtual PRN table 1640. The size of the entries in virtual PRN table 1640 only requires the number of bits needed to store the PRN, and therefore, for the same number of entries, virtual PRN table 1640 is typically smaller than APRF 1640.

[0183] Consumer 1650 (e.g., an execution unit executing instructions to read floating-point registers) that stores data in physical floating-point registers can first check virtual PRN table 1640. If a hit is found, consumer 1650 will read the data from APRF 1620 instead of MPRF 1610. Otherwise, the consumer will read from MPRF 1610.

[0184] Table 3 summarizes the status of PRN entries when using a sliding window.

[0185]

[0186] Table 3

[0187] Table 4 shows an example of a PRN with a sequence of instructions held up. In this example, the free list has 10 entries, and the floating-point physical register file has 5 registers. Therefore, the sliding window size is five entries. In this example, the FP tail pointer points to the PRN entry corresponding to PRN 3. Table 3 indicates the virtual PRN and corresponding floating-point PRN assigned to each of the ten floating-point instructions, and the state of the FP bit of each PRN entry when the FP tail pointer of the sliding window points to PRN entry 3.

[0188]

[0189] Table 4

[0190] In this example, if floating-point PRN3 remains in the AMT, instruction 8 may never be issued because it will never enter the sliding window. In other words, the FP bit of the PRN entry used by instruction 3 will not be cleared until PRN3 returns to the free list, and PRN3 returning to the free list will trigger an update of the FP tail pointer.

[0191] However, by storing the value of PRN3 into the APRF, the system can clear the FP bit of PRN entry 3 for instruction 3. This will trigger an update of the FP tail pointer, instruction 8 will enter the sliding window, and the processor can then issue instruction 8 for execution. Subsequent consumers of the logical register to which PRN3 is allocated will read data from the APRF instead of from the main physical register file.

[0192] Figure 17 This is a flowchart of an example procedure for managing a virtual PRN for managing free lists of multiple different register sets. The procedure can be performed by a processor configured according to this specification. For convenience, the procedure will be described as being performed by the system.

[0193] The system maintains a shared free list for multiple different register sets (1710). As described above, the free list allows multiple different PRN entries to be assigned to floating-point register names by assigning virtual PRNs. The number of virtual PRNs is greater than the actual number of physical floating-point registers.

[0194] The system allocates a first virtual PRN based on the first PRN entry and designates the first PRN entry as ready to execute (1720). The system can then execute the first instruction.

[0195] The system allocates a second virtual PRN (1730) mapped to the same first floating-point PRN based on the second PRN entry. The system can still assign a virtual PRN to the second PRN entry, but may wait until the first PRN returns to the free list before issuing the instruction. In some implementations, the system uses a sliding window and waits until the second PRN entry arrives at the sliding window before issuing the instruction; by the time the second PRN entry arrives at the sliding window, the first PRN will have already returned to the free list.

[0196] The system determines that the first PRN has returned to the free list (1740), and the system designates the second instruction as ready for execution (1750). For example, the system can use a sliding window as a way to determine when a previously allocated PRN has returned to the free list. When the second PRN entry enters the sliding window, the system can issue the corresponding instruction.

[0197] Issuing an instruction with a renamed register may include sending the instruction to a reserved station. The reserved station issuing the floating-point instruction may continuously check a sliding window, as indicated by the FP tail pointer, to determine whether to issue the instruction. Then, once the corresponding PRN entry enters the sliding window, the reserved station can execute the instruction. In some implementations, the system includes modified logic for the reserved station handling floating-point instructions. Reserved stations handling other instructions not written to the floating-point register can issue their instructions regardless of the state of the sliding window.

[0198] This specification also describes how a system with a hybrid free list can use data in the free list to clock-gated regions of the physical register file. Clock gating is a power-saving technique that refers to reducing the power usage of specific regions of an electronic device.

[0199] Figure 18 Clock gating based on a free list is illustrated. One advantage of the hybrid free list described in this specification is that it implicitly includes information about which regions of the physical register file are likely to be active and which are likely to be inactive. In particular, the region between the head pointer and the tail pointer is inherently a highly active region of the corresponding physical register file. The system can use this information to intelligently control which regions of the physical register file can be clock-gated to save power.

[0200] Figure 18 A mixed free list 1830 with 512 entries is shown, the 512 entries being divided into memory banks, each with 16 entries. Free list 1830 allocates PRNs to physical register file 1810, which has lower memory bank 1811 and upper memory bank 1812.

[0201] Free list 1830 has a head pointer 1815 and a tail pointer 1805. Physical registers with names represented by PRN entries falling between head pointer 1815 and tail pointer 1805 are physical registers that are likely to be read and written during program execution. Therefore, the system can designate the region of physical register file 1810 corresponding to the PRN entries between head and tail pointers as active region 1840.

[0202] If other areas of the physical register file meet certain criteria, they can be designated as "low-activity" or "inactive." If so, the system can clock-gated portions of the register file or subsets of the register file. For example, register file 1810 is divided into two banks, 1811 and 1812. The portion following the lower bank 1811 is a low-activity region, and therefore the system can clock-gated the registers belonging to that portion of the lower bank 1811.

[0203] Alternatively or additionally, the system can clock-gated the entire memory bank of physical register file 1810. For example, if the registers of upper memory bank 1812 are in low-activity areas, for example, because they are outside the head and tail pointers, the system can clock-gated the entire upper memory bank 1812.

[0204] The system can use any appropriate criteria to determine whether a physical register is in a low-activity or inactive state based on the free list 1830. For example, the row of PRN entries in the free list 1830 where all valid bits are cleared corresponds to the group of physical registers that are not being used by the program. Therefore, the system can clock-gated that region, including reducing the region's power, clock frequency, voltage, or some combination thereof.

[0205] In some cases, PRN entries may still exist outside the head and tail pointers, containing data that could potentially be read during program execution. For example, if the PRN has been schemated into an AMT, the value of that physical register may still be read during program execution. However, due to the schemated state of the PRN, it is no longer possible to write to it.

[0206] Therefore, the system can also designate regions in the physical register file that are outside the head and tail pointers but still have one or more PRN entries with valid bits set as active regions. Alternatively or additionally, the system can clock-gated these regions less aggressively, for example, by setting them to a lower power mode where they can still be read from but not written to.

[0207] Figure 19 This is a flowchart illustrating an example of clock gating of a physical register file based on a free list. The process can be performed by a processor configured according to this specification. For convenience, the process will be described as being performed by the system.

[0208] The system maintains a mixed free list (1910) for physical register files partitioned into multiple regions. The physical register file can be partitioned into regions that can be clock-gated individually, such as banks, rows, or any other suitable partitions.

[0209] The system determines that the region in the mixed free list represents an inactive region of the physical register file (1920). In some implementations, the system only searches inactive regions outside the head and tail of the free list (e.g., before the head pointer and after the tail pointer). As described above, the region in the physical register file corresponding to the PRN entry between the head and tail pointers is likely to be the region where active reads and writes occur during program execution.

[0210] The system can identify active regions based on PRN entries by examining their valid bits. PRN entries with their valid bits cleared essentially represent unused physical registers. Therefore, if the system can identify the adjacent region of a PRN entry whose valid bits are all cleared, the system can identify the corresponding region of the physical register file as an inactive region. A significant benefit of the hybrid free list described in this specification is that it tends to group large regions of inactive PRN entries, which can then be used to identify inactive regions of the physical register file.

[0211] For PRN entries that are outside the head and tail pointers but whose valid bits are still set, the system can designate those areas as still active because the corresponding physical registers may still be read during program execution.

[0212] The system reduces the power of inactive regions of the physical register file (1930). The system can, for example, reduce the clock frequency, voltage, or completely shut down inactive regions of the physical register file.

[0213] The system can re-evaluate the active and inactive regions of the physical register file at any appropriate time interval or as the head and tail pointers move. For example, if the head pointer moves to the next line of the free list corresponding to an inactive region of the physical register file, the system can stop clock-gating that region and restore the previous power or clock level because its physical registers are about to be allocated for program execution. Similarly, when the tail pointer of the free list moves to the next line, the system can check the valid bits of the PRN entry in the previous line to determine if all valid bits have been cleared, in which case the system can clock-gate a new region of the physical register file.

[0214] These techniques allow the system to activate specific regions of the physical register file before registers are needed for execution, reducing or eliminating any latency associated with clock gating of the physical register file. For example, the system can use a free list to reactivate inactive regions of the physical register file at instruction issue time or register renaming time. In either case, the corresponding regions of the physical register file will be started and running when the instruction is issued.

[0215] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in a combination of one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more executable programs, i.e., one or more modules of instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0216] The term "data processing device" refers to data processing hardware and encompasses all kinds of devices, apparatuses, and machines used for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The device may also be or further include a dedicated logic circuit system, such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0217] A computer program, also referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any programming language (including compiled or interpreted languages, or declarative or procedural languages) and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program may be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as a collection of coordinated files (e.g., a file storing portions of one or more modules, subroutines, or code). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a data communication network.

[0218] The processes and logic flows described in this specification can also be performed by a dedicated logic circuit system (e.g., an FPGA or ASIC) or by a combination of a dedicated logic circuit system and one or more programmed computers.

[0219] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0220] In addition to the embodiments described above, the following embodiments are also innovative: Example 1 is a method comprising: Maintain a hybrid free list, which represents physical registers available for register renaming, wherein the hybrid free list comprises multiple ordered entries, each entry comprising a pair of bits for each of the multiple physical registers. Each entry includes a pair of bits representing the valid and committed states of the physical register, respectively. The valid state indicates whether the physical register is free and can be allocated for a logical register name, and The commit status indicates whether the instruction using the physical register has been committed; and The head and tail pointers of the free list are used to allocate physical register names in order based on the multiple entries in the free list.

[0221] Example 2 is the method as described in Example 1, further comprising: Use the head pointer to find entries that are in an idle state.

[0222] Example 3 is the method as described in Example 2, wherein the entries of the hybrid free list are arranged in multiple storage units, and

[0223] The entries identified as having an idle state include: Read the memory bank corresponding to the head pointer; and Identify entries in the memory that are in an idle state.

[0224] Example 4 is the method as described in Example 3, further comprising: The output temporary buffer is filled based on the valid status of the entries in the memory; and Use the entries in the output temporary buffer to allocate physical register names.

[0225] Example 5 is the method as described in Example 4, wherein filling the output temporary buffer includes: The complete physical register name is generated based on the input temporary buffer that stores the valid bits of the entries from the memory bank.

[0226] Example 6 is the method as described in Example 5, wherein generating the complete physical register name includes expanding each bit in the input temporary buffer to a multi-bit physical register name in the output temporary buffer.

[0227] Example 7 is the method as described in any one of Examples 1 to 6, further comprising: The validity state of the entry corresponding to the physical register name is modified in response to the assignment of the physical register name.

[0228] Example 8 is the method as described in any one of Examples 1 to 7, further comprising: When an instruction that uses a physical register has been submitted, the submission status of the entry corresponding to the physical register is modified.

[0229] Example 9 is the method as described in any one of Examples 1 to 8, further comprising: When a subsequent instruction is written to a logical register name, the validity state and the commit state of the entry corresponding to the physical register name assigned to the same logical register name are modified.

[0230] Example 10 is the method as described in any one of Examples 1 to 9, further comprising: Receive instructions for flushing at a specific flushing point; Modify the valid status of any entry in the mixed idle list that does not have a submitted status between the flush point and the head pointer.

[0231] Example 11 is the method as described in Example 10, further comprising updating the head pointer of the mixed free list based on the flush point.

[0232] Example 12 is the method as described in Example 10, wherein the flushing point is based on the oldest assigned PRN affected by the flushing.

[0233] Example 13 is the method as described in Example 12, further comprising: Store the checkpoints for each PRN-to-LRN mapping in the reorder buffer (ROB); and The checkpoints in the ROB are used as the flushing points.

[0234] Example 14 is a method for implementing a shared free list for multiple different register sets, the method comprising: Maintain a shared free list for multiple different register sets, each register set including one or more registers of a specific type, wherein the free list has multiple Physical Register Name (PRN) entries; Based on the first free list entry, generate a first physical register name for the first physical register of the first type for the first logical register referencing the first type of the first logical register; Based on the second free list entry, a second physical register name is generated for the second physical register of the second type of the second logical register using the second instruction that references the second type of the second logical register. After the first instruction has been executed, the first physical register name is retired to the first architectural mapping table (AMT) for registers of the first type. Retire the second physical register name to the second AMT used for the second type of register; and If the first physical register name is not in the second AMT, then the first physical register name is specified as free.

[0235] Example 15 is the method as described in Example 14, further comprising specifying the second physical register name as free if the first physical register name is not in the first AMT.

[0236] Example 16 is the method as described in Example 15, wherein specifying the second physical register name as free includes: It is determined that the PRN entry for the second physical register name in the free list does not have the commit bit set; and In response, the valid bits of the PRN entry are cleared.

[0237] Example 17 is the method as described in any one of Examples 14 to 16, wherein specifying the first physical register name as free includes: It is determined that the second AMT does not maintain the name of the second physical register; and In response, the valid bits of the PRN entry for the first physical register name are cleared.

[0238] Example 18 is the method as described in any one of Examples 14 to 17, further comprising: Based on the third hybrid free list entry, a third instruction is used to generate a shared physical register name for the third physical register of the first type and the fourth physical register of the second type, referencing both the first logical register of the first type and the second logical register of the second type.

[0239] Example 19 is the method as described in Example 18, further comprising returning the shared physical register name to the free list only if the shared physical register name has been evicted from both the first AMT and the second AMT.

[0240] Example 20 is a method as described in any one of Examples 14 to 19, wherein the first type is an integer register.

[0241] Example 21 is a method as described in any one of Examples 14 to 20, wherein the second type is a flag register.

[0242] Example 22 is a method comprising: Maintain a shared free list for multiple different register sets, each register set including one or more registers of a specific type, wherein the free list has multiple Physical Register Name (PRN) entries; Based on the first free list entry, assign a first physical register name to the first physical register of the first type for a first logical register referencing the first type; Determine that a second instruction referencing a second logical register of a different second type satisfies one or more borrowing criteria with the first instruction; and In response, the name of the first physical register is reassigned to the second logical register referenced by the second instruction.

[0243] Example 23 is the method as described in Example 22, wherein reassigning the first physical register name to the second logical register includes assigning the first physical register name to the second logical register when the first physical register name has already been assigned to the first logical register.

[0244] Example 24 is a method as described in any one of Examples 22 to 23, wherein the reallocation of the first physical register name results in the first physical register name being simultaneously assigned to different types of logical registers.

[0245] Example 25 is a method as described in any one of Examples 22 to 24, wherein determining that the second instruction satisfies one or more borrowing criteria includes determining that the first instruction and the second instruction are consecutive instructions.

[0246] Example 26 is a method as described in any one of Examples 22 to 25, wherein determining that the second instruction satisfies one or more borrowing criteria includes determining that there is no intermediate instruction between the first instruction and the second instruction that requires a register renaming.

[0247] Example 27 is the method as described in Example 26, wherein determining that the second instruction satisfies one or more borrowing criteria includes determining that the intermediate instruction does not require a logic register write.

[0248] Example 28 is a method as described in any one of Examples 26 to 27, wherein the intermediate instruction is a storage instruction.

[0249] Example 29 is a method as described in any one of Examples 22 to 28, wherein determining that the second instruction satisfies one or more borrowing criteria includes determining that the first instruction and the second instruction are renamed in the same cycle.

[0250] Example 30 is a method as described in any one of Examples 22 to 29, wherein determining that the second instruction satisfies one or more borrowing criteria includes determining that an intermediate instruction between the first instruction and the second instruction is not a branch instruction.

[0251] Example 31 is the method as described in any one of Examples 22 to 30, further comprising submitting the first instruction and the second instruction at different periods.

[0252] Example 32 is a method as described in any one of Examples 22 to 31, wherein the first type or the second type is an integer register.

[0253] Example 33 is a method as described in any one of Examples 22 to 32, wherein the first type or the second type is a flag register.

[0254] Example 34 is a method comprising: Maintain a shared free list for multiple different register sets, including a first register set and different second register sets, wherein each register set includes one or more registers of a corresponding type, wherein the second register set has fewer physical registers than the first register set, and wherein the free list has multiple physical register name (PRN) entries, wherein the free list associates a group of one or more PRN entries used to identify a single register in the second register set, wherein at least one group of the multiple PRN entries identifies a single physical register in the second register set; A first physical register name is assigned based on a first PRN entry for a first logical register name of the second type, wherein the first PRN entry belongs to a group of one or more other associated PRN entries that also identify the same physical register name of the second register set; and The one or more other PRN entries in the group are designated as unavailable for allocation to logical registers of the second type until the first PRN entry becomes available again.

[0255] Example 35 is the method as described in Example 34, further comprising: Assign a second physical register name to the second logical register name of the first type based on the second PRN entry belonging to the group.

[0256] Example 36 is a method as described in any one of Examples 34 to 35, wherein multiple PRN entries in the group are assigned to different types of physical registers.

[0257] Example 37 is a method as described in any one of Examples 34 to 36, wherein each PRN entry includes a second register type bit to indicate whether any PRN entry in the group of associated PRN entries has been assigned to a logical register of the second type.

[0258] Example 38 is the method as described in Example 37, wherein designating the one or more other PRN entries in the group as unavailable includes setting the first PRN entry and each corresponding second register type bit of the one or more other PRN entries in the group.

[0259] Example 39 is a method as described in any one of Examples 34 to 38, wherein for a free list having N PRN entries, the group of PRN entries is a pair of PRN entries spaced N / 2 entries apart.

[0260] Example 40 is a method as described in any one of Examples 34 to 39, wherein the second register set has as many physical registers as half the number of the first register set.

[0261] Example 41 is a method as described in any one of Examples 34 to 40, wherein the first register set includes integer physical registers and the second register set includes floating-point registers.

[0262] Example 42 is the method as described in Example 41, wherein the shared free list is configured to maintain PRN entries for the integer register, the flag register, and the floating-point register.

[0263] Example 43 is a method comprising: Maintain a shared free list for multiple different register sets, the multiple different register sets including a first register set and different second register sets, wherein each register set includes one or more registers of a corresponding type, wherein the second register set has fewer physical registers than the first register set, and wherein the free list has multiple physical register name (PRN) entries, wherein the free list maintains multiple PRN entries that identify individual registers in the second register set; The first virtual PRN is allocated according to the first PRN entry for the first logical register name of the second type in the first instruction; A second virtual PRN is allocated to the first physical register name based on the second PRN entry for the second logical register name of the second type in the second instruction; The first instruction is executed using the first physical register name; and The second instruction for execution using the first physical register name is issued only after the first virtual PRN returns to the free list.

[0264] Example 44 is the method as described in Example 43, further comprising: determining that the first virtual PRN has been returned to the free list, including determining that the valid bits in the first PRN entry have been cleared.

[0265] Example 45 is the method as described in any one of Examples 43 to 44, further comprising: Maintain a sliding window for PRN entries of the second type that are ready to be issued for execution. Issuing the second instruction includes issuing the second instruction only after the second PRN entry enters the sliding window.

[0266] Example 46 is the method as described in Example 45, further comprising: Maintain a floating-point tail pointer (FP tail pointer), which indicates the oldest allocated PRN entry of the second type to a physical register in the main physical register set of the second type, and wherein the sliding window is defined by a predetermined number of PRN entries following the FP tail pointer.

[0267] Example 47 is the method as described in Example 46, further comprising: Determine that the PRN entry associated with the PRN entry indicated by the FP tail pointer is free and can be used for renaming; and In response, update the FP tail pointer.

[0268] Example 48 is the method as described in Example 47, further comprising: Clear the floating-point bits of the PRN entry indicated by the FP tail pointer.

[0269] Example 49 is the method as described in Example 45, wherein the sliding window size is based on the number of registers in the second register set.

[0270] Example 50 is the method as described in Example 45, further comprising: It is determined that the PRN assigned to the third logical register name based on the third PRN entry for the third physical register of the second type remains in the Architectural Mapping Table (AMT). In response, data is moved from the third physical register to the structured floating-point physical register file; and Clear the FP bit of the fourth PRN entry mapped to the same PRN that remains in the AMT.

[0271] Example 51 is the method as described in Example 50, wherein determining that the PRN remains in the AMT includes: It is determined that the FP tail pointer points to the third PRN entry, while the PRN has not yet been evicted from the AMT.

[0272] Example 52 is the method as described in Example 50, further comprising: Execute the instruction read from the name of the third logical register; and Values ​​are read from the structured floating-point physical register file instead of from the main floating-point physical register file.

[0273] Example 53 is the method as described in any one of Examples 43 to 52, further comprising: Issue instructions with renamed registers to the reserved station for execution; Continuously check whether the PRN entry for the instruction is within the sliding window; and The instruction is executed after the PRN entry enters the sliding window.

[0274] Example 54 is a method as described in any one of Examples 43 to 53, wherein the first register set includes integer physical registers and the second register set includes floating-point registers.

[0275] Example 55 is the method described in Example 54, wherein the shared free list is configured to maintain PRN entries for the integer register, the flag register, and the floating-point register.

[0276] Example 56 is a method comprising: A free list is maintained in the processor, which tracks the availability of physical registers in a physical register file, wherein the free list is partitioned into multiple regions, each corresponding to a subset of the physical register file; The region in the free list represents the inactive region of the physical register file; and In response, clock gating is performed on one or more physical registers in the physical register file that correspond to the inactive region of the free list.

[0277] Example 57 is the method as described in Example 56, wherein each free list entry in the free list has a valid bit indicating whether the corresponding physical register is in use, and

[0278] Determining that a region of the free list is an inactive region includes determining that all valid bits of the free list entries in that region indicate that the corresponding physical register is not in use.

[0279] Example 58 is a method as described in any one of Examples 56 to 58, further comprising: The region of the free list is determined to be between the head pointer and the tail pointer of the free list; and In response, designate that area as the active area.

[0280] Example 59 is a method as described in any one of Examples 56 to 58, wherein the physical register file comprises a plurality of memory banks, and wherein determining that region of the free list is an inactive region comprises determining that all physical registers in the memory bank corresponding to that region of the free list are unused.

[0281] Example 60 is the method as described in Example 59, wherein clock gating of the one or more physical registers includes clock gating of the memory corresponding to the region of the free list, while maintaining normal clock operation for one or more other memory banks of the physical register file.

[0282] Example 61 is a method as described in any one of Examples 56 to 60, wherein clock gating of the one or more physical registers occurs before the one or more physical registers have been assigned to instructions.

[0283] Example 62 is a method as described in any one of Examples 56 to 61, further comprising activating the one or more physical registers during the instruction issuance phase of using a physical register in one or more physical registers.

[0284] Example 63 is a method as described in any one of Examples 56 to 62, further comprising restoring normal clock operation of the physical registers during the register renaming process using the free list.

[0285] Example 64 is a processor that includes a logic circuit system configured to perform the method as described in any one of Examples 1 to 63.

[0286] Example 64 is a data processing device configured to perform the method as described in any one of Examples 1 to 63.

[0287] While this specification contains numerous details of specific implementations, these should not be construed as limiting the scope of any invention or the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases, one or more features from said combination may be removed from the claimed combination, and the claimed combination may involve sub-combinations or variations thereof.

[0288] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all shown operations to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0289] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims may be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.

Claims

1. A processor, comprising: A logic circuit system configured to maintain a hybrid free list representing physical registers available for register renaming, wherein the hybrid free list comprises multiple ordered entries, each entry comprising a pair of bits for each of the multiple physical registers, and wherein the logic circuit system is configured to generate physical register names for free physical registers, and wherein the logic circuit system maintains head and tail pointers for sequential allocation of free physical register names. Each entry includes a pair of bits representing the valid and committed states of the physical register, respectively. The valid state indicates whether the physical register is free and can be allocated for a logical register name, and The commit status indicates whether the instruction using the physical register has been committed.

2. The processor of claim 1, wherein the logic circuitry is configured to perform operations, the operations including: Use the head pointer to find entries that are in an idle state.

3. The processor of claim 2, wherein the entries of the hybrid free list are arranged in multiple storage units, and The entries identified as having an idle state include: Read the memory bank corresponding to the head pointer; as well as Identify entries in the memory that are in an idle state.

4. The processor of claim 3, wherein the logic circuitry is configured to perform operations, the operations including: The output temporary buffer is filled based on the valid status of the entries in the storage. as well as Use the entries in the output temporary buffer to allocate physical register names.

5. The processor of claim 4, wherein filling the output temporary buffer comprises: The complete physical register name is generated based on the input temporary buffer that stores the valid bits of the entries from the memory bank.

6. The processor of claim 5, wherein generating the complete physical register name includes expanding each bit in the input buffer to a multi-bit physical register name in the output buffer.

7. The processor of any one of claims 1 to 6, wherein the logic circuit system is configured to perform an operation, the operation comprising: The validity state of the entry corresponding to the physical register name is modified in response to the assignment of the physical register name.

8. The processor of any one of claims 1 to 7, wherein the logic circuit system is configured to perform an operation, the operation comprising: When an instruction that uses a physical register has been submitted, the submission status of the entry corresponding to the physical register is modified.

9. The processor of any one of claims 1 to 8, wherein the logic circuit system is configured to perform an operation, the operation comprising: When a subsequent instruction is written to a logical register name, the validity state and the commit state of the entry corresponding to the physical register name assigned to the same logical register name are modified.

10. The processor of any one of claims 1 to 9, wherein the logic circuitry is configured to perform an operation, the operation comprising: Receive instructions for flushing at a specific flushing point; Modify the valid status of any entry in the mixed idle list that does not have a submitted status between the flush point and the head pointer.

11. The processor of claim 10, wherein the operation further includes updating the head pointer of the mixed free list based on the flush point.

12. The processor of claim 10, wherein the flush point is based on the oldest assigned PRN affected by the flush.

13. The processor of claim 12, wherein the operation further comprises: Store the checkpoints for each PRN-to-LRN mapping in the reorder buffer (ROB); as well as The checkpoints in the ROB are used as the flushing points.

14. A method executed by a processor, the method comprising: Maintain a hybrid free list, which represents physical registers available for register renaming, wherein the hybrid free list comprises multiple ordered entries, each entry comprising a pair of bits for each of the multiple physical registers. Each entry includes a pair of bits representing the valid and committed states of the physical register, respectively. The valid state indicates whether the physical register is free and can be allocated for a logical register name, and The commit status indicates whether the instruction using the physical register has been committed; as well as The head and tail pointers of the free list are used to perform the sequential allocation of free physical register names based on the multiple entries of the free list.

15. The method of claim 14, further comprising: Use the head pointer to find entries that are in an idle state.

16. The method of claim 15, wherein the entries of the hybrid free list are arranged in multiple storage units, and The entries identified as having an idle state include: Read the memory bank corresponding to the head pointer; as well as Identify entries in the memory that are in an idle state.

17. The method of claim 16, further comprising: The output temporary buffer is filled based on the valid status of the entries in the storage. as well as Use the entries in the output temporary buffer to allocate physical register names.

18. The method of claim 17, wherein filling the output temporary buffer comprises: The complete physical register name is generated based on the input temporary buffer that stores the valid bits of the entries from the memory bank.

19. The method of claim 18, wherein generating the complete physical register name includes expanding each bit in the input temporary buffer to a multi-bit physical register name in the output temporary buffer.

20. The method of any one of claims 14 to 19, further comprising: The validity state of the entry corresponding to the physical register name is modified in response to the assignment of the physical register name.

21. The method of any one of claims 14 to 20, further comprising: When an instruction that uses a physical register has been submitted, the submission status of the entry corresponding to the physical register is modified.

22. The method of any one of claims 14 to 21, further comprising: When a subsequent instruction is written to a logical register name, the validity state and the commit state of the entry corresponding to the physical register name assigned to the same logical register name are modified.

23. The method of any one of claims 14 to 22, further comprising: Receive instructions for flushing at a specific flushing point; Modify the valid status of any entry in the mixed idle list that does not have a submitted status between the flush point and the head pointer.

24. The method of claim 23, further comprising updating the head pointer of the mixed free list based on the flush point.

25. The method of claim 23, wherein the scour point is based on the oldest assigned PRN affected by the scour.

26. The method of claim 25, further comprising: Store the checkpoints for each PRN-to-LRN mapping in the reorder buffer (ROB); as well as The checkpoints in the ROB are used as the flushing points.