Power reduction using free lists for register renaming

CN122804220APending Publication Date: 2026-09-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480088708.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,对于常规任务,位向量空闲列表需要非常复杂的控制逻辑

Benefits of technology

[0010]头指针可以用于高效刷新恢复。当在刷新恢复期间恢复PRN时,简单地将头指针重定向到刷新点,并且重置刷新点与旧头指针之间的PRN的有效位,如由其提交位所指示处于架构化状态的PRN除外。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122804220A_ABST
    Figure CN122804220A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, for reducing power to physical registers using a free list for register renaming. One of the methods includes maintaining, in a processor, a free list that tracks availability of physical registers of a physical register file, where the free list is partitioned into a plurality of regions, each region corresponding to a subset of the physical register file. Power is reduced to one or more physical registers in the physical register file corresponding to an inactive region of the free list if the region of the free list represents an inactive region of the physical register file.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Modern out-of-order (OOO) processing utilizes register renaming to eliminate pseudo-data dependencies, enabling the parallel execution of more instructions, which typically involves out-of-order instruction execution. In this specification, register renaming refers to assigning a physical register name (PRN) to a logical register name (LRN) within an instruction being executed. Using register renaming, subsequent instructions that write to a logical register can execute before or in parallel with other instructions that use or read from the same logical register name, because the register renaming process assigns a different underlying physical register for that logical register name.

[0002] The "free list" is a microarchitectural component of the processor that maintains a pool of available physical resources (e.g., physical registers). In this specification, available resources (e.g., physical registers) will be described as being in an idle state, or the corresponding PRN will be a free PRN. Therefore, the free list can provide free PRNs to allocate physical registers for register renaming. When a PRN is no longer needed, it can be returned to the free list for reuse.

[0003] The size of the free list and the complexity of the control logic are major considerations in processor design, affecting silicon area, power, and performance in machines using OOO processors.

[0004] Traditionally, there have been two main approaches to implementing free lists in OOO processors—using a First-In-First-Out (FIFO) queue and using a bit vector. In a FIFO free list, PRNs are allocated sequentially from the beginning of the list, and PRNs are returned to the end when no longer needed. Therefore, in a pure FIFO free list, PRNs are allocated based on age, where older PRNs that have not been allocated for a longer period are prioritized over newer PRNs that were recently allocated. However, because entire PRNs, which can be 10 bits or longer, can be stored in the FIFO, returning a PRN to the FIFO free list requires writing a multi-bit PRN entry. Therefore, frequently writing multi-bit PRN entries to the FIFO free list requires significant area and power, posing serious problems for performance, cost, and scalability.

[0005] In contrast, bit-vector free lists use a single bit to represent each PRN, and therefore have smaller storage requirements than FIFO free lists. However, for routine tasks, bit-vector free lists require very complex control logic. For example, performing flush recovery (e.g., after predicting a branch that has been wrong) in a bit-vector free list requires significantly more complex logic compared to a FIFO free list, because the bit-vector free list does not contain age information for each PRN. Furthermore, as the total number of PRNs increases, the complexity of finding available PRNs also increases, further driving area and power requirements. For instance, when multiple available entries are provided from the bit-vector free list, finding the first available PRN or the first pair of available PRNs becomes more challenging as request frequency and vector length increase, representing a significant scalability bottleneck. Summary of the Invention

[0006] This specification describes techniques related to hybrid free lists for fast and efficient register renaming, which offer the benefits of both FIFO and bit-vector free lists, as well as many other additional advantages. The hybrid free list maintains a set of PRN entries that require only a small number of storage bits. As an example, each PRN entry may have: 1) a valid bit, indicating whether the corresponding PRN is free or has been allocated; and 2) a commit bit, indicating whether the PRN is in a schemated state, meaning that the data in the corresponding physical register is part of the machine's permanent state at a given point in time, and the corresponding instructions are no longer being flushed or re-executed.

[0007] In this specification, a PRN entry is a structure that stores data from which a PRN can be generated based on the location of the PRN entry in the free list. In this specification, the free list is described as generating a PRN meaning that a PRN is generated using the location of a PRN entry, and does not imply that the free list itself stores the complete PRN.

[0008] The hybrid free list maintains a head pointer, which indicates the position in the free list for searching the next free PRN. In some implementations, the head pointer advances along free list segments with multiple PRN entries rather than along individual PRN entries. The hybrid free list also maintains a tail pointer, which indicates the entry corresponding to the youngest free PRN that is available for reassignment. PRNs before the head pointer and after the tail pointer can be assigned in order, skipping PRNs that have already been assigned and have not yet been returned.

[0009] The tail pointer is typically used to protect the ordered arrangement of allocated PRN entries. Therefore, when allocating PRN entries, the head pointer is not allowed to exceed the tail pointer. By maintaining the head and tail pointers, the processor can essentially perform cyclic allocation of PRNs by iteratively traversing the free list in an ordered manner, skipping already allocated PRNs.

[0010] Head pointers can be used for efficient refresh recovery. When recovering a PRN during refresh recovery, simply redirect the head pointer to the refresh point and reset the valid bits of the PRN between the refresh point and the old head pointer, except for PRNs in a schematized state as indicated by their commit bits.

[0011] Specific embodiments of the subject matter described herein can be implemented to achieve one or more of the following advantages. The hybrid free list combines the benefits of both FIFO and bit-vector free lists while mitigating their respective disadvantages. For example, during refresh recovery, the hybrid free list can efficiently clear allocated PRNs. Additionally, the hybrid free list contains simplified PRN information that requires less storage space. Therefore, the hybrid free list achieves area and power reductions that would not be achievable in an equivalent FIFO free list. Furthermore, the hybrid free list is fully scalable because refresh recovery does not become disproportionately more challenging as the number of PRNs increases, unlike in a strict bit-vector free list. The free list described herein can also efficiently allocate groups of PRNs to meet the bandwidth requirements of a superscalar processor. For example, the free list can continuously fill input and output buffers such that N available and complete PRNs are queued and ready in each cycle.

[0012] The hybrid free list described in this specification can also be used for register renaming for multiple different types of physical registers. This allows the processor to maintain a single free list instead of multiple free lists, thereby saving power and reducing the required silicon area. For example, the free list can be used to rename registers for both integer registers and flag registers. The free list can also be used to rename registers for both integer registers and floating-point registers, even when the sizes of the corresponding physical register files differ. Furthermore, even when the free list is used to rename registers for multiple different types of physical registers, the basic scalable refresh recovery mechanism remains largely unchanged and can be used to return the names of all types of physical registers to the free list after a refresh.

[0013] This specification also describes how a free list can be used to borrow already allocated physical register names. This arrangement improves processor efficiency, reduces the likelihood of the processor running out of register space, and all without requiring any additional naming storage.

[0014] This specification also describes how a processor can improve floating-point instruction bandwidth by introducing a new microarchitecture structure (i.e., a structured floating-point physical register file), which is an auxiliary physical register file used to maintain the values ​​of floating-point registers whose names are locked in a structured mapping table.

[0015] This specification also describes techniques for saving processor power through clock gating using an ordered structure of hybrid free lists. Because the free list inherently maintains information about which physical registers are in use during execution, the processor can use this information to perform highly targeted clock gating on unused physical registers, thereby saving power without significantly impacting processor performance.

[0016] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the description, drawings, and claims. Attached Figure Description

[0017] Figure 1A This is an overview of the sample lifecycle of a PRN in a mixed free list.

[0018] Figure 1B This is an example overview of mixed free lists.

[0019] Figure 2 This is a sample procedure for allocating PRNs from a mixed free list.

[0020] Figure 3 This is an example procedure for deallocating a PRN to a mixed free list.

[0021] Figure 4 An example refresh recovery using a mixed free list is shown.

[0022] Figure 5 This is a sample procedure for implementing a mixed free list.

[0023] Figure 6 This is a diagram of a processor system that has a shared free list that uses a single physical register namespace for two different physical register files.

[0024] Figure 7 This is an example of how the system can maintain a single shared free list for integer registers and flag registers.

[0025] Figure 8 This is a flowchart of an example process for maintaining a shared list of free spaces.

[0026] Figure 9 An example of PRN borrowing in a system with two sets of physical registers is shown.

[0027] Figure 10 This is a flowchart of an example procedure for borrowing physical register names.

[0028] Figure 11 An example technique for associating PRN entries with floating-point registers is illustrated.

[0029] Figure 12 It is a diagram illustrating how the system can maintain a shared free list between integer registers and floating-point registers.

[0030] Figure 13 This is a flowchart of an example procedure for maintaining virtual PRN entries during refresh recovery when managing multiple different register sets.

[0031] Figure 14 This is a flowchart of an example procedure for allocating virtual PRNs to manage free lists of multiple different register sets.

[0032] Figure 15 This example shows a sliding window for PRNs in the free list.

[0033] Figure 16 The data structure used to handle stuck floating-point PRNs is illustrated.

[0034] Figure 17 This is a flowchart of an example procedure for managing a virtual PRN for managing a free list of multiple different register sets.

[0035] Figure 18 An example of a mixed free list with 512 entries is shown, which are divided into storage bodies, each with 16 entries.

[0036] Figure 19 This is a flowchart illustrating an example of clock gating of a physical register file based on a free list.

[0037] In the various figures, the same reference numerals and designations indicate the same elements. Detailed Implementation

[0038] Figure 1A This is an overview of a sample lifecycle 100 of the PRN maintained by a hybrid free list. Lifecycle 100 consists of three phases maintained by three corresponding structures: Hybrid Free List Phase 102, Reorder Buffer (ROB) Phase 130, and Architectural Map (AMT) Phase 140. Throughout the OOO process, the PRN cycles between these three phases.

[0039] The PRN begins in an idle state in free list stage 102. The idle PRN can then be allocated from the mixed free list, renamed, and transitioned to ROB stage 130, which stores the mapping between the PRN and its corresponding LRN in instruction order. When the PRN transitions from free list stage 120 to ROB stage 130 or AMT stage 140, it is typically no longer idle and therefore cannot be allocated to another logical register until it returns to free list stage 120. However, in certain scenarios described in detail below, the PRN can be reassigned to another logical register before returning to the free list.

[0040] The ROB can also associate each PRN-LRN mapping with a corresponding position in the free list, so that when a refresh occurs, the head pointer of the free list can efficiently and quickly jump back to the refresh point.

[0041] When the instruction corresponding to an entry in the ROB has completed execution and reached the head of the ROB, the instruction is committed by retiring the PRN to the AMT, which also stores the mapping between the PRN and its corresponding LRN. The data stored in the physical registers corresponding to the PRN is now part of the machine's permanent state and is not subject to flushing or re-execution.

[0042] When subsequent instructions are written to the logical register of the PRN maintained in the AMT, the PRN mapping is no longer required, and the PRN is evicted from the AMT and looped back to the hybrid free list, in which the PRN re-enters the hybrid free list stage 102.

[0043] A hybrid free list can maintain PRN entries, each with a pair of values ​​(e.g., a pair of bits) to track the valid and committed status of each PRN. The "valid" bit can be used to track the valid status of the PRN, indicating whether the PRN is free for renaming. Therefore, the valid status of a PRN can be referred to as being in an idle state or an allocated state.

[0044] The "Commit" bit is used to track the commit status of the PRN. This commit status indicates whether instructions using the allocated PRN have been committed. Commit means that the value in the corresponding physical register has become the permanent state of the program being executed by the processor and will no longer be flushed or re-executed. Therefore, the commit status of the PRN can be referred to as committed or uncommitted. The change from uncommitted to committed status typically corresponds to the PRN leaving the ROB and entering the AMT.

[0045] Therefore, in this specification, a PRN or its corresponding PRN entry may be referred to as having an idle state or an allocated state, which may indicate the value of the valid bit of the PRN entry. Similarly, a PRN or its corresponding PRN entry may be referred to as having a committed state or an uncommitted state, which may indicate the value of the committed bit of the PRN entry.

[0046] In some cases, the valid and commit bits have binary values, where "1" indicates a set value and "0" indicates a cleared value. A set value for the valid bit (e.g., "1") can indicate that the PRN has been allocated for renaming, while a set value for the commit bit can indicate that the PRN has been retired to AMT. In some cases, clearing the valid bit value (e.g., "0") and setting the commit bit value (e.g., "1") indicates an illegal or improper state, as this would mean that the PRN was committed before it was allocated. This situation is not expected to occur, but if it does, the processor can enter a fault state to restore the integrity of the free list. Figure 1A A table is presented showing example combinations of valid bit values ​​and committed bit values. In alternative implementations, values ​​that are considered set and clear, or high and low, can be interchanged.

[0047] In some implementations, PRN availability can be dynamically constructed rather than explicitly represented. For example, the free state of each PRN can be determined using a free list head pointer and a free list tail pointer instead of maintaining an explicit commit bit for each PRN. In this case, the head pointer indicates the region in the free list used to search for PRNs to be allocated. Alternatively, the commit state of a PRN can be constructed by examining the AMT instead of maintaining a commit bit. However, when the processor has hundreds of physical registers, simply maintaining an explicit commit bit in the free list instead of searching the AMT is often faster.

[0048] Figure 1B This is an example overview 101 of a hybrid free list 132. Hybrid free list 132 comprises a collection of PRN entries representing PRNs. Each PRN entry has a valid bit 104 and a commit bit 106. Additionally, hybrid free list 132 includes a head pointer 108 and a tail pointer 110.

[0049] When a PRN needs to be allocated in OOO processing, the mixed free list 132 will use the head pointer 108 to find the PRN entry that can be allocated. Although Figure 1B A PRN entry is indicated at the head pointer 108 in the hybrid free list 132, but the head pointer in the hybrid free list can be implemented as a coarse-grained pointer that steps through segments with multiple PRN entries. For example, Figure 1BThe diagram also illustrates how the set of PRN entries in the free list 132 can be implemented as a memory bank 116, with the head pointer stepping through memory banks. In some implementations, the size of each memory bank is based on the processor's bandwidth. For example, if the processor might need to rename ten registers per cycle, the free list could have a memory bank size of ten PRN entries, or, to account for PRNs still in the committed state, a memory bank size of more than ten PRN entries. Typically, each instruction is written to only one logical register, and therefore, the processor's instruction bandwidth typically corresponds to the number of PRNs needed to be generated per cycle.

[0050] When a PRN is retired from the ROB, the tail pointer 110 is updated. The tail pointer 110 is updated when the next PRN entry is retired from the ROB by having a status of {0,0} or {1,1} for its valid and committed bits. This means that the PRN has been retired from the ROB but has not entered the AMT, or the instruction using the PRN has been committed and the PRN has entered the AMT. Alternatively or additionally, as the tail pointer steps through the memory bank of the PRN entry, it is updated when all entries in the next memory bank have a status of {0,0} or {1,1}. Therefore, the tail pointer 110 indicates the youngest PRN entry or the memory bank of the PRN entry that can be reassigned.

[0051] In contrast, the head pointer indicates the oldest PRN entry or the storage of PRN entries that can be allocated, which corresponds to the PRN that has not been used within the longest period number.

[0052] As mentioned above, the system maintains a tail pointer 110 to keep the PRN entries in the free list in an ordered order by preventing the head pointer 108 from overtaking the tail pointer 110. This is important for maintaining or deriving age information about PRN entries, because if a structured PRN returned to the free list is reassigned, determining the age of the corresponding PRN entry during a refresh may be impossible or impractical. Maintaining the tail pointer 110 also improves power performance by preventing the head pointer and the temporary buffer from being constantly updated over many PRN entries that are not free to be assigned.

[0053] When a PRN entry is read from the mixed free list 102, a decoding process is performed to generate the corresponding PRN from the PRN entry's location in the free list. This process saves storage space because it allows the free list to efficiently store two bits per PRN instead of storing all bits of each PRN.

[0054] The process of using pointers to step through the PRN memory bank is highly scalable. In other words, any number of PRNs can be allocated from the mixed free list using the head pointer 108 and the tail pointer 110 without disproportionately increasing processing requirements.

[0055] In some implementations, the free list continuously fills one or more scratch buffers, ensuring that there are always available PRNs ready in each cycle. The input scratch buffer receives PRN entries from the memory bank indicated by the head pointer. The output scratch buffer is then filled with the complete physical register number corresponding to the PRN entries in the input scratch buffer. At either the input or output scratch buffer, PRN entries that have been allocated, as indicated by their valid bits, can be removed. In some cases, the input scratch buffer includes only the valid bit 104 value for each PRN entry already identified from the mixed free list 102. In other words, the input scratch buffer can essentially be a bit vector representing which PRNs are free in the memory bank indicated by the head pointer. The output scratch buffer is then filled with complete PRNs using the positioning of the set bits in the input scratch buffer and the positioning of the head pointer. For example, the output scratch buffer may include a memory bank with 16 different PRNs. In other cases, the buffer may be expanded to include different numbers of PRNs and / or bits. References will follow below. Figure 2 The operation of the temporary buffer will be explained in more detail.

[0056] Figure 2 This is a diagram of an example procedure 200 for allocating a PRN from a mixed free list 202. This example procedure, as well as other procedures described herein, can be implemented using a digital logic circuitry system of a processor (e.g., an out-of-order processor).

[0057] In this example, for ease of illustration, each PRN entry in free list 202 is represented as having a single valid bit, although, as mentioned above, mixed free list entries typically also have a commit bit. The processor includes input buffer 204 and output buffer 206. In some cases, when a PRN memory bank identified by the head pointer is read, the allocatable PRNs within that memory bank (e.g., valid bit 104 is cleared, such as "0", or an alternative availability state) are expanded into a series of PRNs in input buffer 204 and output buffer 206. This is achieved by placing the PRN entry from the memory bank indicated by the head pointer into input buffer 204. For example, as... Figure 2 As shown, the input temporary buffer 204 may include 16 bits, where each bit corresponds to a valid bit value from a different PRN entry in the memory bank indicated by the header pointer 108.

[0058] Next, the physical register names corresponding to the PRN entries in input buffer 204 are populated into output buffer 206. The physical register names can be generated by locating the PRN entry in the free list. In some implementations, this location can be reconstructed by adding the offset of the PRN entry in the input or output buffer, or in the memory bank of the PRN entry, to the location of the head pointer. Assuming a 10-bit PRN length, output buffer 204 can include 10 bits of PRN for each free PRN entry in input buffer 204, thus converting the PRN entry into an actual PRN. For example, the system can generate up to 16 complete 10-bit PRNs in output buffer 206 from 16 bits in input buffer 204.

[0059] As part of this process, “holes” in the input buffer caused by PRNs that are unavailable due to their valid bits are essentially squeezed out, so that the output buffer only stores PRNs with cleared valid bits from the PRN entries.

[0060] The processor can then rename the available PRNs in the output buffer 204, for example, by assigning the PRN to a logical register name and storing each assigned PRN-to-LRN mapping in ROB 302.

[0061] Figure 3 This is an example procedure 300 for releasing a PRN from the ROB and / or from the AMT 304 to the mixed free list 102. As described above, after a PRN has been allocated, the value of the associated PRN entry valid bit 104 can be set to "1".

[0062] During OOO processing, the PRN-to-LRN mapping between the PRN and LRN remains in ROB 302 until the associated instruction is committed, thus retiring the PRN-to-LRN mapping to the AMT, or until the associated instruction is flushed. When a PRN is retired, the commit bit 106 of the associated PRN entry is set to "1" (operation 306), indicating that the PRN instruction has been committed and the PRN has been cycled into AMT 304 (operation 308). As discussed above, AMT 304 maintains the PRN-to-LRN mapping for committed PRNs, and a committed PRN entering AMT 304 for the same logical register will evict any previous PRN mapping for that logical register. Therefore, the size of AMT 304 is typically based on the number of logical registers in the instruction set, while the size of ROB 302 is typically much larger because the number of physical registers in the instruction set is usually far greater than the number of logical register names. When the PRN is evicted from AMT 304, the values ​​of the associated valid bit 104 and commit bit 106 are cleared, for example, returned to "0", thus indicating that the PRN is now available for allocation (operation 310).

[0063] Figure 4 An example refresh recovery using the hybrid free list 102 is illustrated. In some cases, the PRN needs to return from ROB 302 to the hybrid free list 102, or be deallocated from that ROB. This process is called refresh recovery. For example, refresh recovery can be performed when instructions are speculatively executed along an incorrect control flow path (e.g., a branch that is predicted incorrectly).

[0064] During refresh recovery, a refresh point is specified at a point in the mixed free list 102, and the head pointer 108 is redirected to this new point. In some implementations, the refresh point is the point at which the writer of a particular PRN and all updated PRNs is invalidated from the ROB. Any PRN after the refresh point and before the original head pointer 108 (e.g., the head pointer 108 before redirection) will have its valid bit 104 cleared, for example, cleared to "0", except for PRNs in AMT 304 that are still in use as indicated by their commit bit being set.

[0065] In some cases, to support efficient refresh recovery, the PRN pointer can be continuously checked during the dispatch process (operation 402). For example, during dispatch, the current head pointer 108 plus an offset value can be associated with each PRN entering ROB 302. Alternatively or additionally, the PRN pointer can also be stored in other structures linked to the ROB, or in structures tracking instructions in flight, such as reservation stations or dispatch queues, to name just a few. In other words, the system can use any suitable structure to collect PRNs that have been assigned but not yet committed in the AMT and will return the corresponding PRN entries to the free list upon refresh.

[0066] When refresh recovery begins, the current checkpoint is indicated by the oldest PRN to be refreshed from ROB 302. The head pointer of the free list can then be reset to the refresh point using a pointer to that ROB entry (operation 404). If the head pointer points to a bank of PRN entries, offsets can be used to determine which PRN entries in the bank are part of the refresh. In some implementations, the input and output scratch buffers are also cleared, allowing PRN data to be refilled into the input and output scratch buffers from the new refresh point. In some cases, the refresh recovery process is constrained during allocation time to prevent conflicts with ROB 302.

[0067] By maintaining checkpoints, refresh recovery can be performed efficiently without exhaustively searching for affected PRNs during the refresh. The refresh recovery process in the hybrid free list 102 is also highly scalable because the performance of refresh recovery is minimally affected as the free list size increases. In other words, if the free list grows tenfold, the refresh recovery process will still be just as fast. This is very different from bit vector free lists, where refresh recovery becomes more computationally intensive as the number of PRNs increases, and more complex and time-consuming as the ROB size increases. Using checkpoints in the ROB also avoids the need for side-tracking of in-flight PRNs, for example, by the ROB or other structures, to determine and indicate the set of PRNs affected by the refresh (as is the case in bit vector free lists).

[0068] Figure 5 This is a sample procedure 500 for implementing a mixed free list. This procedure can be executed by a processor configured according to this specification. For convenience, this procedure will be described as being executed by the system.

[0069] The system uses the free list header pointer to locate PRNs (510) that have an idle status (e.g., a valid bit with a value of "0"). As described above, a PRN in the mixed free list is associated with both a valid bit and a commit bit indicating which stage the PRN is in. The valid bit indicates whether the PRN is currently in use.

[0070] The hybrid free list also maintains a head pointer and a tail pointer. The head pointer indicates the next PRN entry or set of entries to be allocated, while the tail pointer indicates the next PRN to be decommissioned, currently being written to by the oldest in-flight instruction written to the register. In some implementations, the head pointer points to the memory bank of the PRN entry, and the head pointer steps on each memory bank of the PRN entry.

[0071] The system allocates a PRN and modifies its idle state (520). The PRN is then allocated, and the valid bit is updated to indicate that the PRN is no longer idle, for example, by setting the valid bit value to "1". As described above, once a PRN is allocated from the mixed idle list, it is cyclically passed to the ROB, which maintains an ordered list of instructions to which PRNs have been allocated, along with the PRN-to-LRN mapping for those instructions.

[0072] The system retires the PRN and modifies its commit status (530). Upon commit, the PRN is retired, and the PRN-to-LRN mapping is stored in the AMT. When this occurs, the commit bit of the PRN entry is modified to indicate that the PRN has been retired (e.g., by setting the value of the commit bit to "1"). As described above, when a PRN transitions to the AMT, it can evict older PRNs assigned to the same LRN for older instructions.

[0073] The system evicts the PRN from the AMT and modifies the PRN's idle status (540). When the PRN is evicted from the AMT and returned to the mixed free list, both the valid bit and the commit bit are updated to indicate that the PRN is idle; for example, the values ​​of both the valid bit and the commit bit can be set to "0".

[0074] This specification also describes how free lists can be used for various types of physical registers, each with its own physical register namespace. For example, some processors have different physical register files, each with its own namespace. Instructions can reference different logical register names corresponding to different underlying physical register files. While maintaining a separate free list for each corresponding physical register file would be possible, in many situations this would lead to hardware inefficiency and redundancy.

[0075] Conversely, the techniques described in this specification allow the processor to use a single shared free list to maintain and allocate physical register names for multiple different register files. When an instruction execution requires a physical register name for register renaming, the free list can allocate the physical register name using a single shared namespace applicable to both types of physical register files. Furthermore, for instructions that reference multiple different types of registers, the free list can allocate the same physical register name for both types of logical registers in the instruction.

[0076] A key benefit of using a shared free list for various types of physical register files is that only a flush recovery mechanism needs to be executed once during a flush. Furthermore, sharing the free list among different types of registers saves hardware resources without significantly impacting performance.

[0077] Figure 6 This is a diagram of a processor system 600, which has a shared free list that uses a single physical register namespace for two different physical register files. System 600 includes a first physical register file 610 with physical registers PR1, PR2 through PRN. System 600 also includes a second register file 620 with physical registers FR1, FR2 through FRN.

[0078] System 600 also includes a shared free list 630 that maintains physical register name entries PRN1, PRN2 to PRNN. The PRN entries in this single shared namespace can be used to rename registers for both the first physical register file 610 and the second physical register file 620.

[0079] Typically, each physical register file has its own Schematic Mapping Table (AMT). Therefore, physical register file 610 will have its own AMT, and the second physical register file 620 will have its own AMT. Because the two physical register files share the same namespace used for register renaming, the system can ensure that a physical register name is not returned to the free list 630 unless it is evicted from either AMT or does not exist in either AMT.

[0080] An example of a separate type of physical register file is a special register file. In some processors, special physical registers are used to maintain information about the processor's state. These special registers are typically maintained in a separate physical register file from the physical register files for integers or floating-point numbers.

[0081] An example of a special register is the flag register, which is sometimes referred to as the application status register. In this specification, "flag register" will refer to any special register that can be referenced by a logical register name in an instruction set that has a physical register file different from an integer register file or a floating-point register file.

[0082] In some processors, there may be only one logical flag register that an instruction can write to, but there may be dozens or hundreds of physical flag registers that can be allocated to that single logical flag register. This is because, during OOO execution, flag registers can be used to maintain the conditional state of many different in-flight instruction branches. Therefore, even if the instruction set has only one logical flag register, register renaming can be used to maintain the state of all in-flight conditional branches.

[0083] Figure 7 This is an example of how system 700 can maintain a single shared free list 730 for integer registers and flag registers. System 700 includes a set of integer registers 710 and a set of flag registers 720. The system also includes a hybrid free list 730 that maintains a single physical register namespace for both integer registers 710 and flag registers 720. The hybrid free list 730 includes PRN entries, each including a valid bit and a commit bit. Integer register file 710 has its own integer AMT 740. Flag register file 720 also has its own flag AMT 750.

[0084] As described above, the valid bit indicates whether a physical register name is free for register renaming, or whether a physical register name has already been allocated. The commit bit indicates whether a physical register name is currently still in the AMT. As mentioned above, these two bits determine how the system allocates physical register names, and also how the system handles returning physical register names to the free list upon refresh, which essentially involves returning PRNs allocated after the refresh point but not yet in the AMT.

[0085] In this example system, the integer register and the flag register share the free list 730 by using the valid bits explicitly represented in the PRN entries of the free list 730.

[0086] System 700 can use the integer commit bit 706 of a PRN entry to track the commit status of a PRN allocated to the integer register. However, for flag PRNs, System 700 can reconstruct the virtual flag commit bit from the PRN entry and from the status of the flag AMT 750. This eliminates the need for the system to maintain a separate free list for the flag register and keeps the size of the PRN entry to two bits per PRN entry.

[0087] In this example, the Flag AMT 750 has only a single entry because, in this example, the flag register has only a single logical register in the instruction set. Therefore, the physical register-to-logical register name mapping in the Flag AMT 750 requires only a single entry. Furthermore, when only a single logical flag register exists, the Flag AMT 750 only needs to store the physical register name, without storing the complete mapping from physical register name to logical register name. Therefore, the Flag AMT 750 is exemplified as storing only a single PRN instead of a mapping. In contrast, the Integer AMT 740 has as many integer entries as the existing integer logical registers, and the Integer AMT 740 stores the mapping between physical register name and logical register name for each entry.

[0088] During allocation, physical register names can be assigned based on their valid status. For integer registers and flag registers, this means checking the valid bits in the free list 730. However, to prevent naming conflicts, the system also performs a bit-wise OR operation between the valid status indicated by the PRN entry and the committed bit decoded from the flag AMT 750. This prevents the reassignment of the same PRN until that PRN has been evicted from either of the two AMTs 740 and 750 or otherwise does not exist in either AMT.

[0089] The decoding module 770 can be used to generate a commit bit for decoding for any PRN entry in the free list 730 by checking whether the corresponding PRN is still in the flag AMT 750. In other words, for a given PRN, if the PRN is in the flag AMT 750, the decoding module 770 can generate a 1 for the commit bit of decoding; otherwise, it can generate a 0 for the commit bit of decoding.

[0090] Then, system 700 can generate a virtual valid bit 772 for the PRN from the decoded commit bit by performing a bitwise OR operation between 1) the decoded commit bit and 2) the valid bit of the corresponding PRN entry in the free list 730. In practice, this means that when the integer register and the flag register appear in the same instruction, they can be assigned to the same PRN simultaneously.

[0091] Similarly, system 700 can generate a virtual flag commit bit 774 for the PRN by performing a bitwise OR operation between 1) the decoded commit bit and 2) the commit bit of the corresponding PRN entry in the free list 730.

[0092] As described above, the physical register name can become free again when the PRN is no longer in the integer AMT 740 or the flag AMT 750. For the integer AMT 740, the system can use the commit bit of the PRN entry in the free list 730 to determine whether the physical register name is still maintained in the integer AMT 740.

[0093] On the other hand, because the Flag AMT 750 maintains only one logical flag register, the system can save a considerable amount of storage space by virtually determining the flag commit bit from the decoded commit bit. Therefore, for the flag PRN, the system can determine the commit status by performing a bitwise OR operation between 1) the commit bit of the corresponding PRN entry and 2) the decoded commit bit.

[0094] Table 1 summarizes the status of the PRN entries shared between the integer register and the flag register.

[0095]

[0096] Table 1

[0097] The operation for maintaining the shared free list 730 will now be described. During operation 1, a physical register name can be assigned by identifying PRN entries with a free state in the free list. As mentioned above, PRN entries that can be renamed can be identified by using an input scratch buffer and an output scratch buffer, which can also be used to locate the generated PRN from the free list of the corresponding PRN entry.

[0098] Because the integer register file 710 and the flag register file 720 share a namespace, if an instruction references a logical register for an integer or a logical register for a flag register, the system simply assigns a physical register name to any free PRN entry in the free list 730.

[0099] In some implementations, the system can assign the same PRN to both the integer register and the flag register. For example, if a single instruction references both the integer logical register and the flag logical register, the system can search for an available PRN from the free list and assign the same PRN to both. Although the integer logical register and the flag logical register are assigned the same physical register name, the values ​​written to the respective logical registers will be stored by different physical registers because the logical register names point to different physical register files 710 and 720.

[0100] Then, during instruction execution, the PRN-to-LRN mapping is stored in ROB 760. In some implementations, ROB760 includes fields 711 and 712 indicating whether the mapping is an integer mapping, a flag mapping, or both. Because the system assigns the same physical register names for instructions referencing both the integer register file 710 and the flag register file 720, ROB760 only needs two bits—flag bit 711 and integer bit 712—for each ROB entry to indicate whether the entry is an integer mapping, a flag mapping, or both.

[0101] Operation 2 indicates a refresh recovery. Because the integer register file 710 and the flag register file 720 share the free list, only a single refresh operation is required for the free list 730. As described above, the head pointer of the free list 730 will quickly jump back to the refresh point indicated by the checkpoint in ROB 760, as described above. Any allocated PRN entries that are after the refresh point and before the head pointer and whose commit bits are not set will return to the free list, for example, by resetting their valid and commit bits.

[0102] Operation 3 represents a bypass return to the free list 730. This is a special case to avoid unnecessarily storing the PRN in the AMT. For example, if a physical register mapped to a logical register is decommissioned in the same cycle as an update instruction written to the same logical register, there is no need to add that physical register to the AMT, because the physical register will be overwritten by the update writer. Therefore, the physical register name can bypass the AMT and return directly to the free list. In doing so, the system can clear these bits, for example, by setting the valid bit and the integer commit bit to 0.

[0103] During operation 4, when retiring a PRN from ROB 760, if the PRN is only an integer PRN, it will be retired to integer AMT 740. The system can also set the integer commit bit, for example, by setting the integer commit bit of the corresponding PRN entry to 1. If the PRN is only a flag PRN, it will be retired to flag AMT 750. If the PRN is both an integer PRN and a flag PRN, it will be retired to both AMTs 740 and 750.

[0104] Finally, the PRN can be removed from its corresponding AMT 740 and 750.

[0105] During operation 5, the integer PRN is evicted from the integer AMT 740. The system can clear the valid bit and integer commit bit of the corresponding PRN entry from the free list 730 to indicate that the PRN is no longer in the integer AMT 740.

[0106] However, the PRN will not actually become free for allocation unless and until it is not stored in the flag AMT 750. As described above, the system 700 can enforce this constraint using a virtual valid bit 772, which is a bitwise OR operation between the valid bits of the PRN entry and the commit bits of the decoded data from the flag AMT 750.

[0107] When a flag PRN is evicted from flag AMT 750, the system can reset the valid bits of the corresponding PRN entry. When this happens, flag AMT 750 will store a different PRN, so the decoded commit bit for that PRN will change from 1 to 0. Therefore, at this time, unless the integer PRN is still in integer AMT 740, the virtual flag commit bit of that PRN will be reset to 0.

[0108] Figure 8 This is a flowchart of an example process for maintaining a shared free list. This process can be executed by a processor configured according to this specification. For convenience, the process will be described as being executed by the system.

[0109] The system maintains a shared free list for multiple types of register sets (810). As mentioned above, the processor can have different physical register files, each with its own physical register names. For multiple physical register files, the system can maintain a single physical register namespace. As an example, the system can maintain a single physical register namespace for both the integer register and the flag register, which are maintained in separate physical register files.

[0110] The system generates a first physical register name (820) for a first type of physical register and a second physical register name (830) for a second type of physical register. In other words, even if the physical registers are in different physical register files, the system can generate physical register names from the same shared free list. Furthermore, as described above, if an instruction references a physical register in two physical register files, the system can generate a shared physical register name for the two physical registers in different physical register files.

[0111] The system deregisters the first physical register name and the second physical register name to different corresponding AMTs (840). As mentioned above, different register files have different AMTs for maintaining the processor state for submitted instructions.

[0112] If the first physical register name is not in the second AMT, the system designates the first physical register name as free (850). In other words, because different physical register files are sharing a namespace, a physical register name should not become free if it is still maintained in one of the AMTs. Therefore, in order to return the first physical register name to the free list, the system may first check to ensure that the physical register name is no longer maintained in the second AMT. As mentioned above, in some implementations, the system may perform this check as a bitwise OR operation between the decoded commit bit and the valid bits in the PRN entry maintained in the free list.

[0113] This specification also describes how a shared free list can be used to borrow already allocated physical register names. In other words, in some situations, even if a physical register name has already been assigned to a logical register of an instruction, the same physical register name can be reassigned for another logical register. This arrangement improves processor efficiency, reduces the likelihood of the processor running out of register space, and all without requiring any additional named storage.

[0114] Figure 9 This example illustrates PRN borrowing in a system with two physical register sets 910 and 920. In this example, a shared free list 930 is used to allocate physical register names for a series of instructions 905, 915, and 925. In this example, the same physical register name is used for the logical registers referenced by instructions 905 and 925.

[0115] The addition instruction 905 adds the values ​​of logical registers X2 and X3 and stores the value in logical register X1. This requires allocating a physical register, so the shared free list 930 generates a physical register name (PRN2 in this example), which can be, for example, the binary value 0x10. The processor then uses this physical register name to identify the second physical register in the integer register set 910. Therefore, the result of the addition instruction is written to the second physical integer register 911.

[0116] Next, intermediate store instruction 915 is executed. Because store instructions do not need to write any value to any physical register, they do not need to allocate any physical registers via register renaming. Therefore, if the next instruction only references a logical register of a different type, the physical register name allocated for addition instruction 905 can be reassigned to the logical register for instruction 925.

[0117] Instruction 925 compares the values ​​of logic registers X5 and X6 and stores the result in the flag register. The syntax of instruction 925 implicitly references the flag register simply because it is a comparison instruction. Because instruction 925 needs to write to the physical flag register, the shared free list generates a physical register name that identifies one of the flag registers in the flag register file 920.

[0118] However, because the comparison instruction 925 meets the criteria for borrowing physical register names, the free list 930 reallocates the same physical register names generated by the free list for instruction 905 by regenerating PRN2. Therefore, the result of the comparison instruction will be stored in the second physical flag register (921).

[0119] Figure 10 This is a flowchart of an example procedure for borrowing physical register names. This procedure can be executed by a processor configured according to this specification. For convenience, the procedure will be described as being executed by the system.

[0120] The system maintains a shared free list (1010) for multiple different sets of registers. As mentioned above, the different sets of registers have different types and are referenced by logical register names that reflect those types.

[0121] The system assigns a first physical register name (1020) to the first physical register of the first type. The system can use the free list allocation procedure to generate the next available PRN based on the value of the PRN entry in the free list and the position of the head pointer.

[0122] The system receives a second instruction (1030) for a second type of physical register, and the system determines whether the second instruction satisfies one or more borrowing criteria (1030) with respect to the first instruction. In other words, the system can check the second logical register type against one or more previously received instructions to determine whether those instructions satisfy one or more borrowing criteria.

[0123] The first example of a borrowing criterion is when instructions are sequential. Therefore, in some implementations, the system checks the instructions against previously executed instructions to determine if they reference different types of physical registers. If so, the system can determine that the two instructions satisfy the borrowing criterion.

[0124] A second example of borrowing the standard is when one or more intermediate instructions are instructions that do not require register renaming. For example, a store instruction between two other instructions is an example of an instruction that does not require register renaming. Therefore, any number of store instructions can appear between these two instructions, and these store instructions will still satisfy the borrowing standard.

[0125] A third example of borrowing is whether instructions are renamed within the same cycle. As mentioned above, the use of input and output buffers allows the processor to perform multiple renamings within the same cycle. Therefore, in some implementations, the system can allow borrowing only between instructions that are renamed within the same cycle.

[0126] A fourth example of borrowing is whether a branch instruction exists between the two instructions. In some implementations, to simplify refresh recovery logic, the system can disable borrowing if the intermediate instruction is a branch instruction. However, in some alternative implementations, borrowing can still be allowed even if the instruction exists, if the processor treats the intermediate branch instruction as non-borrowing after initiating refresh recovery. In other words, the system can allow borrowing before refresh recovery, but can disable borrowing for these instructions after refresh.

[0127] If one or more borrowing criteria are not met, the system generates a new physical register name for the second instruction (branching to 1050).

[0128] On the other hand, if one or more of the borrowing criteria are met, the system reallocates the same PRN (branching to 1060) for the second instruction. As mentioned above, when different instructions are executed using the same PRN, the results will actually go to different physical register files.

[0129] When two instructions have been assigned to the same PRN, the system ensures that both instructions are evicted from their respective AMTs before the PRN is returned to the free list. Therefore, when the first PRN is evicted from the AMT to a logical register by the newer writer, the system can check the commit status of the other PRN by checking the explicitly maintained commit bit or by searching its AMT. The shared PRN will only be returned to the free list if both PRNs have been evicted from their respective AMTs.

[0130] This specification also describes how the shared free list can be further extended to also manage physical floating-point registers without significantly increasing hardware cost. By sharing the free list between integer and floating-point registers, the processor can effectively eliminate the need for a separate free list for floating-point registers. Furthermore, the free list can still employ the same allocation and refresh recovery process with only a small amount of additional record keeping.

[0131] The complexity of sharing between integer and floating-point registers is greater than that of sharing with flag registers, because the number of floating-point registers is often less than that of integer or flag registers. The processor design is a result of the careful selection of the number of each type of register to achieve a specific level of performance and efficiency. Therefore, the relative number of physical registers in each register file varies considerably, but floating-point registers tend to be fewer than integer registers, partly because floating-point registers are larger in size. As mentioned above, the same number of integer and flag registers can exist efficiently because flag registers store far less data than integer registers; for example, 4 bits for a flag register versus 64 bits for an integer register. In contrast, the number of physical floating-point registers can be far less than the number of integer registers, for example, only a quarter or half the number of integer registers, to name just a few common examples.

[0132] To support sharing between integer and floating-point registers, the system can essentially allocate virtual PRNs for floating-point registers. In this specification, "virtual" PRN means that the namespace to which the PRN belongs does not correspond to the size of the physical register file. Instead, there is a mapping between each virtual PRN and each floating-point PRN in the floating-point PRN namespace, which does correspond to the size of the physical register file. When a free list is used to maintain the virtual PRN namespace, this means that every PRN entry in the free list can be allocated for a floating-point register, but because there are fewer floating-point registers than free list entries, multiple PRN entries with different virtual PRNs in the free list can be mapped to the same floating-point PRN. To avoid reallocating already allocated floating-point PRNs, the system can track which PRN entries are associated with each other by having virtual PRNs mapped to the same floating-point PRN. Therefore, if a PRN has already been allocated to a floating-point register, other associated PRN entries cannot be allocated to another logical floating-point register name until the allocated PRN is deallocated. However, during this period, the PRN of other associated PRN entries can still be assigned to other register types, such as integer registers. In the following description, PRN entries described as paired or associated PRN entries mean a group of multiple PRN entries that have different virtual PRNs mapped to the same floating-point PRN.

[0133] The techniques described below for sharing between integer and floating-point registers can also be combined with the techniques described above for sharing with flag registers and for PRN borrowing. Therefore, the shared free list described in this specification can be used to maintain PRN allocations for at least three different types of physical registers: integer registers, flag registers, and floating-point registers.

[0134] Figure 11 An example technique for associating PRN entries with floating-point registers is illustrated. In this example, the physical floating-point register file has half the number of physical registers as the integer physical register file. Therefore, the free list can maintain a virtual PRN namespace by associating two PRN entries with each floating-point PRN. To this end, the system can pair PRN entries in any suitable manner such that each pair of PRN entries identifies the same floating-point PRN. Figure 11 This illustrates a method for pairing PRN entries without requiring overly complex logic or record keeping.

[0135] like Figure 11 As shown, the free list with N PRN entries is conceptually divided into two halves: a first half 1110 spanning from PRN0 to PRN N / 2-1 and a second half 1120 spanning from PRN N / 2 to PRN N-1. Corresponding PRN entries in each half are paired together, such that they each identify the same floating-point PRN in the physical floating-point register file 1130 with N / 2 entries.

[0136] For example, both PRN entries for PRN0 and PRN N / 2 recognize the first floating-point register FP PR0. Similarly, both PRN entries for PRN N / 2-1 and PRN N-1 recognize the last floating-point register FP PR N / 2-1.

[0137] This arrangement allows the system to quickly and efficiently check whether paired PRN entries have already been assigned to floating-point registers. If so, the system can reassign the same PRN to a different register type, but not to another floating-point register.

[0138] Figure 12 This diagram illustrates how the system can maintain a shared free list between integer registers and floating-point registers. As shown, the system includes an integer register file 1210 and a floating-point register file 1220. The system also includes a shared free list 1230. The system further has an integer AMT 1240 for integer registers and a floating-point AMT 1250 for floating-point registers. Each AMT stores a mapping between logical register names and physical register names for its corresponding physical register file.

[0139] Unlike the shared free list described above, the shared free list 1230 includes additional bits for each PRN entry to track whether the PRN entry has been allocated to a floating-point register. Therefore, each PRN entry includes a valid bit, a commit bit, and a floating-point bit (“FP bit”).

[0140] Table 2 summarizes the status of PRN entries that still have floating-point bits.

[0141]

[0142] Table 2

[0143] As shown in Table 2, the FP bit of a PRN entry is used to track whether the PRN entry itself, or an associated PRN entry mapped to the same floating-point PRN, has already been allocated to a floating-point register. If the PRN entry itself has already been allocated to a floating-point register, then the PRN cannot be allocated to any other register until the PRN returns to the free list. However, for a given PRN entry, if its associated PRN entry has already been allocated to a floating-point register, that particular PRN entry can still be used to allocate the PRN to a non-floating-point register.

[0144] exist Figure 12 In Operation 1, the free list generates a PRN for the physical register and sets the valid bit of the corresponding PRN entry. If the physical register is a floating-point register, the system sets the FP bit of the PRN entry as well as the FP bits of all other paired or associated PRN entries. This makes the paired or associated PRN entries with the FP bit set only available for allocation to non-floating-point registers until the PRN returns to the free list, such as during a bypass return, at a refresh, or by being evicted from the FP AMT1250.

[0145] Then, the PRN-LRN mapping is entered into the ROB 1260. To track which AMT the PRN is to be retired to, the ROB 1260 can also maintain an integer field 1211 and a floating-point field 1212 to indicate whether the PRN is allocated to a floating-point register or a non-floating-point register. During operation 2, the PRN is retired to its corresponding AMT 1240 or 1250, which can be based on the ROB 1260's integer field 1211 and floating-point field 1212. The corresponding PRN commit bit is then set to indicate that these PRNs have been retired to their AMT, as indicated by the label "Set PRN commit = 1".

[0146] During refresh recovery at Operation 3, the system can use a checkpoint in the ROB to reset the head pointer to the refresh point, and then refresh the affected PRN entries by resetting the valid bits of the affected PRN entries located between the old head pointer and the refresh point, unless these affected PRN entries were stored in the AMT based on their commit bits. Therefore, if the commit bit of a refreshed PRN entry is 0, the system clears both the valid bits and floating-point bits.

[0147] However, the system can perform some additional maintenance to account for virtual PRNs, which allows the recovery of the FP bit in the PRN entries. That is, the system can perform a bitwise OR operation across all associated PRN entries, and if any associated PRN entry has its FP bit set, the system sets the FP bit in all associated PRN entries. Conversely, if none of the associated PRN entries have their FP bit set, the system can refuse to recover any FP bit.

[0148] This additional step considers a scenario where a PRN entry is allocated to a physical floating-point register and its paired or associated PRN entries are flushed. In this case, the system can restore the FP bit of the flushed PRN entry so that the flushed PRN entry is not reassigned to another floating-point register. Therefore, the system can perform a bitwise OR operation between the FP bits of all associated PRN entries and store the resulting value of the FP bits in all associated PRN entries.

[0149] In this example, the bitwise OR operation process includes first identifying paired PRN entries. When the PRN entries are as follows... Figure 11 When paired as illustrated, the paired PRN entries can be obtained using integer addition of N / 2, where N is the size of the free list (1221). The system then (e.g., using decoding modules 1222 and 1223) decodes the values ​​of the FP bits of the paired PRN entries, performs a bitwise OR operation between the FP bits (1224), and writes the result to the two PRN entries in the free list 1230.

[0150] A final step can be performed to consider the opposite situation where paired PRN entries have not yet been allocated (as evidenced by their valid bits). In this case, the system clears the FP bits of both PRN entries. See below for reference. Figure 13 This describes in more detail the maintenance of the FP bit during refresh.

[0151] During operation 4, when a floating-point PRN is evicted from the FP AMT 1250, the system clears the FP bits of that floating-point PRN entry and all other associated PRN entries. To do this, the system may first identify paired PRN entries, which in this example may include performing an integer addition of N / 2 (1231). The system may then clear their corresponding FP bits (1232 and 1233). Following the usual procedure of returning a PRN to the free list, the system may also clear the valid and committed bits of the PRN entry.

[0152] During operation 5 (i.e., the bypass return of the FP PRN without reaching the FP AMT 1250), the system can clear the valid and committed bits of the PRN entry. And because the FP PRN is ready to be reassigned to another floating-point register, the system can also clear all FP bits of all paired or associated PRN entries.

[0153] Figure 13 This is a flowchart of an example procedure for maintaining virtual PRN entries during refresh recovery when managing multiple different register sets. This procedure can be executed by a processor configured according to this specification. For convenience, the procedure will be described as being executed by the system.

[0154] As described above, in order to perform refresh recovery, the system resets the head pointer to the refresh point. All PRN entries after the refresh point and before the old head pointer position will be adjusted accordingly. Figure 13 The system processes these operations. Therefore, the system can perform this example procedure for all refreshed PRN entries.

[0155] The system determines whether the commit bit is set (1310). If the commit bit is not set, the PRN entry's PRN can be returned to the free list. Therefore, the system clears the valid and floating bits of the PRN entry (branch to 1320). If the commit bit is set, the system jumps to the next step (branch to 1330).

[0156] The system performs a bitwise OR operation between associated PRN entries and writes the result to the FP bit (1330) of each associated PRN entry. In other words, if any pair or associated PRN entry has its FP bit set, then all pair or associated PRN entries will have their FP bits set.

[0157] The system determines whether any associated PRN entry in the associated PRN entries has a valid bit set (1340). In this case, if a PRN entry in a pair or group of associated PRN entries has a valid bit set, then the PRN entry has been allocated to an unrefreshed register and will remain unchanged (branch to end).

[0158] However, if all paired or associated PRN entries are not set to valid bits, the FP bits also need to be cleared so that the PRN entries can be used in the floating-point register in the future. Therefore, if all paired or associated PRN entries are not set to valid bits, the system will clear all floating-point bits of the PRN entries in that group (branch to 1350).

[0159] Figure 14This is a flowchart of an example procedure for allocating a virtual PRN to a free list that manages multiple different sets of registers. This procedure can be executed by a processor configured according to this specification. For convenience, the procedure will be described as being executed by the system.

[0160] The system maintains a shared free list for various types of register sets (1410). As described above, the system can allocate virtual PRNs by maintaining a virtual PRN namespace in which multiple PRN entries with different virtual PRNs can be mapped to the same physical register name. Therefore, the system can group multiple PRN entries with the same PRN number mapped to a specific type of physical register together.

[0161] The system assigns a first PRN (1420) to a logical register name of type 2 and designates one or more other PRN entries as unavailable for allocation to logical registers of type 2 (1430). As described above, the system may maintain a separate bit value in each PRN entry to indicate whether any associated registers in the associated registers within the group have already been allocated to a register of a specific type. For example, the system may maintain floating-point bits to indicate whether any paired or associated PRN entries have been allocated to floating-point registers. If so, other PRN entries in the group are unavailable for allocation to floating-point registers, but these other PRN entries can still be allocated to other types of non-floating-point registers.

[0162] Although the above description uses floating-point registers as an example of a register type with a virtual PRN for clarity, the same technique can be used for any other suitable register type.

[0163] This specification also describes another way in which a shared free list can support multiple different types of register files. In the floating-point register example above, because there are fewer floating-point registers than integer registers, the free list ensures that the correct number of floating-point PRNs are allocated only according to the floating-point register file. In this example, for a free list of size N, N / 2 floating-point PRNs are allocated for N / 2 physical floating-point registers.

[0164] However, for some applications that heavily utilize floating-point execution, the renaming process itself can become a performance bottleneck. Therefore, the system described above can be configured or reconfigured, for example, during execution or manufacturing, to allow concurrent allocation of multiple associated or paired PRN entries to different logical floating-point registers. In other words, a group of PRN entries with different virtual PRNs mapped to the same floating-point PRN can be allocated to logical register names within the same time period. In contrast, in the example above, paired or associated PRN entries mapped to the same floating-point PRN as the already allocated PRN can only be allocated to registers of different types, such as integer registers or flag registers.

[0165] To prevent conflicts between multiple instructions assigned to the same floating-point PRN, the system manages renaming and execution at different stages. In the first stage, the system assigns a floating-point PRN to an instruction but does not yet allow that instruction to be executed. In the second stage, the system issues an instruction for execution only when a PRN entry with a different virtual PRN mapped to the same floating-point PRN is available for renaming (e.g., because the PRN entry was never assigned, or because the PRN entry has been returned to the free list). Returning the same floating-point PRN to the free list means that the floating-point PRN has been evicted from the AMT or otherwise returned (e.g., through the bypass return described above), and means that the program no longer needs the data in the physical register. The effect of this is to postpone the resolution of name conflicts from the renaming stage to the execution stage, where actual data dependency conflicts must be resolved anyway.

[0166] To support the assignment of multiple pairs or associated PRN entries to floating-point registers, the system can use modified logic to determine when to set and clear the FP bits of PRN entries in the free list during allocation and deallocation of floating-point registers; and to use modified logic to determine when to allow instruction issuance when an instruction is written to a floating-point register.

[0167] To maintain information about which instructions using floating-point registers are ready to be issued, the system can maintain a sliding window of PRN entries in the free list. The sliding window separates which PRNs are ready to be issued for execution from which PRNs must wait for the floating-point PRNs to return to the free list before being issued. The sliding window is a data item maintained separately from the head and tail pointers of the free list itself.

[0168] This specification also describes techniques and circuit structures for handling stuck floating-point PRNs that might otherwise never be issued for execution within the sliding window. Many situations exist in which a PRN can get stuck in the AMT. As mentioned above, a PRN can remain in the AMT until a logic writer evicts it. However, in many programs, there may be instructions that write to a logic register once, and then either write to that register much later in the program or never write to it again. In this situation, the floating-point PRN may be stuck in the AMT indefinitely, preventing its paired or associated PRN entries from being issued for execution.

[0169] Figure 15 A sliding window of PRNs in the free list is illustrated. In this example, the free list is maintained in storage units, each with 16 PRN entries. Therefore, the first storage unit 1510 contains PRN entries 0 through 15. The total free list size is 512 entries, so the last entry 1520 contains PRN entries 495 through 511.

[0170] In this example, the number of floating-point registers is half the number of integer registers. Therefore, the sliding window 1545, indicated by the dashed line, has 256 entries, which is half the number of integer registers. Using this arrangement, the system can assign two PRN entries to two different instructions, even when the PRN entries are mapped to the same floating-point PRN. In other words, different virtual PRNs for the PRN entries are mapped to the same floating-point PRN.

[0171] To implement the sliding window, the system maintains a floating-point tail pointer (FP tail pointer), which points to the oldest PRN entry that has been allocated to floating-point registers in the main floating-point register file. The FP tail pointer differs from the free list tail pointer mentioned above; for clarity, it can be referred to as the FL tail pointer to distinguish it from the FP tail pointer that defines the sliding window. The system considers the next M PRN entries starting from the FP tail pointer as being within the sliding window, where M is based on the size of the physical floating-point register file. Therefore, the system can use the FP tail pointer to determine which PRN entries are designated as ready for execution by entering the sliding window. Instructions using PRN entries within the sliding window can be safely executed because, since the sliding window is based on the size of the physical floating-point register file, it is guaranteed that instructions using PRN entries within the sliding window have corresponding physical floating-point registers at execution time.

[0172] Therefore, in order to move the sliding window, the next PRN entry or the next group of PRN entries (e.g., bank or row) must have its FP bit cleared.

[0173] References above Figures 11 to 14 In the previous example, the FP bit was set when a pair or associated PRN entry was assigned to the floating-point register. However, in this example, this requirement is modified, and the setting and clearing of each FP bit depends solely on the state of the PRN entry itself, not on the associated or paired PRN entries. Instead, the system uses a sliding window instead of the FP bit to avoid name conflicts.

[0174] Therefore, when a PRN entry becomes free for renaming (e.g., by returning to the free list after being evicted from the FP AMT, by being part of a refresh, or by entering the structured floating-point register file (described in more detail below)), the system can clear the FP bit of the PRN entry. The system can check the valid bits of the PRN entry to determine whether the PRN entry is free for renaming.

[0175] The FP tail pointer can be stepped by PRN entries or by each memory bank of a PRN entry. When the FP tail pointer is stepped by PRN entries, it can be advanced when the FP bit of the oldest allocated PRN entry is cleared. For example, the system can continuously check the valid bit of the PRN entry indicated by the FP tail pointer, and if the valid bit indicates that the PRN entry is available for renaming, the system can clear the FP bit of the PRN entry and advance the FP tail pointer. When the FP tail pointer is stepped by memory banks of PRN entries, it can advance the FP tail pointer when all FP bits in the oldest memory bank have been cleared.

[0176] When the FP tail pointer is updated, thereby moving the sliding window, the system can issue instructions sequentially or concurrently assigned to PRN entries that have already entered the sliding window.

[0177] Figure 16 This example illustrates the data structure used to handle stuck floating-point PRNs. A PRN stuck in the AMT can delay the execution of floating-point instructions because the FP bit of the PRN entry will never be cleared. Furthermore, when the FP bit is never cleared, the FP tail pointer cannot be advanced, and new floating-point instructions will not enter the sliding window.

[0178] Therefore, to address this issue, the system can use a new auxiliary physical register file that stores the structured floating-point values ​​of the physical registers that the PRN still contains in the AMT. The Structured Floating-Point Physical Register File (APRF) 1620 is a separate physical register file from the Main Floating-Point Physical Register File (MPRF) 1610.

[0179] The APRF 1620 requires only as many physical registers as the processor's logical registers, while the MPRF 1610 typically has far more. For example, if the processor has 32 logical registers, the APRF only needs 32 physical registers. Meanwhile, the MPRF typically has hundreds or more physical registers, such as 256 or 512 registers.

[0180] The APRF 1620 allows the processor to clear the FP bit of a floating-point PRN entry while still retaining the possibility of reading the value of a stuck PRN in the future. The processor cannot simply remove a stuck PRN from the AMT because it cannot know whether the logic register will be read at some future time.

[0181] APRF 1620 essentially provides another way to clear the FP bit of a PRN entry. Therefore, the ways to clear the FP bit of a PRN entry when using APRF include: the PRN being evicted from the FP AMT, returning to the free list by bypassing or refreshing, and the PRN entering the APRF.

[0182] Each entry in APRF 1620 has a corresponding PRN in virtual PRN table 1640. The size of the entries in virtual PRN table 1640 only requires the number of bits needed to store the PRN, and therefore, for the same number of entries, virtual PRN table 1640 is typically smaller than APRF 1640.

[0183] Consumer 1650 (e.g., an execution unit that executes instructions to read floating-point registers) that stores data in physical floating-point registers can first check virtual PRN table 1640. If a hit is found, consumer 1650 will read data from APRF 1620 instead of MPRF 1610. Otherwise, the consumer will read from MPRF 1610.

[0184] Table 3 summarizes the status of PRN entries when using a sliding window.

[0185]

[0186] Table 3

[0187] Table 4 illustrates an example of a stuck PRN using a series of instructions. In this example, the free list has 10 entries, and the floating-point physical register file has 5 registers. Therefore, the sliding window size is five entries. In this example, the FP tail pointer points to the PRN entry corresponding to PRN 3. Table 3 indicates the virtual PRN and the corresponding floating-point PRN assigned to each of the ten floating-point instructions, and the state of the FP bit of each PRN entry when the FP tail pointer of the sliding window points to PRN entry 3.

[0188]

[0189] Table 4

[0190] In this example, if floating-point PRN3 is stuck in the AMT, instruction 8 will never be issued because it will never enter the sliding window. In other words, the FP bit of the PRN entry used by instruction 3 will not be cleared until PRN3 returns to the free list, and PRN3 returning to the free list will trigger an update of the FP tail pointer.

[0191] However, by storing the value of PRN3 in the APRF, the system can clear the FP bit of PRN entry 3 for instruction 3. This will trigger an update of the FP tail pointer, instruction 8 will enter the sliding window, and the processor can then issue instruction 8 for execution. Subsequent consumers of the logical register to which PRN 3 is allocated will read data from the APRF instead of from the main physical register file.

[0192] Figure 17 This is a flowchart of an example procedure for managing a virtual PRN for a free list that manages multiple different sets of registers. This procedure can be executed by a processor configured according to this specification. For convenience, the procedure will be described as being executed by the system.

[0193] The system maintains a shared free list for multiple different sets of registers (1710). As mentioned above, the free list allows multiple different PRN entries to be assigned for floating-point register names by assigning virtual PRNs. The number of virtual PRNs exceeds the actual number of physical floating-point registers.

[0194] The system allocates a first virtual PRN from the first PRN entry and designates the first PRN entry as ready for execution (1720). The system can then execute the first instruction.

[0195] The system allocates a second virtual PRN (1730) mapped to the same first floating-point PRN from the second PRN entry. The system can still assign a virtual PRN to the second PRN entry, but can wait until the first PRN returns to the free list before issuing a command. In some implementations, the system uses a sliding window and waits until the second PRN entry reaches the sliding window (at which point the first PRN will have already returned to the free list) before issuing a command.

[0196] The system determines that the first PRN has been returned to the free list (1740), and the system designates the second instruction as ready for execution (1750). For example, the system can use a sliding window as a way to determine when a previously allocated PRN has been returned to the free list. When the second PRN entry enters the sliding window, the system can issue the corresponding instruction.

[0197] Issuing an instruction with a renamed register may include sending the instruction to a reserved station. The reserved station issuing the floating-point instruction may continuously check the sliding window (as indicated by the FP tail pointer) to determine whether to issue the instruction. Once the corresponding PRN entry enters the sliding window, the reserved station can execute the instruction. In some implementations, the system includes logic for modifications to the reserved station handling the floating-point instruction. Reserved stations handling other instructions not written to floating-point registers may issue their instructions regardless of the state of the sliding window.

[0198] This specification also describes how a system with a hybrid free list can use data from the free list to clock-gated regions of the physical register file. Clock gating is a power-saving technique that reduces power usage to specific regions of an electronic device.

[0199] Figure 18 Clock gating based on a free list is illustrated. One advantage of the hybrid free list described in this specification is that it implicitly includes information about which regions of the physical register file are likely to be active and which are likely to be inactive. Specifically, the region between the head and tail pointers is inherently a highly active region of the corresponding physical register file. The system can use this information to intelligently control which regions of the physical register file can be clock-gated to save power.

[0200] Figure 18 An example of a mixed free list 1830 with 512 entries is shown, which are divided into memory banks each with 16 entries. Free list 1830 allocates PRNs for physical register file 1810, which has lower memory bank 1811 and upper memory bank 1812.

[0201] The free list 1830 has a head pointer 1815 and a tail pointer 1805. Physical registers with names represented by PRN entries falling between the head pointer 1815 and the tail pointer 1805 are physical registers that are highly likely to be read from and written to during program execution. Therefore, the system can designate the region of the physical register file 1810 corresponding to the PRN entries located between the head and tail pointers as the active region 1840.

[0202] If other areas of the physical register file meet certain criteria, those areas can be designated as "low activity" or "inactive." If so, the system can clock-gated portions of the register file or subsets thereof. For example, register file 1810 is divided into two banks, 1811 and 1812. The latter part of the lower bank 1811 is a low-activity region, so the system can clock-gated the registers belonging to that portion of the lower bank 1811.

[0203] Alternatively or additionally, the system can also clock-gated the entire memory bank of physical register file 1810. For example, if the registers of upper memory bank 1812 are in low-activity areas, for example, because these registers are outside the head and tail pointers, the system can clock-gated the entire upper memory bank 1812.

[0204] The system can use any appropriate criterion to determine whether a physical register is in a low-activity or inactive state from the free list 1830. For example, a row of PRN entries in the free list 1830 where all valid bits are cleared corresponds to a group of physical registers not used by the program. Therefore, the system can clock-gated that region, including reducing the power, clock frequency, voltage, or some combination thereof to that region.

[0205] In some situations, a PRN entry may exist outside the head and tail pointers, but despite this, the data in that PRN entry may still be read during program execution. For example, if the PRN has been schemated into an AMT, the value of that physical register may still be read during program execution. However, due to the schemated state of the PRN, it is no longer possible to write to the PRN.

[0206] Therefore, the system can also designate regions of the physical register file that are outside the head and tail pointers but still have one or more PRN entries with valid bits set as active regions. Alternatively or additionally, the system can perform less aggressive clock gating on these regions, for example, by setting them to a low-power mode where they can still be read from but not written to.

[0207] Figure 19 This is a flowchart illustrating an example of clock gating of a physical register file based on a free list. This process can be performed by a processor configured according to this specification. For convenience, the process will be described as being performed by the system.

[0208] The system maintains a hybrid free list (1910) for a physical register file that is partitioned into multiple regions. The physical register file can be partitioned into regions that can be clock-gated individually, such as banks, rows, or any other suitable partitions.

[0209] This system determines that the region in the mixed free list represents the inactive region of the physical register file (1920). In some implementations, the system searches only the inactive regions outside the head and tail of the free list (e.g., before the head pointer and after the tail pointer). As mentioned above, the region in the physical register file corresponding to the PRN entry located between the head and tail pointers is likely to be the region where active reads and writes occur during program execution.

[0210] This system can identify active regions from PRN entries by examining their valid bits. PRN entries with their valid bits cleared essentially represent unused physical registers. Therefore, if the system can identify a contiguous region of PRN entries with all their valid bits cleared, it can identify the corresponding region of the physical register file as an inactive region. A key benefit of the hybrid free list described in this specification is that it tends to group large areas of inactive PRN entries, which can then be used to identify inactive regions of the physical register file.

[0211] For PRN entries that are outside the head and tail pointers but have their valid bits set, the system can designate these regions as still active because the corresponding physical registers may still be read during program execution.

[0212] The system reduces the power to inactive regions of the physical register file (1930). For example, the system can reduce the clock frequency, voltage, or completely shut down inactive regions of the physical register file.

[0213] The system can re-evaluate the active and inactive regions of the physical register file at any appropriate time interval or as the head and tail pointers move. For example, if the head pointer moves to the next line in the free list corresponding to an inactive region of the physical register file, the system can stop clock-gating that region and restore the previous power or clock level, as the physical registers in that region will soon be allocated for program execution. Similarly, when the tail pointer of the free list moves to the next line, the system can examine the valid bits of the PRN entry in the previous line to determine if all valid bits have been cleared. If so, the system can clock-gate a new region of the physical register file.

[0214] These techniques allow the system to activate specific regions of the physical register file before execution requires registers, reducing or eliminating any latency associated with clock gating of the physical register file. For example, the system can use a free list to reactivate inactive regions of the physical register file at instruction issuance time or register renaming time. In either case, the corresponding region of the physical register file will be started and running at the time of instruction issuance.

[0215] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in one or more of these. Embodiments of the subject matter described in this specification may be implemented as one or more executable programs, i.e., one or more modules of instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of the data processing device. The storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, the instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals) generated to encode information for transmission to a suitable receiver device for execution by the data processing device.

[0216] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. The device may also be or further include special-purpose logic circuit systems, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, or combinations thereof.

[0217] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages ​​or declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not necessarily, correspond to a file in a file system. A program may be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code sections). A computer program can be deployed to execute on a single computer or on multiple computers located at a site or distributed across multiple sites and interconnected via a data communication network.

[0218] The processes and logic flows described in this specification can also be executed by a dedicated logic circuit system, such as an FPGA or ASIC, or by a combination of a dedicated logic circuit system and one or more programming computers.

[0219] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices); magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0220] In addition to the embodiments described above, the following embodiments are also innovative: Example 1 is a method comprising: Maintaining a hybrid free list of physical registers that can be used for register renaming, wherein the hybrid free list includes multiple ordered entries, each entry including a pair of bits for each of the multiple physical registers. Each entry includes a valid bit and a committed bit that respectively represent the valid and committed states of the physical register. The valid state indicates whether the physical register is free to be allocated for a logical register name, and The commit status indicates whether the instruction using the physical register has been committed; and Use the head and tail pointers of the free list to perform the ordered allocation of free physical register names from the multiple entries in the free list.

[0221] Example 2 is the method as described in Example 1, further comprising: Use the head pointer to find entries that are free.

[0222] Example 3 is the method as described in Example 2, wherein the entries of the hybrid free list are arranged in multiple storage units, and

[0223] The search for entries with an idle state includes: Read the memory bank corresponding to the head pointer; and Search for entries with a free state in this storage.

[0224] Example 4 is the method as described in Example 3, further comprising: The output temporary buffer is filled based on the valid status of the entries in the memory; and Use the entries in the output temporary buffer to assign physical register names.

[0225] Example 5 is the method as described in Example 4, wherein filling the output temporary buffer includes: The complete physical register name is generated from the input temporary buffer that stores the valid bits of the entries from the memory bank.

[0226] Example 6 is the method as described in Example 5, wherein generating the complete physical register name includes: expanding each bit in the input temporary buffer into a multi-bit physical register name in the output temporary buffer.

[0227] Example 7 is the method as described in any one of Examples 1 to 6, further comprising: In response to the assignment of a physical register name, the validity status of the entry corresponding to that physical register name is modified.

[0228] Example 8 is the method as described in any one of Examples 1 to 7, further comprising: When an instruction that uses a physical register has been committed, modify the commit status of the entry corresponding to that physical register.

[0229] Example 9 is the method as described in any one of Examples 1 to 8, further comprising: When subsequent instructions are written to a logical register name, the validity and commit status of the entry corresponding to the physical register name assigned to the same logical register name are modified.

[0230] Example 10 is the method as described in any one of Examples 1 to 9, further comprising: Receive instructions to refresh at a specific refresh point; Modify the valid status of any entries in the mixed free list that are located between the refresh point and the head pointer and do not have a committed status.

[0231] Example 11 is the method as described in Example 10, further comprising: updating the head pointer of the mixed free list based on the refresh point.

[0232] Example 12 is the method as described in Example 10, wherein the refresh point is based on the oldest allocated PRN affected by the refresh.

[0233] Example 13 is the method as described in Example 12, further comprising: The checkpoints for each PRN-to-LRN mapping are stored in the reorder buffer (ROB); and Use the checkpoint in the ROB as the refresh point.

[0234] Example 14 is a method for implementing a shared free list for multiple different sets of registers, the method comprising: A shared free list is maintained for multiple different register sets, each of which includes one or more registers of a specific type, and the free list has multiple Physical Register Name (PRN) entries; For a first instruction that references a first logical register of a first type, a first physical register name is generated from a first free list entry for the first physical register of that first type; For a second instruction that references a second logical register of the second type, a second physical register name is generated from a second free list entry for that second type of second physical register; After the first instruction has been executed, the name of the first physical register is retired to the first architectural mapping table (AMT) for the register of the first type. The second physical register name is retired to the second AMT used for the second type of register; and If the first physical register name is not in the second AMT, the first physical register name is designated as idle.

[0235] Example 15 is the method as described in Example 14, further comprising: if the first physical register name is not in the first AMT, then designating the second physical register name as free.

[0236] Example 16 is the method as described in Example 15, wherein designating the second physical register name as free includes: It was determined that the PRN entry for the second physical register name in the free list was not set to the commit bit; and In response, the valid bits of the PRN entry are cleared.

[0237] Example 17 is the method as described in any one of Examples 14 to 16, wherein designating the first physical register name as free includes: It is determined that the second AMT does not maintain the name of the second physical register; and In response, the valid bits of the PRN entry for that first physical register name are cleared.

[0238] Example 18 is the method as described in any one of Examples 14 to 17, further comprising: For a third instruction that references both the first logical register of the first type and the second logical register of the second type, a shared physical register name is generated from a third mixed free list entry for the third physical register of the first type and the fourth physical register of the second type.

[0239] Example 19 is the method as described in Example 18, further comprising: returning the shared physical register name to the free list only if and only if the shared physical register name has been evicted from both the first AMT and the second AMT.

[0240] Example 20 is a method as described in any one of Examples 14 to 19, wherein the first type is an integer register.

[0241] Example 21 is a method as described in any one of Examples 14 to 20, wherein the second type is a flag register.

[0242] Example 22 is a method comprising: A shared free list is maintained for multiple different register sets, each of which includes one or more registers of a specific type, and the free list has multiple Physical Register Name (PRN) entries; For a first instruction that references a first logical register of a first type, a first physical register name is assigned from a first free list entry for the first physical register of that first type; Determines that a second instruction referencing a second logical register of a different second type satisfies one or more borrowing criteria with the first instruction; and In response, the name of the first physical register is reassigned to the second logical register referenced by the second instruction.

[0243] Example 23 is the method as described in Example 22, wherein reassigning the first physical register name to the second logical register includes: assigning the first physical register name to the second logical register when the first physical register name has already been assigned to the first logical register.

[0244] Example 24 is a method as described in any one of Examples 22 to 23, wherein the first physical register name is reassigned such that the first physical register name is simultaneously assigned to different types of logical registers.

[0245] Example 25 is a method as described in any one of Examples 22 to 24, wherein determining that the second instruction satisfies the one or more borrowing criteria includes: determining that the first instruction and the second instruction are consecutive instructions.

[0246] Example 26 is a method as described in any one of Examples 22 to 25, wherein determining that the second instruction satisfies one or more borrowing criteria includes: determining that there is no intermediate instruction between the first instruction and the second instruction that requires register renaming.

[0247] Example 27 is the method as described in Example 26, wherein determining that the second instruction satisfies one or more borrowing criteria includes: determining that the intermediate instruction does not require a logic register write.

[0248] Example 28 is a method as described in any one of Examples 26 to 27, wherein the intermediate instruction is a storage instruction.

[0249] Example 29 is a method as described in any one of Examples 22 to 28, wherein determining that the second instruction satisfies the one or more borrowing criteria includes: determining that the first instruction and the second instruction were renamed in the same cycle.

[0250] Example 30 is a method as described in any one of Examples 22 to 29, wherein determining that the second instruction satisfies one or more borrowing criteria includes: determining that an intermediate instruction between the first instruction and the second instruction is not a branch instruction.

[0251] Example 31 is the method as described in any one of Examples 22 to 30, further comprising: submitting the first instruction and the second instruction at different intervals.

[0252] Example 32 is a method as described in any one of Examples 22 to 31, wherein the first type or the second type is an integer register.

[0253] Example 33 is a method as described in any one of Examples 22 to 32, wherein the first type or the second type is a flag register.

[0254] Example 34 is a method comprising: A shared free list is maintained for multiple distinct register sets, including a first register set and different second register sets, wherein each register set includes one or more registers of a corresponding type, wherein the second register set has fewer physical registers than the first register set, and wherein the free list has multiple physical register name (PRN) entries, wherein the free list associates groups of one or more PRN entries used to identify a single register in the second register set, wherein at least one group with multiple PRN entries identifies a single physical register in the second register set; Assign a first physical register name from a first PRN entry used for the first logical register name of the second type, wherein the first PRN entry belongs to a group of one or more other associated PRN entries that also identify the same physical register name of the second register set; and Designate the one or more other PRN entries in the group as unavailable for allocation of the second type of logical register until the first PRN entry becomes available again.

[0255] Example 35 is the method as described in Example 34, further comprising: The second PRN entry belonging to the group assigns a second physical register name to the second logical register name of the first type.

[0256] Example 36 is the method as described in any one of Examples 34 to 35, wherein multiple PRN entries in the group are assigned to different types of physical registers.

[0257] Example 37 is a method as described in any one of Examples 34 to 36, wherein each PRN entry includes a second register type bit to indicate whether any PRN entry in the group of associated PRN entries has been assigned to the logical register of the second type.

[0258] Example 38 is the method as described in Example 37, wherein designating the one or more other PRN entries in the group as unavailable includes setting the first PRN entry and each of the corresponding second register type bits of the one or more other PRN entries in the group.

[0259] Example 39 is the method as described in any one of Examples 34 to 38, wherein these groups of PRN entries are pairs of PRN entries spaced N / 2 apart in the case of a free list with N PRN entries.

[0260] Example 40 is a method as described in any one of Examples 34 to 39, wherein the second register set has half the number of physical registers as the first register set has.

[0261] Example 41 is a method as described in any one of Examples 34 to 40, wherein the first set of registers includes integer physical registers and the second set of registers includes floating-point registers.

[0262] Example 42 is the method as described in Example 41, wherein the shared free list is configured to maintain PRN entries for the integer register, flag register, and floating-point register.

[0263] Example 43 is a method comprising: A shared free list is maintained for multiple distinct register sets, including a first register set and different second register sets, wherein each register set includes one or more registers of a corresponding type, wherein the second register set has fewer physical registers than the first register set, and wherein the free list has multiple physical register name (PRN) entries, wherein the free list maintains multiple PRN entries that identify individual registers in the second register set; For the first logical register name of the second type in the first instruction, allocate a first virtual PRN from the first PRN entry; For the second logical register name of the second type in the second instruction, a second virtual PRN is assigned from the second PRN entry to the first physical register name; The first instruction is executed using the name of the first physical register; and The second instruction is issued only after the first virtual PRN returns to the free list to be used for execution using the first physical register name.

[0264] Example 44 is the method as described in Example 43, further comprising: determining that the first virtual PRN has been returned to the free list includes: determining that the valid bits in the first PRN entry have been cleared.

[0265] Example 45 is the method as described in any one of Examples 43 to 44, further comprising: Maintain a sliding window for the second type of PRN entry that is ready to be issued for execution. Issuing the second instruction includes issuing the second instruction only after the second PRN entry enters the sliding window.

[0266] Example 46 is the method as described in Example 45, further comprising: Maintain a floating-point tail pointer (FP tail pointer) that indicates the oldest allocated PRN entry of the second type to a physical register in the set of main physical registers of the second type, and wherein the sliding window is limited by a predetermined number of PRN entries following the FP tail pointer.

[0267] Example 47 is the method as described in Example 46, further comprising: Determine that the associated PRN entry of the PRN entry indicated by the FP tail pointer is free for renaming; and In response, update the FP tail pointer.

[0268] Example 48 is the method as described in Example 47, further comprising: Clear the floating-point bits of the PRN entry indicated by the FP tail pointer.

[0269] Example 49 is the method as described in Example 45, wherein the size of the sliding window is based on the number of registers in the second register set.

[0270] Example 50 is the method as described in Example 45, further comprising: Determine in the Architectural Mapping Table (AMT) the PRN card for which the third logical register name is assigned from the third PRN entry to the third physical register of the second type; In response, data is moved from the third physical register to the structured floating-point physical register file; and Clear the FP bit of the fourth PRN entry mapped to the same PRN in the AMT.

[0271] Example 51 is the method as described in Example 50, wherein determining the PRN card in the AMT includes: It is determined that the FP tail pointer points to the third PRN entry, which has not yet been evicted from the AMT.

[0272] Example 52 is the method as described in Example 50, further comprising: Execute instructions that read from the name of the third logical register; and Read values ​​from the structured floating-point physical register file instead of the main floating-point physical register file.

[0273] Example 53 is the method as described in any one of Examples 43 to 52, further comprising: Issue an instruction with a renamed register to the reserved station for execution; Continuously check whether the PRN entry used for this instruction is within the sliding window; and The instruction is executed after the PRN entry enters the sliding window.

[0274] Example 54 is a method as described in any one of Examples 43 to 53, wherein the first set of registers includes integer physical registers and the second set of registers includes floating-point registers.

[0275] Example 55 is the method described in Example 54, wherein the shared free list is configured to maintain PRN entries for the integer register, flag register, and floating-point register.

[0276] Example 56 is a method comprising: The processor maintains a free list that tracks the availability of physical registers in the physical register file, where the free list is partitioned into multiple regions, each corresponding to a subset of the physical register file; The region in the free list represents the inactive region of the physical register file; and In response, clock gating is applied to one or more physical registers in the physical register file that correspond to the inactive region of the free list.

[0277] Example 57 is the method as described in Example 56, wherein each free list entry has a valid bit indicating whether the corresponding physical register is being used; and

[0278] Determining that a region of the free list is an inactive region includes: determining that all valid bits of the free list entries in that region indicate that the corresponding physical register is not in use.

[0279] Example 58 is the method as described in any one of Examples 56 to 58, further comprising: Determine the region of the free list between the head and tail pointers of the free list; and In response, designate that area as the active area.

[0280] Example 59 is a method as described in any one of Examples 56 to 58, wherein the physical register file includes a plurality of memory banks, and wherein determining that the region of the free list is an inactive region includes: determining that all physical registers in the memory bank corresponding to the region of the free list are unused.

[0281] Example 60 is the method described in Example 59, wherein clock gating of the one or more physical registers includes: clock gating of the memory corresponding to the region of the free list, while maintaining normal clock operation for one or more other memory banks of the physical register file.

[0282] Example 61 is a method as described in any one of Examples 56 to 60, wherein clock gating of the one or more physical registers occurs before the one or more physical registers have been assigned to an instruction.

[0283] Example 62 is a method as described in any one of Examples 56 to 61, further comprising: activating the one or more physical registers during the instruction issuance phase of using the physical registers in the one or more physical registers.

[0284] Example 63 is a method as described in any one of Examples 56 to 62, further comprising: resuming normal clock operation of the physical register during a register renaming process using the free list.

[0285] Example 64 is a processor including a logic circuit system configured to perform the method as described in any one of Examples 1 to 63.

[0286] Example 64 is a data processing device configured to perform the method as described in any one of Examples 1 to 63.

[0287] While this specification contains numerous details of specific implementations, these details should not be construed as limiting the scope of any invention or the scope of any claims, but rather as descriptions of features that may be characteristic of particular embodiments of a particular invention. Certain features described in this specification within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may involve sub-combinations or variations thereof.

[0288] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all illustrated operations to be performed to achieve the desired result. In some contexts, multitasking and parallel processing can be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0289] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions recited in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A method comprising: The processor maintains a free list that tracks the availability of physical registers in the physical register file, wherein the free list is partitioned into multiple regions, each corresponding to a subset of the physical register file; The region defined in the free list represents the inactive region of the physical register file; as well as In response, the power of one or more physical registers in the physical register file corresponding to the inactive region of the free list is reduced.

2. The method of claim 1, wherein each free list entry in the free list has a valid bit indicating whether the corresponding physical register is being used; and Determining that the region in the free list is an inactive region includes: It is determined that all valid bits of the free list entry in the region indicate that the corresponding physical register is not in use.

3. The method according to any one of claims 1 to 2, further comprising: The region of the free list is determined to be between the head pointer and the tail pointer of the free list; as well as In response, the area is designated as the active area.

4. The method of any one of claims 1 to 3, wherein the physical register file comprises a plurality of storage banks, and wherein determining that the region of the free list is an inactive region comprises: It is determined that all physical registers in the memory corresponding to the region in the free list are unused.

5. The method of claim 4, wherein reducing the power to the one or more physical registers comprises: Clock gating is performed on the memory bank corresponding to the region of the free list, while normal clock operation is maintained for one or more other memory banks of the physical register file.

6. The method of any one of claims 1 to 5, wherein the power reduction to the one or more physical registers occurs before the one or more physical registers have been assigned to an instruction.

7. The method of any one of claims 1 to 6, further comprising: The one or more physical registers are activated during the instruction issuance phase of the one or more physical registers.

8. The method of any one of claims 1 to 7, further comprising: During the register renaming process using the free list, the normal clock operation of the physical registers is resumed.

9. A processor including a logic circuit system configured to perform operations including: The processor maintains a free list that tracks the availability of physical registers in the physical register file, wherein the free list is partitioned into multiple regions, each corresponding to a subset of the physical register file; The region defined in the free list represents the inactive region of the physical register file; as well as In response, the power of one or more physical registers in the physical register file corresponding to the inactive region of the free list is reduced.

10. The processor of claim 9, wherein each free list entry in the free list has a valid bit indicating whether the corresponding physical register is being used; and Determining that the region in the free list is an inactive region includes: It is determined that all valid bits of the free list entry in the region indicate that the corresponding physical register is not in use.

11. The processor of any one of claims 9 to 10, wherein the operation further comprises: The region of the free list is determined to be between the head pointer and the tail pointer of the free list; as well as In response, the area is designated as the active area.

12. The processor of any one of claims 10 to 11, wherein the physical register file comprises a plurality of memory banks, and wherein determining that the region of the free list is an inactive region comprises: It is determined that all physical registers in the memory corresponding to the region in the free list are unused.

13. The processor of claim 12, wherein the power reduction to the one or more physical registers comprises: Clock gating is performed on the memory bank corresponding to the region of the free list, while normal clock operation is maintained for one or more other memory banks of the physical register file.

14. The processor of any one of claims 10 to 13, wherein the power reduction to the one or more physical registers occurs before the one or more physical registers have been assigned to instructions.

15. The processor of any one of claims 10 to 14, wherein the operation further comprises: The one or more physical registers are activated during the instruction issuance phase of the one or more physical registers.

16. The processor of any one of claims 10 to 15, wherein the operation further comprises: During the register renaming process using the free list, the normal clock operation of the physical registers is resumed.