Power reduction using a prelist for register renaming

KR1020260139786APending Publication Date: 2026-09-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
KR1020267027660
Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-09-22

Smart Images

  • Figure P1020267027660_ABST
    Figure P1020267027660_ABST
Patent Text Reader

Abstract

A method, system, and apparatus are provided for reducing power to physical registers using a freelist used for register renaming. One method comprises the step of maintaining a freelist in a processor that tracks the availability of physical registers in a physical register file, wherein the freelist is partitioned into multiple regions, each region corresponding to a subset of the physical register file. If a region of the freelist represents an inactive region of the physical register file, power is reduced for one or more physical registers in the physical register file corresponding to the inactive region of the freelist.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Modern out-of-order (OOO) processing utilizes register renaming to eliminate false data dependencies, allowing more instructions to be executed in parallel; this often entails executing them out of order. In this specification, register renaming refers to assigning a physical register name (PRN) to a logical register name (LRN) within an instruction currently being executed. When register renaming is used, subsequent instructions that write to a logical register may be executed before or in parallel with other instructions that use or read from the same logical register name, because the register renaming process has assigned different underlying physical registers to the logical register name.

[0002] A "freelist" is a microarchitecture component of a processor that maintains a pool of available physical resources, e.g., physical registers. In this specification, available resources, e.g., physical registers, will be described as being in a free state or the corresponding PRN being a free PRN. Thus, the freelist may provide free PRNs to allocate physical registers for register renaming. When a PRN is no longer needed, it may be returned to the freelist for reuse.

[0003] Freelist size and control logic complexity are major considerations in processor design that affect silicon area, power, and performance in machines using OOO processing.

[0004] Traditionally, there have been two main methods for implementing free lists in OOO processors: using a First-In, First-Out (FIFO) queue and using bit vectors. In a FIFO free list, PRNs are allocated one by one starting from the beginning of the free list and returned to the end when they are no longer needed. Therefore, in a pure FIFO free list, PRNs are allocated based on age, and older PRNs that have not been allocated for a longer time are allocated before younger PRNs that have been allocated more recently. However, since the entire PRN, which can be up to 10 bits or longer, is stored in the FIFO, returning a PRN to the FIFO free list requires writing multi-bit PRN entries. Consequently, frequently writing multi-bit PRN entries in a FIFO free list entails large area and power requirements, which raises serious issues regarding performance, cost, and scalability.

[0005] In contrast, a bit vector freelist uses a single bit to represent each PRN and therefore has smaller storage requirements than a FIFO freelist. However, bit vector freelists require very complex control logic for routine tasks. For example, flush recovery, such as performing it after a mis-predicted branch, requires substantially more complex logic in a bit vector freelist than in a FIFO freelist because the bit vector freelist does not contain age information for each PRN. Furthermore, as the total number of PRNs increases, the complexity of finding available PRNs also increases, which further increases area and power requirements. For example, when providing multiple available entries from a bit vector freelist, finding the first available entries, or the first pair of available PRNs, becomes more difficult as the frequency of requests and vector lengths increase, which is a significant scalability bottleneck.

[0006] This specification describes techniques related to a hybrid freelist for fast and efficient register renaming that provide the benefits of both FIFO and bit vector freelists, as well as many other additional benefits. A hybrid freelist maintains a set of PRN entries that require only a few bits of storage. As an example, each PRN entry may have 1) a valid bit indicating whether the corresponding PRN is in a free state or is allocated, and 2) a commit bit indicating whether the PRN is in an architectured state, which means that the data in the corresponding physical register is part of the machine's permanent state at a given time and its corresponding instruction is no longer subject to being flushed or re-executed.

[0007] In this specification, a PRN entry is a structure that stores data for which a PRN can be generated based on the location of the PRN entry within a freelist. In this specification, when a freelist is described as generating a PRN, it means generating a PRN using the location of the PRN entry, and does not mean that the freelist stores the entire PRN itself.

[0008] The hybrid freelist maintains a head pointer indicating the location within the freelist to search for the next free PRN. In some implementations, the head pointer is advanced by freelist segments of multiple PRN entries rather than by individual PRN entries. The hybrid freelist also maintains a tail pointer indicating the entry corresponding to the most recent PRN that is free to be reallocated. The PRNs before the head pointer and after the tail pointer can be allocated in order, skipping over PRNs that have already been allocated and have not yet been returned.

[0009] The tail pointer is generally used to protect the sequential array of allocated PRN entries. Therefore, the head pointer is not allowed to overtake the tail pointer when allocating a PRN entry. By maintaining the head and tail pointers, the processor can essentially perform circular allocation of PRNs by cycling through the freelist in order while skipping over already allocated PRNs.

[0010] The head pointer can be used for efficient flush recovery. When recovering PRNs during flush recovery, the head pointer is simply redirected to the flush point, and the valid bits for PRNs between the flush point and the previous head pointer are reset, except for PRNs in an architectured state as indicated by their commit bits.

[0011] Specific embodiments of the subject of the invention described herein may be implemented to realize one or more of the following advantages. A hybrid freelist achieves the advantages of both a FIFO freelist and a bit vector freelist while mitigating their individual disadvantages. For example, a hybrid freelist can efficiently clear allocated PRNs during flush recovery. Additionally, a hybrid freelist contains simplified PRN information that requires less storage space. Thus, a hybrid freelist achieves area and power reductions that would otherwise not be achievable in an equivalent FIFO freelist. Furthermore, a hybrid freelist is fully scalable in that flush recovery does not become disproportionately more difficult as the number of PRNs increases, as is the case with a strict bit vector freelist. The freelist described herein can also efficiently allocate groups of PRNs to meet the bandwidth requirements of a hyperscalar processor. For example, the freelist can continuously fill input staging buffers and output staging buffers, so that N available and all PRNs are already queued at every cycle.

[0012] The hybrid freelist described herein may also be used for register renaming of multiple different types of physical registers. This allows the processor to maintain a single freelist instead of multiple freelists, which saves power and reduces the required silicon area. For example, the freelist may be used for register renaming of both integer registers and flag registers. The freelist may also be used for register renaming of integer registers and floating-point registers, even if the sizes of the individual physical register files differ. Furthermore, even when the freelist is used for register renaming of multiple different types of physical registers, the underlying scalable flush recovery mechanism remains largely unchanged and can be used to return all types of physical register names to the freelist after a flush.

[0013] This specification also describes a method in which a freelist can be used to borrow already allocated physical register names. This arrangement increases processor efficiency and mitigates the possibility of the processor running out of register space, all without adding any additional naming storage.

[0014] This specification also describes a method by which a processor can improve floating-point instruction bandwidth by introducing a new microarchitecture structure and an architectured floating-point physical register file, which is an auxiliary physical register file for maintaining the values ​​of floating-point registers whose names are stuck in an architectured map table.

[0015] This specification also describes techniques for using a sequential structure of a hybrid freelist to reduce processor power through clock gating. Since the freelist inherently maintains information about which physical registers are in use during execution, the processor can utilize this information to perform highly targeted clock gating of unused physical registers, thereby reducing power without substantially affecting processor performance.

[0016] Details regarding one or more embodiments of the subject matter of the invention described herein are provided in the accompanying drawings and the following description. Other features and advantages of the subject matter of the invention described herein will become apparent from the description, drawings, and claims. Brief explanation of the drawing

[0017] Figure 1a is an overview of the exemplary lifecycle of a PRN within a hybrid freelist. Figure 1b is an exemplary overview of a hybrid freelist. Figure 2 is an exemplary process for allocating PRNs from a hybrid freelist. Figure 3 is an exemplary process for deallocating PRNs into a hybrid freelist. Figure 4 illustrates an exemplary flush recovery using a hybrid freelist. Figure 5 is an exemplary process for implementing a hybrid freelist. Figure 6 is a diagram of a processor system having a shared freelist using a single namespace of physical register names for two different physical register files. Figure 7 is an example of how a system can maintain a single shared free list for integer registers and flag registers. Figure 8 is a flowchart of an exemplary process for maintaining a shared freelist. Figure 9 illustrates PRN borrowing in a system having two sets of physical registers. Figure 10 is a flowchart of an exemplary process for borrowing physical register names. FIG. 11 illustrates an exemplary technique for associating PRN entries for floating-point registers. Figure 12 is a diagram of how a system can maintain a shared freelist between integer registers and floating-point registers. FIG. 13 is a flowchart of an exemplary process for maintaining a virtual PRN entry during flush recovery when managing multiple different sets of registers. FIG. 14 is an exemplary process flowchart for allocating virtual PRNs to a freelist that manages multiple different sets of registers. Figure 15 illustrates a sliding window for PRNs in a freelist. Figure 16 illustrates a data structure for handling fixed floating-point PRNs. FIG. 17 is an exemplary flowchart of a process for managing virtual PRNs for a freelist that manages multiple different sets of registers. FIG. 18 illustrates a hybrid freelist having 512 entries divided into banks of 16 entries each. Figure 19 is a flowchart of an example for clock gating a physical register file based on a freelist. Similar reference numbers and names in various drawings indicate similar elements. Specific details for implementing the invention

[0018] FIG. 1a is an overview of an exemplary lifecycle (100) of a PRN maintained by a hybrid freelist. The lifecycle (100) includes three phases maintained by three corresponding structures: a hybrid freelist phase (102), a reorder buffer (ROB) phase (130), and an architectured map table (AMT) phase (140). Throughout OOO processing, the PRN cycles between these three phases.

[0019] PRNs start in a free state in the freelist phase (102). Then, free PRNs can be allocated from the hybrid freelist, renamed, and transitioned to the ROB phase (130), which stores the mapping between the PRNs and their corresponding LRNs in instruction order. When a PRN exits the freelist phase (120) and transitions to the ROB phase (130) or the AMT phase (140), as a general rule, the PRN is no longer in a free state and therefore cannot be allocated to another logical register until the PRN is returned to the freelist phase (120). However, in certain situations described in detail below, the PRN may be reassigned to another logical register before being returned to the freelist.

[0020] ROB can also associate each PRN-LRN mapping with a corresponding position within the freelist, so that when a flush occurs, the head pointer of the freelist can be efficiently snapped back to the flush point.

[0021] When the instruction corresponding to an entry in the ROB completes execution and reaches the ROB head, the instruction is committed by retiring the PRN to the AMT, which also saves the mapping between the PRN and its corresponding LRN. The data stored by the physical register corresponding to the PRN is now part of the machine's permanent state and is not subject to being flushed or re-executed.

[0022] When a subsequent instruction is written to the logical register of the PRN maintained within the AMT, the PRN mapping is no longer needed, the PRN is evicted from the AMT and cycled back into the hybrid freelist, re-entering the hybrid freelist phase (102).

[0023] A hybrid freelist may maintain PRN entries, each having a pair of values, for example, a pair of bits, to track the active and committed states of each PRN. The "active" bit can be used to track the active state of a PRN, indicating whether the PRN is in a free state for renaming. Thus, the active state of a PRN can be referred to as being in a free state or an allocated state.

[0024] The "Commit" bit can be used to track the commit state of a PRN, indicating whether the instructions using the assigned PRN have been committed. This means that the value in the corresponding physical register has become the permanent state of the program being executed by the processor and is no longer subject to being flushed or re-executed. Therefore, the commit state of a PRN can be referred to as being in a committed state or an uncommitted state. A change in the commit state from an uncommitted state to a committed state generally corresponds to the PRN leaving the ROB and entering the AMT.

[0025] Accordingly, a PRN or its corresponding PRN entry may be referred to herein as having a free state or an assigned state capable of indicating the value of the effective bit of the PRN entry. Similarly, a PRN or its corresponding PRN entry may be referred to as having a committed state or an uncommitted state capable of indicating the value of the commit bit of the PRN entry.

[0026] In some cases, the valid and commit bits have binary values, where "1" represents a set value and "0" represents a clear value. A set value for the valid bit, e.g., "1", may indicate that the PRN has been allocated for renaming, whereas a set value for the commit bit may indicate that the PRN has been retired to the AMT. In some cases, a cleared valid bit value, e.g., "0", and a set commit bit value, e.g., "1", indicate an illegal or improper state because they imply that the PRN was committed before being allocated. Although such a situation is not expected to occur, if it does, the processor may enter a fault state to restore the integrity of the freelist. A table presenting exemplary combinations of valid and commit bit values ​​is shown in FIG. 1a. In an alternative implementation, the values ​​of what are considered to be set and clear, or high and low, may be inverted.

[0027] In some implementations, the availability of PRNs may be dynamically configured rather than explicitly expressed. For example, the free state for each PRN may be determined using a freelist head pointer and a tail pointer instead of maintaining an explicit valid bit for each PRN. In this case, the head pointer indicates the area of ​​the freelist to search for PRNs to allocate. Additionally, the commit state of a PRN may be configured by checking the AMT rather than maintaining a commit bit. However, when the processor has hundreds of physical registers, simply maintaining an explicit commit bit in the freelist may often be faster than searching the AMT.

[0028] FIG. 1b is an exemplary overview (101) of a hybrid freelist (132). The hybrid freelist (132) includes a set of PRN entries representing PRNs. Each PRN entry has a valid bit (104) and a commit bit (106). Additionally, the hybrid freelist (132) includes a head pointer (108) and a tail pointer (110).

[0029] When PRNs need to be allocated in OOO processing, the hybrid freelist (132) will use a head pointer (108) to find PRN entries for the PRNs that can be allocated. Although FIG. 1b illustrates a single PRN entry at the head pointer (108) within the hybrid freelist (132), the head pointer within the hybrid freelist can be implemented as a coarse-grained pointer that steps per segment of multiple PRN entries. For example, FIG. 1b also illustrates a method in which a set of PRN entries within the freelist (132) can be implemented as banks (116), and the head pointer steps per bank. In some implementations, the size of each bank is based on the processor's bandwidth. For example, if the processor needs to rename 10 registers per cycle, the freelist may have a bank size of 10 PRN entries, or more than 10 PRN entries to account for PRNs that are still in a committed state. Typically, each instruction writes to only one logical register, and thus, the processor's instruction bandwidth typically corresponds to the PRNs that need to be generated in each cycle.

[0030] The tail pointer (110) is updated as PRNs retire from the ROB. The tail pointer (110) is updated when the next PRN entry retires from the ROB by having a state of {0,0} or {1,1} for its valid and commit bits. This means that the PRN retired from the ROB without entering the AMT, or that an instruction using the PRN was committed and the PRN entered the AMT. Alternatively or additionally, when the tail pointer is stepped by a bank of PRN entries, the tail pointer is updated when all entries in the next bank have a state of {0,0} or {1,1}. Thus, the tail pointer (110) indicates the most recent PRN entry or bank of PRN entries that can be reassigned.

[0031] In contrast, the head pointer indicates the oldest PRN that can be allocated, or a bank of PRN entries, which corresponds to the PRN that has not been used for the longest number of cycles.

[0032] As mentioned above, the system maintains the tail pointer (110) to preserve the sequential arrangement of PRN entries in the free list by preventing the head pointer (108) from overtaking the tail pointer (110). This is important for maintaining or deriving age information regarding PRN entries, because when the architectured PRNs returned to the free list are reallocated, it may be impossible or impractical to determine the age of the corresponding PRN entries on the flush. Maintaining the tail pointer (110) also improves power performance by preventing continuous updates of the head pointer and staging buffers across many PRN entries that are not free to be allocated.

[0033] When PRN entries are read from the hybrid freelist (102), a decoding process is performed to generate a corresponding PRN from the location of the PRN entry in the freelist. This process saves storage space because it allows the freelist to efficiently store 2 bits per PRN instead of all bits of each PRN.

[0034] The process of stepping through PRN banks using pointers is highly scalable. In other words, head (108) and tail (110) pointers can be used to allocate any number of PRNs from the hybrid freelist without disproportionately increasing processing requirements.

[0035] In some implementations, the freelist continuously fills one or more staging buffers so that available PRNs are always ready in each cycle. The input staging buffer receives PRN entries from the bank indicated by the head pointer. Then, the output staging buffer is filled with all physical register numbers corresponding to the PRN entries in the input staging buffer. In either the input staging buffer or the output staging buffer, already allocated PRN entries may be removed, for example, as indicated by their effective bits. In some cases, the input staging buffer contains only the effective bit (104) values ​​of each PRN entry identified from the hybrid freelist (102). In other words, the input staging buffer may essentially be a bit vector representing which PRNs are free in the bank indicated by the head pointer. The position of the set bit in the input staging buffer and the position of the head pointer are used to fill the output staging buffer with all PRNs. For example, the output staging buffer may contain a bank of 16 different PRNs. In other cases, the buffers may be extended to include different numbers of PRNs and / or bits. The operation of the staging buffers will be described in more detail below with reference to FIG. 2.

[0036] FIG. 2 is a diagram of an exemplary process (200) for allocating PRNs from a hybrid freelist (202). The exemplary process, as well as other processes described herein, may be implemented using a digital logic circuit of a processor, for example, a non-sequential processor.

[0037] In this example, for the sake of simplification of the example, each PRN entry in the freelist (202) is represented as having a single valid bit, but as described above, hybrid freelist entries will typically also have a commit bit. The processor includes an input staging buffer (204) and an output staging buffer (206). In some cases, when the PRN bank identified by the head pointer is read, the PRNs in the bank that can be assigned (e.g., valid bits (104) are cleared, e.g., "0", or an alternative availability state) are extended from a series of input staging (204) and output staging (206) buffers to the PRNs. This is achieved by placing the PRN entry from the bank indicated by the head pointer into the input staging buffer (204). For example, as illustrated in FIG. 2, the input staging buffer (204) may contain 16 bits, where each bit corresponds to the valid bit value of a different PRN entry from the bank indicated by the head pointer (108).

[0038] Next, physical register names corresponding to the PRN entries in the input staging buffer (204) are filled into the output staging buffer (206). The physical register names can be generated by the positions of the PRN entries in the free list. In some embodiments, the positions can be reconstructed by adding the offset of the PRN entry in the input or output staging buffer or in the bank of PRN entries to the position of the head pointer. Assuming a PRN length of 10 bits, the output staging buffer (204) may contain 10-bit PRNs for all free PRN entries in the input staging buffer (204), thus converting the PRN entries into actual PRNs. For example, from 16 bits of the input staging buffer (204), the system may generate up to 16 complete 10-bit PRNs in the output staging buffer (206).

[0039] As part of this process, "holes" in the input staging buffer caused by PRNs that are unavailable according to their valid bits are essentially squeezed out, and accordingly, the output staging buffer stores only the PRNs of PRN entries that have cleared valid bits.

[0040] Then, the processor can use the available PRNs in the output staging buffer (204) for renaming by, for example, assigning the PRNs to logical register names and storing each assigned PRN to LRN mapping in the ROB (302).

[0041] FIG. 3 is an exemplary process (300) for deallocating PRNs into a hybrid freelist (102) after they have been retired from the ROB and / or evicted from the AMT (304). As described above, after a PRN is allocated, the value of the associated PRN entry valid bit (104) can be set to "1".

[0042] During OOO processing, the PRN mapping between PRN and LRN will remain in ROB (302) until the associated instruction is committed and the PRN to LRN mapping is retired or flushed to AMT. When the PRN is retired, the value of the commit bit (106) of the associated PRN entry is set to “1” (operation 306), which indicates that the instruction of the PRN has been committed and the PRN has been cycled to AMT (304) (operation 308). As discussed above, AMT (304) maintains the PRN to LRN mapping for the committed PRNs, and any committed PRNs entering AMT (304) for the same logical register will evict any previous PRN mapping for that logical register. Therefore, typically the size of the AMT (304) is based on the number of logical registers in the instruction set, whereas the size of the ROB (302) is generally much larger because there are more physical registers than logical register names in the instruction set. When the PRN is evicted from the AMT (304), the values ​​of the associated valid bit (104) and commit bit (106) are cleared, for example, back to "0" to indicate that the PRN is now available for allocation (operation 310).

[0043] FIG. 4 illustrates an exemplary flush recovery using a hybrid freelist (102). In some cases, PRNs need to be returned from ROB (302) to the hybrid freelist (102) or deallocated. This process is referred to as flush recovery. Flush recovery may be performed when instructions are speculatively executed along an incorrect control flow path, for example, an incorrectly predicted branch.

[0044] During flush recovery, the flush point is designated at a point within the hybrid freelist (102), and the head pointer (108) will be steered to this new point. In some embodiments, the flush point is the point where the writer for a specific PRN and all more recent PRNs are invalidated from the ROB. Any PRN after the flush point and before the original head pointer (108) (e.g., the head pointer (108) prior to steering) will have its valid bit (104) cleared, for example, to “0”, except for PRNs in the AMT (304) that are still in use, as indicated by their commit bits being set.

[0045] In some cases, to support efficient flush recovery, PRN pointers may be continuously checkedpointed during the dispatch process (Operation 402). For example, in the dispatch process, the current head pointer (108) plus an offset value may be associated with each PRN entering the ROB (302). Alternatively, or additionally, PRN pointers may be stored in other structures linked to the ROB or structures tracking in-flight instructions, such as a reservation station or an issue queue, these are just a few examples. In other words, the system may use any appropriate structure to collect PRNs that have been allocated but not yet committed within the AMT, and the corresponding PRN entries will be returned to the freelist on the flush.

[0046] When flush recovery begins, the current checkpoint is indicated by the oldest PRN to be flushed from the ROB (302). The head pointer of the freelist can then be reset to the flush point by using the pointer of the corresponding ROB entry (Operation 404). If the head pointer points to banks of PRN entries, an offset can be used to determine which of the PRN entries in the bank are part of the flush. In some implementations, the input and output staging buffers are also cleared so that they can be refilled with PRN data from the new flush point. In some cases, the flush recovery procedures are limited to an allocated time to prevent collisions with the ROB (302).

[0047] By maintaining checkpoints, flush recovery can be performed efficiently without the need to thoroughly search for PRNs affected during the flush. The flush recovery process in the hybrid freelist (102) is also highly scalable because the performance of flush recovery is minimally affected by an increase in the freelist size. In other words, even if the freelist increases tenfold, the flush recovery process will still be that fast. This is very different from the bit vector freelist, where flush recovery becomes more computationally intensive as the number of PRNs increases and becomes more complex and time-consuming as the ROB size increases. Using checkpoints within the ROB also avoids side-tracking in-flight PRNs by the ROB or other structures to determine and represent the set of PRNs affected by the flush (as in the case of the bit vector freelist).

[0048] FIG. 5 is an exemplary process (500) for implementing a hybrid freelist. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0049] The system uses the free list head pointer to find PRNs that are in a free state (e.g., a valid bit with a value of "0") (510). As described above, PRNs in the hybrid free list are associated with both a valid bit and a commit bit, which indicate which phase the PRN is in. The valid bit indicates whether the PRN is currently in use.

[0050] The hybrid freelist also maintains head and tail pointers. The head pointer indicates the next PRN entry or set of entries to be allocated, while the tail pointer indicates the next PRN to be retired, which is being written to by the oldest in-flight instruction writing to the register. In some implementations, the head pointer points to a bank of PRN entries, and the head pointer steps through each bank of PRN entries.

[0051] The system allocates a PRN and modifies the free status of the PRN (520). Then, the PRN is allocated, and the valid bit is updated to a valid bit value, for example, “1,” to indicate that the PRN is no longer in a free state. As described above, when a PRN is allocated from the hybrid freelist, it is cycled to the ROB, which maintains the sequential list of instructions to which the PRNs are allocated, as well as the PRN to LRN mappings for those instructions.

[0052] The system retires the PRN and modifies the commit state (530). At commit time, the PRN is retired and the PRN to LRN mapping is saved in the AMT. When this happens, the commit bit of the PRN entry is modified to indicate that the PRN has retired (e.g., by setting the value of the commit bit to "1"). As described above, when the PRN is switched to the AMT, it can evict the older PRN assigned to the same LRN for the older instruction.

[0053] The system evicts PRN from AMT and modifies the free state of PRN (540). When PRN is evicted from AMT and returned to the hybrid freelist, both the valid bit and the commit bit are updated to indicate that PRN is in a free state, for example, both values ​​of the valid bit and the commit bit can be set to “0”.

[0054] This specification also describes how freelists are used for multiple different types of physical registers, each having its own physical register namespace. For example, some processors have different physical register files, each with its own namespace. Instructions may reference different logical register names corresponding to different underlying physical register files. While it is possible to maintain a separate freelist for each individual physical register file, in many situations this would be inefficient and result in redundant hardware.

[0055] Instead, the techniques described herein enable the processor to use a single shared freelist to maintain and allocate physical register names for multiple different register files. When executing an instruction requires a physical register name for register renaming, the freelist may allocate the physical register name using a single shared namespace applicable to both types of physical register files. Additionally, for some instructions that reference multiple different types of registers, the freelist may allocate the same physical register name for both types of logical registers within the instruction.

[0056] One major advantage of using a shared freelist for multiple different types of physical register files is that only one flush recovery mechanism needs to be executed on a flush. Additionally, sharing a freelist between different types of registers can save hardware resources without substantially affecting performance.

[0057] FIG. 6 is a diagram of a processor system (600) having a shared freelist using a single namespace of physical register names for two different physical register files. The system (600) includes a first physical register file (610) having physical registers (PR1, PR2, to PRN). The system (600) also includes a second register file (620) having physical registers (FR1, FR2, to FRN).

[0058] The system (600) also includes a shared freelist (630) that maintains physical register name entries (PRN1, PRN2 to PRNN). The PRN entries for this single shared namespace can be used for register renaming for both the first physical register file (610) and the second physical register file (620).

[0059] Typically, each physical register file has its own built-in map table (AMT). Thus, the physical register file (610) will have its own AMT, and the second physical register file (620) will have its own AMT. Since the two physical register files share the same namespace for register renaming, the system can ensure that a physical register name is not returned to the freelist (630) unless it is expelled from both AMTs or is absent.

[0060] One example of a distinct type of physical register file is a special register file. In some processors, special physical registers can be used to maintain information regarding the processor's state. These special registers are often maintained in a physical register file separate from the physical register file for integers or floating-point numbers.

[0061] An example of a special register is a flag register, which is sometimes referred to as an application program status register. In this specification, a flag register will refer to any special register that has a physical register file distinct from an integer register file or a floating-point register file and can be referenced by a logical register name within an instruction set.

[0062] In some processors, there is only a single logical flag register to which instructions can write, but there are tens or hundreds of physical flag registers that can be assigned to that single logical flag register. This is because, in OOO execution, flag registers can be used to maintain the conditional state of multiple different in-flight instruction branches. Therefore, even if an instruction set has only a single logical flag register, register renaming can be used to maintain the state of all in-flight conditional branches.

[0063] FIG. 7 is an example of a method in which a system (700) can maintain a single shared freelist (730) for integer registers and flag registers. The system (700) includes a set of integer registers (710) and a set of flag registers (720). The system also includes a hybrid freelist (730) that maintains a single physical register namespace for both the integer registers (710) and the flag registers (720). The hybrid freelist (730) includes PRN entries that include a valid bit and a commit bit, respectively. The integer register file (710) has its own integer AMT (740). The flag register file (720) also has its own flag AMT (750).

[0064] As described above, the valid bit indicates whether a physical register name is free for register renaming or if it has already been allocated. The commit bit indicates whether the physical register name is currently still in the AMT. As described above, these two bits govern how the system allocates physical register names and how the system handles returning physical register names to the freelist on a flush, which essentially involves returning PRNs that were allocated after the flush point but are not yet in the AMT.

[0065] In this exemplary system, the integer registers and flag registers share the freelist (730) by using the valid bits explicitly expressed in the PRN entries of the freelist (730).

[0066] The system (700) may use the integer commit bits (706) of the PRN entries to track the commit status of the PRNs assigned to the integer registers. However, for the flag PRNs, the system (700) may reconstruct the virtual flag commit bit from the PRN entries and from the state of the flag AMT (750). This relieves the system of having to maintain a separate free list for the flag registers and keeps the size of the PRN entries 2 bits each.

[0067] In this example, the flag AMT (750) has only one entry because the flag registers in this example have only one logical register within the instruction set. Therefore, the physical-to-logical register name mapping within the flag AMT (750) requires only a single entry. Furthermore, when only a single logical flag register exists, the flag AMT (750) needs to store only the physical register name, rather than the entire mapping from the physical register name to the logical register name. Thus, the flag AMT (750) is exemplified by storing only a single PRN instead of the mapping. In contrast, the integer AMT (740) has as many integer entries as there are integer logical registers, and the integer AMT (740) stores the mappings between the physical register name and the logical register name for each entry.

[0068] At allocation time, physical register names may be assigned according to their valid state. For integer registers and flag registers, this means checking the valid bit in the freelist (730). However, to prevent naming conflicts, the system also performs a bitwise OR between the valid state, as indicated by the PRN entry, and the decoded commit bit from the flag AMT (750). This prevents the reassignment of the same PRN until the PRN is evicted from both AMTs (740 and 750) or otherwise absent.

[0069] The decode module (770) can be used to generate a decoded commit bit for any PRN entry in the freelist (730) by checking whether the corresponding PRN is still in the flag AMT (750). In other words, for a specific PRN, the decode module (770) can generate a decoded commit bit that is 1 if the PRN is in the flag AMT (750), and a decoded commit bit that is 0 otherwise.

[0070] Then, from the commit bit decoded for the PRN, the system (700) can generate a virtual valid bit (772) for the PRN by using a bitwise OR between 1) the decoded commit bit and 2) the valid bit of the corresponding PRN entry in the freelist (730). In effect, this means that an integer register and a flag register can be simultaneously assigned to the same PRN when they occur in the same instruction.

[0071] Similarly, the system (700) can generate a virtual flag commit bit (774) for the PRN by using a bitwise OR between 1) the decoded commit bit and 2) the commit bit of the corresponding PRN entry in the freelist (730).

[0072] As described above, the physical register name can be freed again when the PRN is not still within the integer AMT (740) or flag AMT (750). In the case of the integer AMT (740), the system may use the commit bits of the PRN entries in the freelist (730) to determine whether the physical register name is still within the integer AMT (740).

[0073] On the other hand, since there is only one logical flag register maintained by the flag AMT (750), the system can save significant storage space by virtually determining the flag commit bit from the decoded commit bit. Thus, for flag PRNs, the system can determine the commit state by 1) performing a bitwise OR between the commit bit of the corresponding PRN entry and 2) the decoded commit bit.

[0074] Table 1 summarizes the states of PRN entries shared between integer registers and flag registers.

[0075] effectiveness Integer commit Flag commit (virtual) meaning 0 0 0 PRN is free for any register type. 1 0 0 PRN is assigned to an integer register, a flag register, or both within the ROB, and has not yet been retired to the AMT. 0 1 0 Invalid state 1 1 0 PRN was retired as an integer AMT. 0 0 1 The integer PRN was ejected from the integer AMT, and the flag PRN is still within the flag AMT. 1 0 1 PRN was retired by Flag AMT. 0 1 1 Invalid state 1 1 1 PRN was retired with both integer AMT and flag AMT.

[0076] Operations for maintaining a shared freelist (730) will now be described. During operation 1, a physical register name can be assigned by identifying a PRN entry in the freelist that has a free state. As described above, PRN entries available for renaming can be identified through the use of input and output staging buffers, which can also be used to generate a PRN from the freelist location of the corresponding PRN entry.

[0077] Because the integer register file (710) and the flag register file (720) share a namespace, when an instruction refers to a logical register for integers or a logical register for flag registers, the system simply assigns a physical register name to any free PRN entry in the free list (730).

[0078] In some implementations, the system may assign the same PRN to both the integer register and the flag register. For example, if a single instruction references both the integer logical register and the flag logical register, the system may find an available PRN from the freelist and assign the same PRN to both the integer logical register and the flag logical register. Although the same physical register name is assigned, the values ​​written to the individual logical registers will be stored by different physical registers because the logical register names reference different physical register files (710 and 720).

[0079] Then, the PRN to LRN mapping is stored in the ROB (760) while the instructions are being executed. In some implementations, the ROB (760) includes fields (711 and 712) indicating whether the mapping is an integer mapping, a flag mapping, or both. Because the system assigns the same physical register name to instructions that reference both the integer register file (710) and the flag register file (720), the ROB (760) requires only 2 bits—a flag bit (711) and an integer bit (712)—for each ROB entry to indicate whether the entry is an integer mapping, a flag mapping, or both.

[0080] Operation 2 represents a flush recovery. Since the integer register file (710) and the flag register file (720) share the freelist, only a single flush operation needs to be performed for the freelist (730). As described above, the head pointer of the freelist (730) will be snapped back to the flush point indicated by the checkpoint within the ROB (760) as described above. Any assigned PRN entries after the flush point and before the head pointer, whose commit bits are not set, will be returned to the freelist, for example, by resetting their valid and commit bits.

[0081] Operation 3 indicates a bypass return to the freelist (730). This is a special case to avoid unnecessary storage of PRNs within the AMTs. For example, if a physical register mapped to a logical register retires in the same cycle as a more recent instruction writing to the same logical register, there is no need to add the physical register to the AMT because the physical register will be overwritten by the more recent writer. Therefore, the physical register name can bypass the AMT and return directly to the freelist. By doing so, the system can clear them, for example, by setting the valid and integer commit bits to 0.

[0082] During operation 4, when PRN is retired from ROB (760), if PRN was only an integer PRN, PRN will be retired to an integer AMT (740). The system may also set this, for example, by setting the integer commit bit of the corresponding PRN entry to 1. If PRN was only a flag PRN, PRN will be retired to a flag AMT (750). If PRN was both an integer PRN and a flag PRN, PRN will be retired to both AMTs (740 and 750).

[0083] The PRNs can eventually be expelled from their individual AMTs (740 and 750).

[0084] During operation 5, the integer PRN is evicted from the integer AMT (740). The system may clear the valid bit and the integer commit bit from the corresponding PRN entry in the freelist (730) to indicate that the PRN is no longer in the integer AMT (740).

[0085] However, until the PRN is also stored within the flag AMT (750), and unless it is stored, the PRN will not actually be free for allocation. As described above, the system (700) can enforce this restriction using a virtual valid bit (772) which is a bitwise OR between the valid bit of the PRN entry and the decoded commit bit from the flag AMT (750).

[0086] When flag PRN is evicted from flag AMT (750), the system may reset the effective bit of the corresponding PRN entry. When this occurs, since a different PRN will be stored by flag AMT (750), the decoded commit bit will be changed from 1 to 0 for that PRN. Therefore, at that point, unless the integer PRN is still in integer AMT (740), the virtual flag commit bit for the PRN will be reset to 0.

[0087] FIG. 8 is a flowchart of an exemplary process for maintaining a shared freelist. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0088] The system maintains a shared freelist for multiple types of register sets (810). As described above, the processor may have different physical register files, each having its own physical register name. For multiple physical register files, the system may maintain a single physical register namespace. As an example, the system may maintain a single physical register namespace for both integer registers and flag registers maintained in separate physical register files.

[0089] The system generates a first physical register name for a first type of physical register (820) and generates a second physical register name for a second type of physical register (830). In other words, the system can generate physical register names from the same shared freelist even if the physical registers are in different physical register files. And, as described above, if an instruction references physical registers in both physical register files, the system can generate a shared physical register name for both physical registers in the different physical register files.

[0090] The system retires the first physical register name and the second physical register name to different individual AMTs (840). As described above, different register files have different AMTs to maintain the processor's state for committed instructions.

[0091] The system designates the first physical register name as free if the first physical register name is not in the second AMT (850). In other words, because different physical register files share namespaces, the physical register name should not be free if it is still maintained in one of the AMTs. Therefore, to return the first physical register name to the freelist, the system can first check whether the physical register name is still maintained in the second AMT. As described above, in some implementations, the system can perform this check as a bitwise OR between the decoded commit bit and the valid bit maintained in the PRN entry of the freelist.

[0092] This specification also describes a method in which a shared freelist can be used to borrow an already allocated physical register name. In other words, in some situations, even though a physical register name has already been allocated to a logical register of an instruction, the same physical register name may be reallocated to another logical register under certain circumstances. This arrangement increases processor efficiency and mitigates the possibility of the processor running out of register space, all of which is accomplished without adding any additional naming storage.

[0093] FIG. 9 illustrates PRN borrowing in a system having two sets of physical registers (910 and 920). In this example, a shared freelist (930) will be used to assign physical register names for a sequence of instructions (905, 915, and 925). In this example, the same physical register names are used for the logical registers referenced by the instructions (905 and 925).

[0094] The add instruction (905) adds the values ​​of the logical registers (X2 and X3) and uses the logical register (X1) to store the value. Since this situation requires the allocation of a physical register, the shared freelist (930) generates a physical register name, PRN2 in this example, which may be, for example, a binary value 0x10. Thus, the processor will use that physical register name to identify the second physical register within the integer register set (910). Consequently, the result of the add instruction is written to the second physical integer register (911).

[0095] Next, the interposed store instruction (915) is executed. Since the store instruction does not require writing arbitrary values ​​to arbitrary physical registers, the store instruction does not require the allocation of arbitrary physical registers through register renaming. For this reason, if the next instruction refers only to logical registers of a different type, the physical register name assigned to the add instruction (905) can be reassigned to the logical register of the instruction (925).

[0096] The instruction (925) compares the values ​​of the logical registers (X5 and X6) and stores the result in the flag register. The syntax of the instruction (925) refers to the flag register only implicitly because it is a comparison instruction. Since the instruction (925) needs to write to the physical flag register, the shared freelist will generate a physical register name that identifies one of the flag registers in the flag register file (920).

[0097] However, because the comparison instruction (925) satisfies the criteria for borrowing physical register names, the freelist (930) reassigns the same physical register name previously created for the instruction (905) by regenerating PRN2. As a result, the result of the comparison instruction will be stored in the second physical flag register (921).

[0098] FIG. 10 is a flowchart of an exemplary process for borrowing physical register names. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0099] The system maintains a shared freelist for multiple different sets of registers (1010). As described above, the different sets of registers have different types and are referenced by logical register names that reflect these types.

[0100] The system assigns a first physical register name to a first physical register of a first type (1020). The system can use a freelist allocation process to generate the next available PRN based on the values ​​of the PRN entries and the position of the head pointer.

[0101] The system receives a second instruction for a second type of physical register (1030), and the system determines whether the second instruction satisfies one or more borrowing criteria for the first instruction (1030). In other words, the system can check the second logical register type for one or more previously received instructions to determine whether the instructions satisfy one or more borrowing criteria.

[0102] A first example of a borrowing criterion is that the instructions are sequential. Therefore, in some implementations, the system checks the instructions against previously executed instructions to determine whether they refer to physical registers of different types. In such cases, the system may determine that two instructions satisfy the borrowing criterion.

[0103] A second example of the borrowing criterion is that one or more intervening instructions are instructions that do not require register renaming. For example, a store instruction between two other instructions is an example of an instruction that does not require register renaming. Therefore, any number of store instructions can occur between two instructions, and they can still satisfy the borrowing criterion.

[0104] A third example of a borrowing criterion is whether instructions are renamed within the same cycle. As described above, the use of input and output staging buffers allows the processor to perform multiple renamings within the same cycle. Therefore, in some implementations, the system may allow borrowing only between instructions that are renamed within the same cycle.

[0105] A fourth example of a borrowing criterion is whether there is a branch instruction between two instructions. In some implementations, to simplify the logic of flush recovery, the system may not allow borrowing if the intervening instruction is a branch instruction. However, in some alternative implementations, if the processor processes the instructions as a non-borrowing case after flush recovery is initiated, borrowing may still be allowed despite the intervening branch instruction. In other words, before flush recovery, the system may allow borrowing, but after flush, the system may not allow borrowing for the instructions.

[0106] If one or more borrowing criteria are not met, the system generates a new physical register name for the second instruction (branch to 1050).

[0107] On the other hand, if one or more borrowing criteria are met, the system reallocates the same PRN for the second instruction (branching to 1060). As described above, when the same PRN is used to execute different instructions, the results will actually go to different physical register files.

[0108] When two instructions are assigned the same PRN, the system can ensure that both are evicted from their respective AMTs before returning the PRN to the freelist. Thus, when the first PRN is evicted from the AMT by the more recent writer to the logical register, the system can check the commit status of the other PRN by checking the explicitly maintained commit bit or by retrieving that AMT. Only when both PRNs are evicted from their respective AMTs will the system return the shared PRN to the freelist.

[0109] This specification also describes how a shared freelist can be further extended to manage physical floating-point registers without a substantial increase in hardware cost. By sharing a freelist between integer registers and floating-point registers, the processor can effectively eliminate a separate freelist for the floating-point registers. Additionally, the freelist can still employ the same allocation and flush recovery processes with only an insubstantial amount of additional record keeping.

[0110] Sharing between integer registers and floating-point registers is more complex than sharing with flag registers, because floating-point registers tend to be fewer than integer or flag registers. Processor design is the result of a carefully selected number of each type of register to achieve a specific level of performance and efficiency. Therefore, while the relative number of physical registers within each register file is quite variable, floating-point registers tend to be fewer than integer registers, partly because floating-point registers are larger in size. As explained above, since flag registers store much less data than integer registers—for example, 4 bits for a flag register versus 64 bits for an integer register—an equal number of integer and flag registers can exist efficiently. In contrast, to give just a few common examples, the number of physical floating-point registers can be much smaller, for instance, only one-quarter or half the number of integer registers.

[0111] To support sharing between integer registers and floating-point registers, the system may allocate virtual PRNs for floating-point registers by nature. In this specification, a PRN being virtual means that the PRN belongs to a namespace that does not correspond to the size of the physical register file. Instead, a mapping exists between each virtual PRN and each floating-point PRN in a floating-point PRN namespace that corresponds to the size of the physical register file. When a freelist is used to maintain a virtual PRN namespace, this means that while all PRN entries within the freelist can be allocated to floating-point registers, multiple PRN entries within the freelist with different virtual PRNs can be mapped to the same floating-point PRN because there are fewer floating-point registers than freelist entries. To avoid reallocating an already allocated floating-point PRN, the system can track which PRN entries are associated with each other by having virtual PRNs mapped to the same floating-point PRN. Then, if a PRN is already assigned to a floating-point register, none of the other associated PRN entries can be assigned to another logical floating-point register name until the assigned one is deallocated. However, during this time, the PRNs of the other associated PRN entries can still be assigned to other register types, e.g., integer registers. In the following description, PRN entries described as paired or associated PRN entries refer to a group of multiple PRN entries having different virtual PRNs mapped to the same floating-point PRN.

[0112] The techniques described below for sharing between integer registers and floating-point registers can also be combined with the techniques described above for sharing with flag registers and for PRN borrowing. Thus, the shared free list described herein can be used to maintain PRN allocations for at least three different types of physical registers: integer, flag, and floating-point registers.

[0113] FIG. 11 illustrates an exemplary technique for associating PRN entries with floating-point registers. In this example, the physical floating-point register file has half the number of physical registers as the integer physical register file. Therefore, the freelist can maintain a virtual PRN namespace by associating two PRN entries with all floating-point PRNs. To do so, the system can pair PRN entries in any appropriate manner so that each pair of PRN entries identifies the same floating-point PRN. FIG. 11 illustrates an example of pairing PRN entries in a manner that does not require overly complex logic or write maintenance.

[0114] As illustrated in FIG. 11, a free list of N PRN entries is conceptually divided into two halves: a first half (1110) spanning from PRN0 to PRN N / 2 - 1 and a second half (1120) spanning from PRN N / 2 to PRN N-1. Corresponding PRN entries within each half are paired with each other so that they identify the same floating-point PRN in a physical floating-point register file (1130) having N / 2 entries each.

[0115] For example, the PRN entries for PRN0 and PRN N / 2 both identify the first floating-point register (FP PR0). Similarly, the PRN entries for PRN N / 2-1 and PRN N-1 both identify the last floating-point register (FP PR N / 2-1).

[0116] With this arrangement, the system can quickly and efficiently check whether a paired PRN entry has been assigned to a floating-point register. If so, the system can reassign the same PRN to a different register type, but not to a different floating-point register.

[0117] FIG. 12 is a diagram of a method by which a system can maintain a shared freelist between integer registers and floating-point registers. As illustrated, the system includes an integer register file (1210) and a floating-point register file (1220). The system also includes a shared freelist (1230). The system also has an integer AMT (1240) for the integer registers and a floating-point AMT (1250) for the floating-point registers. Each AMT stores a mapping between logical register names and physical register names for their respective physical register files.

[0118] Unlike the shared freelists described above, the shared freelist (1230) includes an additional bit for each PRN entry to track whether the PRN entry has been assigned to a floating-point register. Thus, each PRN entry includes a valid bit, a commit bit, and a floating-point bit ("FP bit").

[0119] Table 2 also summarizes the states of PRN entries with floating-point bits.

[0120] effectiveness commit Floating point meaning 0 0 0 PRN is free for any register type. 1 0 0 PRN is allocated to a non-floating-point register and has not yet been retired to AMT. 0 1 0 Invalid state 1 1 0 PRN was retired as the AMT of the non-floating-point register. 0 0 1 The associated PRN entry has been assigned to a floating-point register. The PRN is still free for the floating-point register. 1 0 1 PRN was assigned to a floating-point register and has not yet been retired to AMT. 0 1 1 Invalid state 1 1 1 PRN was retired as floating-point AMT.

[0121] As shown in Table 2, the FP bit of a PRN entry is used to track whether the PRN for the PRN entry, or the PRN for an associated PRN entry mapped to the same floating-point PRN, has been assigned to a floating-point register. If the PRN entry itself is assigned to a floating-point register, the PRN cannot be assigned to any other register until it is returned to the freelist. However, for a specific PRN entry, if its associated PRN entry is assigned to a floating-point register, that specific PRN entry can still be used to assign the PRN to a non-floating-point register.

[0122] In operation 1 of FIG. 12, the freelist generates a PRN for a physical register and sets the valid bit for the corresponding PRN entry. If the physical register is a floating-point register, the system sets the FP bit for the PRN entry as well as all other paired or associated PRN entries. It makes the paired or associated PRN entries with the FP bit set available for allocation to non-floating-point registers only, until the PRN is returned to the freelist, for example, during a bypass return, on a flush, or by being evicted from the FP AMT (1250).

[0123] Then, the PRN-LRN mapping enters the ROB (1260). To track which AMT the PRN will retire to, the ROB (1260) may also retain integer and floating-point fields (1211 and 1212) to indicate whether the PRN is assigned to a floating-point register or a non-floating-point register. During operation 2, the PRNs are retired to their respective AMTs (1240 or 1250), which may be based on the integer and floating-point fields (1211 and 1212) of the ROB (1260). Then, the corresponding PRN commit bits will be set to indicate that they have retired to their AMTs, as indicated by the label "Set PRN commit = 1".

[0124] During the flush recovery in operation 3, the system may use a checkpoint within the ROB to reset the head pointer to the flush point, and then, for the affected PRN entries between the previous head pointer and the flush point, flush them by resetting their valid bits unless they are stored in the AMT according to the commit bit. Thus, if the commit bit is 0 for a flushed PRN entry, the system clears both the valid and floating-point bits.

[0125] However, the system may perform some additional maintenance to account for virtual PRNs, which may result in restoring FP bits from PRN entries. That is, the system may perform a bitwise OR operation between all associated PRN entries, and if any of the associated PRN entries still have an FP bit set, the system sets the FP bits within all associated PRN entries. On the other hand, if none of the associated PRN entries have a set FP bit, the system may not restore any of the FP bits.

[0126] This additional step considers a situation where a PRN entry is assigned to a physical floating-point register and its paired or associated PRN entry is flushed. In such a situation, the system can restore the FP bits of the flushed PRN entry to prevent them from being reallocated to another floating-point register. Thus, the system can perform a bitwise OR operation between the FP bits of all associated PRN entries and store the resulting values ​​for the FP bits in all associated PRN entries.

[0127] In this example, the bitwise OR process first involves determining the paired PRN entries. When the PRN entries are paired as exemplified in FIG. 11, the paired PRN entries can be obtained using an integer addition of N / 2, where N is the size of the freelist (1221). Then, the system decodes the values ​​of the FP bits of the paired PRN entries (e.g., using decode modules (1222 and 1223)), performs a bitwise OR between the FP bits (1224), and writes the result to both PRN entries in the freelist (1230).

[0128] The final step can be performed to account for the opposite situation where the paired PRN entry is not assigned, as evidenced by its valid bit. In such a situation, the system clears the FP bits of both PRN entries. The retention of these FP bits during the flush is described in more detail below with reference to FIG. 13.

[0129] During operation 4, when the floating-point PRN is evicted from the FP AMT (1250), the system will clear the FP bits of the floating-point PRN entry and all other associated PRN entries. To do so, the system may first determine the paired PRN entries, which may include performing an integer addition of N / 2 in this example (1231). Then, the system may clear their individual FP bits (1232 and 1233). According to the usual operation of returning the PRN to the freelist, the system may also clear the valid and commit bits of the PRN entries.

[0130] During operation 5, which is a bypass return of FP PRN without reaching FP AMT (1250), the system may clear the valid and commit bits of the PRN entry. And because FP PRN is ready to be reallocated to another floating-point register, the system may also clear all FP bits for all paired or associated PRN entries.

[0131] FIG. 13 is a flowchart of an exemplary process for maintaining a virtual PRN entry during flush recovery when managing a number of different register sets. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0132] As described above, to perform flush recovery, the system resets the head pointer to the flush point. All PRN entries after the flush point and before the previous head pointer position will be processed according to the operation of FIG. 13. Thus, the system can perform an exemplary process for all flushed PRN entries.

[0133] The system determines whether the commit bit is set (1310). If the commit bit is not set, the PRN of the PRN entry can be returned to the freelist. Therefore, the system clears the valid bit and floating-point bit of the PRN entry (branch to 1320). If the commit bit is set, the system skips to the next step (branch to 1330).

[0134] The system performs bitwise OR between the associated PRN entries and writes the result to the FP bit of each associated PRN entry (1330). In other words, if any of the paired or associated PRN entries has its FP bit set, all paired or associated PRN entries will have their FP bits set.

[0135] The system determines whether any of the associated PRN entries has a set valid bit (1340). In this situation, if a PRN entry within a paired or associated group of PRN entries has a set valid bit, it is assigned to a register that is not subject to flushing and will remain as is (branch to exit).

[0136] However, if none of the paired or associated PRN entries have a set valid bit, the FP bits also need to be cleared to make the PRN entries available to floating-point registers in the future. Therefore, if none of the paired or associated PRN entries have a set valid bit, the system clears all floating-point bits for the PRN entries within the group (branch to 1350).

[0137] FIG. 14 is a flowchart of an exemplary process for allocating virtual PRNs to a freelist that manages a plurality of different register sets. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0138] The system maintains a shared freelist for multiple sets of different types of registers (1410). As described above, the system can allocate virtual PRNs by maintaining a virtual PRN namespace in which multiple PRN entries having different virtual PRNs can be mapped to the same physical register name. Thus, the system can group together multiple PRN entries that are mapped to the same PRN number for specific types of physical registers.

[0139] The system assigns a first PRN to a logical register name of type 2 (1420) and designates one or more other PRN entries as unavailable for assignment to logical registers of type 2 (1430). As described above, the system may maintain a distinct bit value in each PRN entry to indicate whether any of the associated registers within the group are assigned to a specific type of register. For example, the system may maintain a floating-point bit to indicate whether any paired or associated PRN entries are assigned to floating-point registers. In such a case, other PRN entries within the group cannot be assigned to floating-point registers, but they can still be assigned to other types of non-floating-point registers.

[0140] For clarity, the above description used floating-point registers as examples of register types with virtual PRNs, but the same techniques can also be used for any other suitable register types.

[0141] This specification also describes another way in which a shared freelist can support multiple different types of register files. In the floating-point register example described above, because there are fewer floating-point registers than integer registers, the freelist ensured that only the exact number of floating-point PRNs were allocated according to the floating-point register file. In that example, N / 2 floating-point PRNs were allocated for N / 2 physical floating-point registers for a freelist of size N.

[0142] However, for some applications that heavily utilize floating-point execution, the renaming process itself can become a performance bottleneck. Therefore, the system described above may be configured or reconfigured to allow the simultaneous assignment of multiple associated or paired PRN entries to different logical floating-point registers, for example, during execution or manufacturing. In other words, a group of PRN entries having different virtual PRNs mapped to the same floating-point PRN may be assigned to logical register names over the same time period. In contrast, in the example described above, a paired or associated PRN entry mapped to the same floating-point PRN as an already assigned PRN may be assigned only to different types of registers, for example, integer or flag registers.

[0143] To prevent conflicts between multiple instructions assigned the same floating-point PRN, the system may manage renaming and execution at different stages. In the first stage, the system assigns a floating-point PRN to the instruction but does not yet allow the instruction to be executed. In the second stage, the system issues an instruction for execution only if a PRN entry with a different virtual PRN mapped to the same floating-point PRN is available for renaming, for example, because it has never been assigned or has been returned to the freelist. Returning the same floating-point PRN to the freelist means that the floating-point PRN has been evicted from the AMT or returned in another way, for example, by the bypass return described above, and that the data in the physical register is no longer needed by the program. The effect of this is to push the resolution of name clashes from the renaming stage to the execution stage, where actual data dependency conflicts must be resolved in any case.

[0144] To support multiple paired or associated PRN entries allocated to floating-point registers, the system may use modified logic to determine when to set and clear the FP bits of PRN entries in the free list when the floating-point registers are allocated and deallocated, as well as to determine when to issue instructions when writing to the floating-point registers.

[0145] To maintain information about which instructions are ready to be issued using floating-point registers, the system may maintain a sliding window of PRN entries in the freelist. The sliding window separates which PRNs are ready to be issued for execution from which floating-point PRNs must wait to return to the freelist before being issued. The sliding window is a data entry maintained separately from the head and tail pointers of the freelist itself.

[0146] This specification also describes techniques and circuit structures for handling stuck floating-point PRNs that can never enter the sliding window for execution. There are many situations in which a PRN can become stuck within the AMT. As described above, a PRN can remain in the AMT until a logic writer evicts it. However, in many programs, there may be instructions that write to a logic register once and then write to the program much later or never write to it again. In such situations, a floating-point PRN can become stuck indefinitely within the AMT, preventing its paired or associated PRN entry from being issued for execution.

[0147] FIG. 15 illustrates a sliding window for PRNs within a freelist. In this example, the freelist is maintained as a bank of 16 PRN entries each. Thus, the first bank (1510) contains PRN entries 0 through 15. Since the total freelist size is 512 entries, the last entry (1520) contains PRN entries 495 through 511.

[0148] In this example, the number of floating-point registers is half the number of integer registers. Thus, the sliding window (1545) indicated by the dashed line has 256 entries, which is half the number of integer registers. Using this arrangement, the system can assign two PRN entries to two different instructions even if the PRN entries are mapped to the same floating-point PRN. In other words, different virtual PRNs of PRN entries are mapped to the same floating-point PRN.

[0149] To implement a sliding window, the system may maintain a floating-point tail pointer (FP tail pointer) that points to the oldest PRN entry assigned to a floating-point register in the main floating-point register file. The FP tail pointer is distinct from the freelist tail pointer described above, and for clarity, it may be referred to as the FL tail pointer to distinguish it from the FP tail pointer defining the sliding window. From the FP tail pointer, the system may consider the next M PRN entries to be within the sliding window, where M is based on the size of the physical floating-point register file. Thus, the system may use the FP tail pointer to determine which PRN entries to designate as ready for execution by entering the sliding window. Because the sliding window is based on the size of the physical floating-point register file, instructions using PRN entries within the sliding window can be safely executed, as they are guaranteed to have the corresponding physical floating-point register at the time of execution.

[0150] Therefore, to move the sliding window, the next PRN entry or next group, for example, the bank or row of PRN entries, must have its FP bits cleared.

[0151] In the previous example described above with reference to FIGS. 11 through 14, the FP bit is set when a paired or associated PRN entry is assigned to a floating-point register. However, in this example, this requirement is modified, and the setting and clearing of each FP bit depend solely on the state of the PRN entry itself, and not on the associated or paired PRN entry. Instead, the system uses a sliding window instead of the FP bit to avoid name conflicts.

[0152] Therefore, the system may clear the FP bit for a PRN entry when the PRN entry is free for renaming, for example, by being evicted from the FP AMT and returned to the freelist, by becoming part of the flush, or by entering the architectured floating-point register file described in more detail below. The system may check the valid bit of the PRN entry to determine whether it is free for renaming.

[0153] The FP tail pointer can be stepped per PRN entry or by each bank of PRN entries. When the FP tail pointer is stepped per PRN entry, the FP tail pointer can be advanced when the FP bit of the oldest allocated PRN entry is cleared. For example, the system can continuously check the valid bit of the PRN entry indicated by the FP tail pointer, and if the valid bit indicates that the PRN entry is available for renaming, the system can clear the FP bit of the PRN entry and advance the FP tail pointer. When the FP tail pointer is stepped per bank of PRN entries, the system can advance the FP tail pointer when all FP bits in the oldest bank are cleared.

[0154] When the FP tail pointer is updated to move the sliding window, the system can issue instructions assigned to the PRN entries that have entered the sliding window sequentially or simultaneously.

[0155] FIG. 16 illustrates a data structure for handling stuck floating-point PRNs. A PRN stuck in the AMT can delay the execution of floating-point instructions because the FP bit of the PRN entry will never be cleared. And when the FP bit is never cleared, the FP tail pointer cannot be advanced, and new floating-point instructions will not be able to enter the sliding window.

[0156] Therefore, to solve this problem, the system may use a new auxiliary physical register file that stores the architectured floating-point values ​​for the physical registers having PRNs that are still within the AMT. The architectured floating-point physical register file (APRF) (1620) is a physical register file separate from the main floating-point physical register file (MPRF) (1610).

[0157] APRF (1620) requires only as many physical registers as there are logical registers for the processor, whereas MPRF (1610) typically has many more physical registers. For example, if the processor has 32 logical registers, APRF only needs to have 32 physical registers. Meanwhile, MPRF typically has hundreds or more physical registers, for example, 256 or 512 registers.

[0158] APRF (1620) allows the processor to clear the FP bit for a floating-point PRN entry while still preserving the possibility that the value of the stuck PRN will be read in the future. Because the processor cannot know whether the logic register will be read at some point in the future, it cannot simply remove the stuck PRN from the AMT.

[0159] APRF (1620) essentially provides another path through which the FP bit of a PRN entry can be cleared. Thus, the ways in which the FP bit for a PRN entry can be cleared when using APRF include PRNs that enter the APRF as well as PRNs that are returned to the freelist by flushing, bypassing, or eviction from the FP AMT.

[0160] Each entry in the APRF (1620) has a corresponding PRN in the virtual PRN table (1640). The size of the entries in the virtual PRN table (1640) only needs to be able to store the bits of the PRN, and thus the virtual PRN table (1640) is typically smaller than the APRF (1640) for the same number of entries.

[0161] A consumer (1650) of data stored in a physical floating-point register, for example, an execution unit executing an instruction to read a floating-point register, may first check the virtual PRN table (1640). If a hit exists, the consumer (1650) will read data from the APRF (1620) rather than from the MPRF (1610). Otherwise, the consumer will read from the MPRF (1610).

[0162] Table 3 summarizes the states of PRN entries when using a sliding window.

[0163] effectiveness commit FP meaning 0 0 0 PRN is free for any register type. 1 0 0 PRN is allocated to a non-floating-point register and has not yet been retired to AMT. 0 1 0 Invalid state 1 1 0 PRN is designed for integer or flag registers, or for the floating-point registers of APRF. 0 0 1 Invalid state 1 0 1 PRN was assigned to a floating-point register and has not yet been architectured. 0 1 1 Invalid state 1 1 1 PRN has been retired and is being architectured in MPRF.

[0164] Table 4 illustrates an example of a stuck PRN through a sequence of instructions. In this example, the freelist has 10 entries, and the floating-point physical register file has 5 registers. Therefore, the sliding window size is 5 entries. In this example, the FP tail pointer points to the PRN entry corresponding to PRN 3. Table 3 shows the virtual PRN assigned to each of the 10 floating-point instructions and the corresponding floating-point PRN, as well as the state of the FP bit for each PRN entry when the sliding window's FP tail pointer points to PRN entry 3.

[0165] Command number Virtual PRN of the assigned PRN entry Floating-point PRN of the assigned PRN entry FP bit of the PRN entry allocated at time T12 note 1 1 1 Not applicable 2 2 2 Not applicable 3 3 3 1 The PRN entries assigned to FP tail pointer instructions 3 through 7 are in a 5-entry sliding window, and thus the instructions are ready to be issued. 4 4 4 1 5 5 5 1 6 6 1 1 7 7 2 1 8 8 3 1 When PRN3 returns to the freelist, instruction 8 (outside the sliding window) cannot enter the sliding window until the FP bit is cleared. 9 9 4 1 10 10 5 1

[0166] In this example, if the floating-point PRN3 gets stuck in the AMT, instruction 8 will never be issued because it will never enter the sliding window. In other words, the FP bit for the PRN entry used by instruction 3 will not be cleared until PRN3 returns to the freelist, which will trigger an update of the FP tail pointer.

[0167] However, by storing the value of PRN3 in the APRF, the system can clear the FP bit for PRN entry 3 for instruction 3. This will trigger an update of the FP tail pointer, instruction 8 will enter the sliding window, and the processor can then issue instruction 8 for execution. Subsequent consumers of the logical register to which PRN 3 is assigned will read data from the APRF rather than from the main physical register file.

[0168] FIG. 17 is a flowchart of an exemplary process for managing virtual PRNs for a freelist that manages a number of different register sets. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0169] The system maintains a shared freelist for multiple different sets of registers (1710). As described above, the freelist can allow multiple different PRN entries to be assigned to floating-point register names by allocating virtual PRNs. The virtual PRNs are greater than the actual number of physical floating-point registers.

[0170] The system allocates a first virtual PRN from the first PRN entry and designates the first PRN entry as ready to execute (1720). Then the system can execute the first instruction.

[0171] The system assigns a second virtual PRN from the second PRN entry that maps to the same first floating-point PRN (1730). The system may still assign a virtual PRN to the second PRN entry, but may wait to issue an instruction until the first PRN is returned to the freelist. In some implementations, the system uses a sliding window and waits to issue an instruction until the second PRN entry reaches the sliding window, at which point the first PRN would have been returned to the freelist.

[0172] The system determines that the first PRN has been returned to the freelist (1740), and the system designates the second instruction as ready for execution (1750). For example, the system may use a sliding window as a method to determine when previously allocated PRNs have been returned to the freelist. When the second PRN entry enters the sliding window, the system may issue a corresponding instruction.

[0173] Issuing an instruction with a renamed register may involve sending the instruction to a reservation station. A reservation station issuing a floating-point instruction may continuously check the sliding window, as indicated by the FP tail pointer, to determine whether to issue the instruction. The reservation station may then execute the instruction as soon as the corresponding PRN entry enters the sliding window. In some implementations, the system includes modified logic for reservation stations handling floating-point instructions. Reservation stations handling other instructions that do not write to floating-point registers may issue their instructions regardless of the state of the sliding window.

[0174] This specification also describes a method in which a system having a hybrid freelist can clock-gate regions of physical register files using data within the freelist. Clock-gating is a power-saving technique that refers to reducing power usage for specific regions of an electronic device.

[0175] FIG. 18 illustrates clock gating based on a freelist. One advantage of the hybrid freelist described herein is that it implicitly contains information regarding which regions of the physical register file are likely to be active and which regions are likely to be inactive. In particular, it is essentially certain that the regions between the head and tail pointers are the highly active regions of the corresponding physical register file. The system can utilize this information to intelligently control which regions of the physical register files may be clock-gated in order to save power.

[0176] FIG. 18 illustrates a hybrid freelist (1830) having 512 entries divided into banks of 16 entries each. The freelist (1830) assigns PRNs to a physical register file (1810) having a lower bank (1811) and an upper bank (1812).

[0177] The freelist (1830) has a head pointer (1815) and a tail pointer (1805). Physical registers having names represented by PRN entries belonging between the head and tail pointers (1815 and 1805) are physical registers likely to be read and written as the program is executed. Thus, the system can designate an area of ​​the physical register file (1810) corresponding to the PRN entry between the head power and the tail power as an active area (1840).

[0178] Other regions of the physical register file may be designated as "low active" or "inactive" if they meet certain criteria. In such cases, the system may clock-gate a portion of the register file or a subset of the register file. For example, the register file (1810) is divided into two banks (1811 and 1812). Since the latter part of the lower bank (1811) is a low-active region, the system may clock-gate the registers belonging to that portion of the lower bank (1811).

[0179] Alternatively or additionally, the system may clock-gate entire banks of the physical register file (1810). For example, if the registers of the upper bank (1812) are in a low-activity region because they are, for example, outside the head and tail pointers, the system may clock-gate the entire upper bank (1812).

[0180] The system may use any appropriate criterion from the freelist (1830) to determine that physical registers are in a low-activity or inactive state. For example, a row of PRN entries in the freelist (1830) having all cleared valid bits corresponds to a group of physical registers that are not used by the program. Accordingly, the system may clock gate the area, including reducing the power, clock frequency, voltage, or some combination thereof for that area.

[0181] In some situations, there may exist PRN entries containing data that is located outside the head and tail pointers but can nevertheless be read during program execution. For example, if the PRN is architectured into the AMT, the values ​​of the corresponding physical registers can still be read during program execution. However, writing to the PRN is no longer possible due to its architectured state.

[0182] Accordingly, for regions of the physical register file that have one or more PRN entries outside the head and tail pointers but still have valid bits set, the system may also designate these regions as active regions. Alternatively, or additionally, the system may perform less aggressive clock gating on these regions by, for example, setting these regions to a low-power mode where they can still be read but not written.

[0183] FIG. 19 is a flowchart of an example for clock gating a physical register file based on a freelist. The process may be performed by a processor configured according to the present specification. For convenience, the process is described as being performed by a system.

[0184] The system maintains a hybrid freelist for a physical register file partitioned into multiple regions (1910). The physical register file may be partitioned into regions that can be individually clock-gated, e.g., banks, rows, or any other suitable partitioning.

[0185] The system determines that the area of ​​the hybrid freelist represents an inactive area of ​​the physical register file (1920). In some implementations, the system searches only for inactive areas outside the head and tail of the freelist, for example, before the head pointer and after the tail pointer. As described above, the areas of the physical register file corresponding to the PRN entries between the head and tail pointers are likely to be areas with active reads and writes that occur as the program is executed.

[0186] The system can identify active regions from PRN entries by examining the valid bits of the PRN entries. A PRN entry with cleared valid bits indicates a physical register that is not being used. Therefore, if the system can identify adjacent regions of PRN entries that all have cleared valid bits, the system can identify the corresponding regions of the physical register file as inactive regions. One significant advantage of the hybrid freelist described herein is that it tends to group large regions of inactive PRN entries, which can then be used to identify inactive regions of the physical register file.

[0187] For PRN entries that are outside the head and tail pointers but nevertheless have their effective bits set, the system can designate that area as still active because the corresponding physical register can still be read during program execution.

[0188] The system reduces power to the inactive area of ​​the physical register file (1930). The system can, for example, reduce the clock frequency, voltage, or completely turn off the power to the inactive area of ​​the physical register file.

[0189] The system may re-evaluate the active and inactive regions of the physical register file at any appropriate time interval or when the head and tail pointers move. For example, if the head pointer moves to the next row of the freelist corresponding to an inactive region of the physical register file, the system may stop clock gating that region and restore the previous power or clock level because the physical registers in that region will soon be allocated for program execution. Similarly, when the tail pointer of the freelist moves to the next row, the system may examine the valid bits of the PRN entries of the previous row to determine whether all valid bits are cleared, in which case the system may clock gate the new region of the physical register file.

[0190] These techniques cause the system to activate specific regions of physical register files before the registers are required for execution, which reduces or eliminates any latency associated with clock gating the physical register files. For example, the system can use a freelist to reactivate inactive regions of physical register files at instruction issue time or register renaming time. In either case, the corresponding regions of the physical register files will be operational and executed until instructions are issued.

[0191] Embodiments of the subject matter and functional operations of the invention described herein may be implemented in digital electronic circuits, in tangibly implemented computer software or firmware, in computer hardware comprising the structures disclosed herein and their structural equivalents, or in combinations of one or more of these. Embodiments of the subject matter of the invention described herein may be implemented as one or more executable programs, namely, as one or more modules of instructions encoded on a tangible non-transient storage medium to be executed by a data processing device or to control the operation thereof. The storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of these. Alternatively or additionally, instructions may be encoded in artificially generated radio signals, for example, machine-generated electrical, optical, or electromagnetic signals generated to encode information for transmission to a receiver device suitable for execution by a data processing device.

[0192] The term “data processing device” refers to data processing hardware and encompasses all kinds of devices, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. Additionally, the device may be a special-purpose logic circuit, for example, a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC), or may further include such. Optionally, in addition to the hardware, the device may include code that creates an execution environment for computer programs, for example, processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.

[0193] A computer program, which may also be referred to or described as a program, software, software application, app, module, software module, script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; it may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may correspond to a file within a file system, but is not required to do so. A program may be stored as part of a file containing other programs or data, for example, one or more scripts stored in a markup language document, a single file dedicated to the program, or multiple coordinated files, for example, files storing one or more modules, subprograms, or parts of code. A computer program may be deployed to be executed on a single computer or a single site, or on multiple computers distributed across multiple sites and interconnected by a data communication network.

[0194] The processes and logic flows described in this specification may be performed by special-purpose logic circuits, e.g., FPGAs or ASICs, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0195] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.

[0196] In addition to the embodiments described above, the following embodiments are also innovative:

[0197] Example 1 is a method,

[0198] A step of maintaining a hybrid freelist representing physical registers available for register renaming—the hybrid freelist comprises a plurality of ordered entries, each entry comprising a pair of bits for each physical register of the plurality of physical registers, and

[0199] A pair of bits for each entry includes an effective bit and a commit bit representing the effective state and commit state for the physical register, respectively, and

[0200] The effective state indicates whether the physical register is free to be allocated for a logical register name, and

[0201] The commit status indicates whether instructions using physical registers have been committed -; and

[0202] It includes the step of using the head pointer and tail pointer of the freelist to perform sequential allocation of free physical register names from multiple entries of the freelist.

[0203] Example 2 is the method of Example 1,

[0204] It further includes a step of using the head pointer to find an entry that has a free state.

[0205] Example 3 is the method of Example 2, wherein the entries of the hybrid freelist are arranged in a plurality of banks, and,

[0206] The step of finding an entry that has a free state is,

[0207] A step of reading the bank corresponding to the head pointer; and

[0208] Includes the step of finding an entry in the bank that has a free state.

[0209] Example 4 is the method of Example 3,

[0210] A step of filling the output staging buffer based on the valid states of the bank entries; and

[0211] It further includes a step of assigning physical register names using entries in the output staging buffer.

[0212] Example 5 is the method of Example 4, wherein the step of filling the output staging buffer is,

[0213] It includes the step of generating all physical register names from an input staging buffer that stores valid bits from the entries of the bank.

[0214] Example 6 is a method of Example 5, wherein the step of generating a whole physical register name includes the step of extending each bit of the input staging buffer into a multi-bit physical register name of the output staging buffer.

[0215] Example 7 is a method of any one of Examples 1 to 6,

[0216] It further includes a step of modifying the valid state of an entry corresponding to a physical register name in response to an assigned physical register name.

[0217] Example 8 is a method of any one of Examples 1 to 7,

[0218] It further includes a step of modifying the commit state of the entry corresponding to the physical register when an instruction using the physical register is committed.

[0219] Example 9 is a method of any one of Examples 1 to 8,

[0220] It further includes a step of modifying the effective state and commit state of the entry corresponding to the physical register name assigned to the logical register name when a subsequent instruction writes to the same logical register name.

[0221] Example 10 is a method of any one of Examples 1 to 9,

[0222] A step of receiving a flush indication at a specific flush point;

[0223] It further includes a step of modifying the valid state of any entry that does not have a committed state in the hybrid freelist between the flush point and the head pointer.

[0224] Example 11 is a method of Example 10, further comprising the step of updating the head pointer of a hybrid freelist based on a flush point.

[0225] Example 12 is the method of Example 10, where the flush point is based on the oldest assigned PRN affected by the flush.

[0226] Example 13 is the method of Example 12,

[0227] A step of storing checkpoints for each PRN to LRN mapping in a reorder buffer (ROB); and

[0228] It further includes a step of using ROB's checkpoint as a flush point.

[0229] Example 14 is a method for implementing a shared freelist for a plurality of different register sets, wherein the method comprises:

[0230] A step of maintaining a shared freelist for multiple different register sets - each register set includes one or more registers of a specific type, and the freelist has multiple physical register name (PRN) entries -;

[0231] A step of generating a first physical register name for a first physical register of a first type for a first instruction referencing a first logical register of a first type from a first freelist entry;

[0232] A step of generating a second physical register name for a second physical register of type 2 for a second instruction referencing a second logical register of type 2 from a second freelist entry;

[0233] A step of retiring the first physical register name to the first architectural map table (AMT) for registers of the first type after the execution of the first instruction is completed;

[0234] A step of retiring the second physical register name for the second type of registers to the second AMT; and

[0235] It includes the step of designating the first physical register name as free if the first physical register name is not in the second AMT.

[0236] Example 15 is a method of Example 14, further comprising the step of designating a second physical register name as free when the first physical register name is not in the first AMT.

[0237] Example 16 is a method of Example 15, wherein the step of designating the second physical register name to a free state is:

[0238] A step of determining that the PRN entry of the freelist for the second physical register name does not have a set commit bit; and

[0239] In response to this, the step of clearing valid bits for the PRN entry is included.

[0240] Example 17 is a method of any one of Examples 14 to 16, wherein the step of designating the first physical register name to a free state is

[0241] A step of determining that the second AMT does not retain the second physical register name; and

[0242] In response to this, the method includes the step of clearing the valid bit for the PRN entry for the first physical register name.

[0243] Example 18 is a method of any one of Examples 14 to 17,

[0244] The method further includes the step of generating shared physical register names for a third physical register of type 1 and a fourth physical register of type 2 for a third instruction that references both a first logical register of type 1 and a second logical register of type 2 from a third hybrid freelist entry.

[0245] Example 19 is a method of Example 18, further comprising the step of returning a shared physical register name to a freelist only when the shared physical register name is expelled from both the first AMT and the second AMT.

[0246] Example 20 is a method of any one of Examples 14 to 19, wherein the first type is an integer register.

[0247] Example 21 is a method of any one of Examples 14 to 20, wherein the second type is a flag register.

[0248] Example 22 is a method,

[0249] A step of maintaining a shared freelist for multiple different register sets - each register set includes one or more registers of a specific type, and the freelist has multiple physical register name (PRN) entries -;

[0250] A step of assigning a first physical register name for a first physical register of a first type for a first instruction referencing a first logical register of a first type from a first freelist entry;

[0251] A step of determining that a second instruction referencing a second logical register of a different second type satisfies one or more borrowing criteria with a first instruction; and

[0252] In response to this, the method includes the step of reassigning the first physical register name to the second logical register referenced by the second instruction.

[0253] Example 23 is a method of Example 22, wherein the step of reassigning a first physical register name to a second logical register includes the step of assigning the first physical register name to a second logical register while the first physical register name is already assigned to the first logical register.

[0254] Example 24 is a method of any one of Examples 22 to 23, wherein the step of reassigning the first physical register name results in the first physical register name being simultaneously assigned to different types of logical registers.

[0255] Example 25 is a method of any one of Examples 22 to 24, wherein the step of determining that a second instruction satisfies one or more borrowing criteria includes the step of determining that the first instruction and the second instruction are consecutive instructions.

[0256] Example 26 is a method of any one of Examples 22 to 25, wherein the step of determining that the second instruction satisfies one or more borrowing criteria includes the step of determining that there are no intervening instructions requiring register renaming between the first instruction and the second instruction.

[0257] Example 27 is a method of Example 26, wherein the step of determining that a second instruction satisfies one or more borrowing criteria includes the step of determining that an intervening instruction does not require a logical register write.

[0258] Example 28 is a method of any one of Examples 26 to 27, and the intervening instruction is a store instruction.

[0259] Example 29 is a method of any one of Examples 22 to 28, wherein the step of determining that a second instruction satisfies one or more borrowing criteria includes the step of determining that the first instruction and the second instruction are renamed in the same cycle.

[0260] Example 30 is a method of any one of Examples 22 to 29, wherein the step of determining that a second instruction satisfies one or more borrowing criteria includes the step of determining that an intervening instruction between a first instruction and a second instruction is not a branch instruction.

[0261] Example 31 is a method of any one of Examples 22 to 30, further comprising the step of committing the first instruction and the second instruction in different cycles.

[0262] Example 32 is a method of any one of Examples 22 to 31, wherein the first type or the second type is an integer register.

[0263] Example 33 is a method of any one of Examples 22 to 32, wherein the first type or the second type is a flag register.

[0264] Example 34 is a method,

[0265] A step of maintaining a shared freelist for a plurality of different register sets, including a first register set and a different second register set—each register set includes one or more registers of a distinct type, the second register set has fewer physical registers than the first register set, the freelist has a plurality of physical register name (PRN) entries, the freelist associates one or more groups of PRN entries to identify single registers of the second register set, and at least one group of the plurality of PRN entries identifies a single physical register of the second register set—;

[0266] A step of assigning a first physical register name from a first PRN entry for a first logical register name of a second type, wherein the first PRN entry belongs to one or more other associated PRN entry groups that also identify the same physical register name of a second register set; and

[0267] It includes the step of designating one or more other PRN entries in the group as unavailable for allocation to logical registers of type 2 until the first PRN entry becomes available again.

[0268] Example 35 is the method of Example 34,

[0269] The method further includes the step of assigning a second physical register name for a second logical register name of a first type from a second PRN entry belonging to a group.

[0270] Example 36 is a method of any one of Examples 34 to 35, wherein a plurality of PRN entries within the group are assigned to physical registers of different types.

[0271] Example 37 is a method of any one of Examples 34 to 36, wherein each PRN entry includes a second register type bit indicating whether any PRN entries within a group of associated PRN entries are assigned to a second type of logical register.

[0272] Example 38 is a method of Example 37, wherein the step of designating one or more other PRN entries in a group as unavailable includes the step of setting each individual second register type bit for the first PRN entry and one or more other PRN entries in the group.

[0273] Example 39 is a method of any one of Examples 34 to 38, wherein the groups of PRN entries are pairs of PRN entries separated by N / 2 entries for a free list having N PRN entries.

[0274] Example 40 is a method of any one of Examples 34 to 39, wherein the second register set has half the number of physical registers of the first register set.

[0275] Example 41 is a method of any one of Examples 34 to 40, wherein the first register set includes integer physical registers and the second register set includes floating-point registers.

[0276] Example 42 is a method of Example 41, wherein the shared freelist is configured to maintain PRN entries for integer registers, flag registers, and floating-point registers.

[0277] Example 43 is a method,

[0278] A step of maintaining a shared freelist for a plurality of different register sets, including a first register set and a different second register set - each register set includes one or more registers of a distinct type, the second register set has fewer physical registers than the first register set, the freelist has a plurality of physical register name (PRN) entries, and the freelist maintains a plurality of PRN entries identifying a single register of the second register set -;

[0279] A step of allocating a first virtual PRN from a first PRN entry for a first logical register name of type 2 in a first instruction;

[0280] A step of allocating a second virtual PRN mapped to a first physical register name from a second PRN entry for a second logical register name of a second type within a second instruction;

[0281] A step of executing a first instruction using a first physical register name; and

[0282] It includes the step of issuing a second instruction for execution using the first physical register name only after the first virtual PRN has returned to the freelist.

[0283] Example 44 is a method of Example 43, wherein the step of determining that the first virtual PRN has been returned to the freelist includes the step of determining that the valid bit in the first PRN entry has been cleared.

[0284] Example 45 is a method of any one of Examples 43 to 44,

[0285] It further includes the step of maintaining a sliding window of Type 2 PRN entries ready to be issued for execution,

[0286] The step of issuing the second instruction includes the step of issuing the second instruction only after the second PRN entry has entered the sliding window.

[0287] Example 46 is the method of Example 45,

[0288] The method further includes the step of maintaining a floating-point tail pointer (FP tail pointer) representing the oldest assigned PRN entry of type 2 assigned to a physical register within a set of main physical registers of type 2, wherein the sliding window is defined by a predetermined number of PRN entries after the FP tail pointer.

[0289] Example 47 is the method of Example 46,

[0290] A step of determining that the associated PRN entry for the PRN entry indicated by the FP tail pointer is free for renaming; and

[0291] In response to this, it further includes a step to update the FP tail pointer.

[0292] Example 48 is the method of Example 47,

[0293] It further includes a step of clearing floating-point bits for the PRN entry represented by the FP tail pointer.

[0294] Example 49 is a method of Example 45, wherein the sliding window size is based on the number of registers in the second register set.

[0295] Example 50 is the method of Example 45,

[0296] A step of determining that the PRN assigned to the third logical register name from the third PRN entry to the third physical register of the second type is stuck in the architectured map table (AMT);

[0297] In response to this, a step of moving data from a third physical register to an architectured floating-point physical register file; and

[0298] It further includes the step of clearing the FP bit for the fourth PRN entry mapped to the same PRN stuck in the AMT.

[0299] Example 51 is the method of Example 50, wherein the step of determining that PRN is fixed to AMT is:

[0300] It includes a step of determining that the FP tail pointer is pointing to the third PRN entry while the PRN is not evicted from the AMT.

[0301] Example 52 is the method of Example 50,

[0302] Step of executing an instruction to read from a third logical register name; and

[0303] It further includes a step of reading values ​​from an architectured floating-point physical register file instead of the main floating-point physical register file.

[0304] Example 53 is a method of any one of Examples 43 to 52,

[0305] A step of issuing an instruction having a renamed register for execution to a reservation station;

[0306] A step of continuously checking whether the PRN entry used in the instruction is within the sliding window; and

[0307] It further includes a step of executing an instruction after the PRN entry enters the sliding window.

[0308] Example 54 is a method of any one of Examples 43 to 53, wherein the first register set includes integer physical registers and the second register set includes floating-point registers.

[0309] Example 55 is a method of Example 54, wherein the shared freelist is configured to maintain PRN entries for integer registers, flag registers, and floating-point registers.

[0310] Example 56 is a method,

[0311] In a processor, a step of maintaining a freelist that tracks the availability of physical registers in a physical register file - the freelist is partitioned into multiple regions, each region corresponding to a subset of the physical register file -;

[0312] A step of determining that the freelist area represents an inactive area of ​​the physical register file; and

[0313] In response to this, the method includes the step of clock-gating one or more physical registers within a physical register file corresponding to an inactive area of ​​the freelist.

[0314] Example 57 is a method of Example 56, wherein each freelist entry of the freelist has a valid bit indicating whether the corresponding physical register is in use, and

[0315] The step of determining that a freelist area is an inactive area includes the step of determining that all valid bits of a freelist entry within the area indicate that the corresponding physical register is not in use.

[0316] Example 58 is a method of any one of Examples 56 to 58,

[0317] A step of determining that the area of ​​the freelist is between the head pointer and the tail pointer of the freelist; and

[0318] In response to this, the method further includes a step of designating the region as an active region.

[0319] Example 59 is a method of any one of Examples 56 to 58, wherein a physical register file comprises a plurality of banks, and the step of determining that a freelist area is an inactive area includes the step of determining that all physical registers in the bank corresponding to the freelist area are not used.

[0320] Example 60 is a method of Example 59, wherein the step of reducing power for one or more physical registers includes the step of clock-gating a bank corresponding to an area of ​​a freelist while maintaining normal clock operation for one or more other banks of a physical register file.

[0321] Example 61 is a method of any one of Examples 56 to 60, wherein the step of clock-gating one or more physical registers occurs before one or more physical registers are assigned to an instruction.

[0322] Example 62 is a method of any one of Examples 56 to 61, further comprising the step of activating one or more physical registers during an instruction issue stage using a physical register among one or more physical registers.

[0323] Example 63 is a method of any one of Examples 56 to 62, further comprising the step of resuming the normal clock operation of a physical register during a register renaming process using a freelist.

[0324] Example 64 is a processor comprising a logic circuit configured to perform any one of the methods of Examples 1 to 63.

[0325] Example 64 is a data processing device configured to perform any one of the methods of Examples 1 to 63.

[0326] Although this specification contains many specific implementation details, they should not be interpreted as limitations on the scope of any invention or claimables, but rather as descriptions of features that may be assigned to specific embodiments of specific inventions. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while features may be described as operating in specific combinations and may even be initially claimed as such, one or more features from the claimed combination may be omitted from the combination in some cases, and the claimed combination may relate to a sub-combination or a variation of a sub-combination.

[0327] Similarly, although operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in the specific order depicted or in a sequential order, or that all illustrated operations be performed, in order to achieve desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the aforementioned embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.

[0328] Specific embodiments of the subject matter of the invention have been described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims may be performed in different orders and still achieve desirable results. As an example, the processes illustrated in the accompanying drawings do not necessarily require the specific order or sequential order illustrated to achieve desirable results. In certain cases, multitasking and parallel processing may be advantageous.

[0329] The scope of the claim is as follows.

Claims

Claim 1 A method comprising: maintaining a freelist in a processor that tracks the availability of physical registers in a physical register file, wherein the freelist is partitioned into a plurality of regions, each region corresponding to a subset of the physical register file; determining that a region of the freelist represents an inactive region of the physical register file; and in response thereto, reducing power to one or more physical registers in the physical register file corresponding to the inactive region of the freelist. Claim 2 A method according to claim 1, wherein each freelist entry of the freelist has a valid bit indicating whether the corresponding physical register is in use, and the step of determining that the area of ​​the freelist is an inactive area includes the step of determining that all valid bits of the freelist entries within the area indicate that the corresponding physical register is not in use. Claim 3 A method according to claim 1 or 2, further comprising the step of determining that the area of ​​the freelist is between the head pointer and the tail pointer of the freelist; and in response thereto, designating the area as an active area. Claim 4 A method according to any one of claims 1 to 3, wherein the physical register file comprises a plurality of banks, and the step of determining that the freelist area is an inactive area includes the step of determining that all physical registers within the bank corresponding to the freelist area are not used. Claim 5 A method according to claim 4, wherein the step of reducing power for one or more physical registers comprises the step of clock-gating a bank corresponding to an area of ​​a freelist while maintaining normal clock operation for one or more other banks of a physical register file. Claim 6 A method according to any one of claims 1 to 5, wherein the step of reducing power to one or more physical registers occurs before one or more physical registers are assigned to an instruction. Claim 7 A method according to any one of claims 1 to 6, further comprising the step of activating one or more physical registers during an instruction issuance stage using one of one or more physical registers. Claim 8 A method according to any one of claims 1 to 7, further comprising the step of resuming normal clock operation of a physical register during a register renaming process using a freelist. Claim 9 A processor comprising a logic circuit portion configured to perform operations, wherein the operations include: an operation of maintaining a freelist that tracks the availability of physical registers in a physical register file in the processor—the freelist is partitioned into a plurality of regions, each region corresponding to a subset of the physical register file—; an operation of determining that a region of the freelist represents an inactive region of the physical register file; and, in response thereto, an operation of reducing power to one or more physical registers in the physical register file corresponding to the inactive region of the freelist. Claim 10 A processor according to claim 9, wherein each freelist entry of the freelist has a valid bit indicating whether the corresponding physical register is in use, and the operation of determining that the area of ​​the freelist is an inactive area includes the operation of determining that all valid bits of the freelist entries within the area indicate that the corresponding physical register is not in use. Claim 11 A processor, wherein, in claim 9 or 10, the operations further include: an operation to determine that the area of ​​the freelist is between the head pointer and the tail pointer of the freelist; and, in response thereto, an operation to designate the area as an active area. Claim 12 A processor according to claim 10 or 11, wherein the physical register file comprises a plurality of banks, and the operation of determining that the freelist area is an inactive area includes the operation of determining that all physical registers within the bank corresponding to the freelist area are not used. Claim 13 A processor according to claim 12, wherein the operation of reducing power for one or more physical registers includes the operation of clock-gating a bank corresponding to an area of ​​a freelist while maintaining normal clock operation for one or more other banks of a physical register file. Claim 14 In any one of claims 10 to 13, the operation of reducing power to one or more physical registers occurs before one or more physical registers are assigned to an instruction, in a processor. Claim 15 A processor according to any one of claims 10 to 14, wherein the operations further include activating one or more physical registers during an instruction issue stage using one of one or more physical registers. Claim 16 A processor according to any one of claims 10 to 15, wherein the operations further include an operation to resume the normal clock operation of a physical register during a register renaming process using a freelist.