Register renaming cache

By introducing register renaming cache in the processor, learning and storing the register renaming attributes of the instruction stream, the problems of register renaming complexity and power cost in the prior art are solved, and more efficient performance and power consumption management are achieved.

CN120233944APending Publication Date: 2025-07-01INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411819429.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-11
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the complexity and power cost of register renaming increase rather than linearly as the distribution width of the core increases, and RAT consumes a large amount of power in a small area, affecting the core frequency.

Method used

By introducing a register rename cache, learn the register rename attributes of the instruction stream and store them in the cache, rename the learned registers from the cache when the same instruction trace reappears, reducing complexity and power consumption.

Benefits of technology

Effectively reduces the complexity and power consumption of register renaming, alleviates the timing path pressure of RAT, and improves overall performance and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233944A_ABST
    Figure CN120233944A_ABST
Patent Text Reader

Abstract

Techniques for register renaming caches are described. In an embodiment, an apparatus includes a register rename cache, a front-end circuit, a lookup circuit, and an execution circuit. The register renaming cache is to store register renaming information associated with the instruction trace. Register renaming information is learned from a first execution of the instruction trace and used to perform register renaming in conjunction with a second execution of the instruction trace. The front-end circuitry is to provide operations for execution based on the instruction trace. The lookup circuit is used for lookup of an entry corresponding to the operation in the register renaming cache. The execution circuitry is to execute the instruction trace.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Processors in computers and other information processing systems can use register renaming techniques to convert an in-order instruction stream into an out-of-order instruction stream by mapping logical registers to physical registers. Brief Description of the Drawings

[0002] Examples in accordance with the present disclosure will be described with reference to the accompanying drawings, in which:

[0003] Figure 1A Illustrates an instruction trace constructed for register renaming according to an embodiment.

[0004] Figure 1B Illustrates a register renaming pipeline according to an embodiment.

[0005] Figure 2 Illustrates a register renaming circuit according to an embodiment.

[0006] Figure 3 Illustrates a method for register renaming according to an embodiment.

[0007] Figure 4 Illustrates an example computing system.

[0008] Figure 5 Illustrates a block diagram of an example processor and / or system-on-a-chip (SoC) that may have one or more cores and an integrated memory controller.

[0009] Figure 6A Is a block diagram illustrating an example in-order pipeline and both an example register renaming, out-of-order issue / execution pipeline according to an example.

[0010] Figure 6B Is a block diagram illustrating an example in-order architecture core and both an example register renaming, out-of-order issue / execution architecture core to be included in a processor according to an example.

[0011] Figure 7 Illustrates an example of an execution unit circuit(s).

[0012] Figure 8 Is a block diagram illustrating an example of converting binary instructions in a source instruction set architecture into binary instructions in a target instruction set architecture using a software instruction converter according to an example. Detailed Description

[0013] The present disclosure relates to methods, apparatuses, systems, and non-transitory computer-readable storage media for register renaming caches. According to some examples, an apparatus includes a register renaming cache, a front-end circuit, a lookup circuit, and an execution circuit. The register renaming cache is used to store register renaming information associated with an instruction trace. The register renaming information is learned from a first execution of the instruction trace and is used to perform register renaming in combination with a second execution of the instruction trace. The front-end circuit is used to provide operations for execution based on the instruction trace. The lookup circuit is used to look up an entry corresponding to the operation in the register renaming cache. The execution circuit is used to execute the instruction trace.

[0014] As described in the background section, a processor or a processor core in a computer and other information processing systems can use register renaming techniques to convert an in-order instruction stream into an out-of-order instruction stream by mapping logical registers to physical registers. For example, register renaming techniques can use a structure (e.g., a register alias table, i.e., RAT) to maintain the mapping of logical registers to physical registers. The RAT can be read for each instruction to rename the instruction source to the physical mapping. Similarly, the logical destination register produced by an instruction can be mapped to a new physical register and written to the RAT as the latest mapping.

[0015] However, the complexity and power cost of register renaming increase non-linearly with the increase in the allocation width of the core. In addition, the RAT can create a power hot spot in the core because it consumes a large amount of power in a small area. Additionally, the use of the RAT is time-critical and affects the core frequency.

[0016] Due to in-line dependency calculations and / or the number of ports in the RAT array, the complexity and power cost of a RAT with a high allocation width can also be high. Regarding in-line dependency calculations, the complexity of this part of the register renaming logic is a quadratic function of the allocation width of the core. The logic level increases linearly with the increase in the allocation width from one generation of core to the next generation of core. Regarding the ports in the RAT array, an integer instruction can have at most three sources and one destination, so each integer instruction will use three times the allocation width of read ports and the number of write ports equal to the allocation width. Therefore, the total number of ports for an N-wide allocation is four times the allocation width. With the increase in the allocation width, the linear addition of ports comes at the cost of a super-linear increase in the dynamic capacitance and the timing complexity of accessing the RAT array for reading and writing.

[0017] Methods, apparatuses, systems, non-transitory computer-readable storage media, etc. according to embodiments can be provided to reduce the complexity and power consumption of register renaming. As an example, an embodiment can reduce the logic level for register renaming, thereby reducing the complexity of the in-line dependency circuitry. As another example, an embodiment can be implemented in combination with the use of a RAT (e.g., by reducing the number of RAT ports to reduce complexity and power consumption, and can be referred to as a RAT trap; however, embodiments are not limited to implementations including a RAT). As another example, an embodiment can allow for more complex elimination, thereby improving performance. As another example, an embodiment can allow a core to scale to a higher performance level with fewer register renaming (e.g., RAT) constraints. As another example, an embodiment can facilitate (e.g., via a compiler and a binary editor) reorganizing code to reduce register renaming (e.g., RAT) constraints.

[0018] Embodiments can be based on: treating register renaming as a property of a sequence of instructions in an allocation window rather than a property of an individual micro-operation (uop), and learning the register renaming properties of an instruction stream (e.g., the instructions between a taken branch and the next taken branch) and storing them in a cache (which can be referred to as a RAT trap cache, or more generally, a register renaming cache). When the same taken-taken trace reappears from a prediction (e.g., by a branch prediction unit or BPU, which can correspond to the branch prediction circuit 632 in Figure 6B as described below), the learned register renaming for the entire trace can be replayed from the register renaming cache because the sequence of uops is static between taken-taken traces. According to an embodiment, register renaming can perform register renaming across traces of taken-taken branches.

[0019] In an embodiment, the register renaming cache can be implemented to read out (live-in) and write back (live-out) only unique logical registers to, for example, a RAT for each trace. The use of embodiments can alleviate the critical timing path (e.g., of a RAT) and improve overall performance and power consumption.

[0020] The benefits of an embodiment can depend on the hit rate in the register renaming cache. Depending on constraints such as the number of sequential reservation registers, branch density, code size, etc., a register renaming cache organized as a taken-taken trace may have significant underutilization. Thus, embodiments can allow a compiler, binary editor (e.g., Binary Optimization Layout Tool (BOLT) or Propeller), etc. to reorganize the code so as to minimize the number of sequential register reservations and retirements and the taken-taken length of the frequent (profile-generated) path per register renaming instruction trace, which can improve the space utilization in the register renaming cache and allow for the construction of a power- and area-efficient register renaming cache (e.g., for large code footprint server applications).

[0021] In an embodiment, branch correlations (e.g., taken-taken branch correlations from the BPU) are learned and used through simple lookups (e.g., rather than invoking a complex RAT circuit). In an embodiment, the register renaming attributes of traces across a program are learned and stored in the register renaming cache, where the trace of the program spans one taken branch to the next taken branch (e.g., from the BPU, such as Figure 6B the branch prediction circuit 632 in, as described below). When the same trace in the program (e.g., predicted by the BPU) reappears, the register renaming attributes of the trace are read out from the cache, and renaming occurs through a simple circuit.

[0022] In an embodiment, renaming operations are learned across uops (which can be fused uops, i.e., fuops (fused uops) within an instruction trace; embodiments can be described based on fuops, but are not limited to fuop-based implementations) because the renaming of a fuop is not just an attribute of the fuop itself; rather, it depends on the sequence of instructions that are renamed along with it in the allocation window. When a fuop from a given basic block is the first fuop in the allocation window, its sources will be renamed (e.g., read through the RAT). When the same fuop is not the first fuop in the allocation window, its sources may be forwarded from the destinations of another previous fuop in the allocation window (inline dependencies). Thus, to learn register renamings and store them in the register renaming cache, deterministic bounds are set for each renaming operation that may occur within a cycle. In an embodiment, renaming operations are learned across fuops within a last_taken_target (last_taken_target) to next_taken_branch (next_taken_branch) (LTT-NTB) trace.

[0023] For example, Figure 1AFIG. illustrates constructing instruction traces for register renaming according to an embodiment. Given LTT-NTB traces can be very long (in terms of the number of fuops), so embodiments can include breaking them into capsules. For example, the capsule size can be based on the median trace length, such as sixteen fuops. In an embodiment, each trace longer than sixteen fuops in length is broken into capsules of sixteen fuops. Each capsule of a trace goes to a different entry in the register renaming cache, all entries belonging to the same set but in different ways and having the same tag. Thus, a register renaming cache hit can mean that multiple ways of the set are hit.

[0024] For example, as Figure 1A shown, the LTT-NTB trace A (100A) from A1 to An is divided into x + 1 capsules or chunks, each capsule or chunk having sixteen fuops (102A, 104A to 106A), and the LTT-NTB trace B (100B) from B1 to Bn is divided into y + 1 capsules or chunks, each capsule or chunk having sixteen fuops (102B, 104B to 106B).

[0025] Figure 1B FIG. illustrates a register renaming pipeline 110 according to an embodiment. The register renaming pipeline 110 can be implemented in a processor, a processor core, an execution core, etc., which can be any type of processor / core, including general microprocessors / cores, such as a processor family or processors / cores from a company or other processor families from another company, a dedicated processor or microcontroller, or any other device or component in an information processing system in which embodiments can be implemented. For example, the register renaming pipeline 110 can be implemented in any one of the processors 470, 480, or 415, Figure 4 as described below respectively, Figure 5 in a processor 500 or one of cores 502(A) to 502(N), and / or Figure 6B in core 690, implemented with circuits, logic gates, structures, hardware, etc., all or part of which can be included in discrete components and / or integrated into the circuitry of a processing device or any other device in a computer or other information processing system.

[0026] As shown in the figure, the register renaming pipeline 110 includes a front end (FE) 112, an instruction decode queue (IDQ) 114, a conventional renaming block 116, a lookup block 120, a register renaming cache 122, and an out-of-order (OOO) block 130. In an embodiment, the FE 112 may correspond to the front end unit 630 as described below Figure 6B In an embodiment, the conventional renaming block 116, the lookup block 120, and / or the register renaming cache 122 may be implemented in the renaming / allocator unit 652 as described below Figure 6B In an embodiment, the OOO block 130 may correspond to the scheduler 656 as described below Figure 6B

[0027] The FE 112 may represent circuitry, logic gates, structures, hardware, etc. for fetching instructions, scanning instructions, decoding instructions, etc., and generating as output one or more micro-operations, microcode entry points, microinstructions, other instructions, or other control signals decoded from the original instructions, or otherwise reflecting or derived from the original instructions. Decoding may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), microcode read only memories (ROMs), etc.

[0028] The FE 112 is coupled to the IDQ 114 for providing, for example, fuops at two fetch branches per cycle, and is coupled to the lookup block 120. According to an embodiment, the lookup block 120 may represent circuitry, logic gates, structures, hardware, etc. for performing lookups in the register renaming cache 122 (e.g., RAT trap). Note that in Figure 1B , the lookup block 120 is labeled RAT trap lookup and the register renaming cache 122 is labeled RAT trap renaming, but is shown only as a specific example of a more general register renaming method according to an embodiment. Two packages (124 and 126) are shown within the register renaming cache 122, and each package has, for example, ten to sixteen fuops as described above.

[0029] The conventional renaming block 116 may represent a pipeline implemented with circuitry, logic gates, structures, hardware, etc. for performing register renaming according to existing methods such as using a RAT. The conventional renaming block 116 and the register renaming cache 122 are coupled to provide up to 10 or 16 fuops respectively to the OOO block 130 for execution. ​

[0030] As shown, for example Figure 1B in FIG. 1, the register renaming cache 122 is looked up by the lookup block 120 in parallel with the write to the IDQ 114. The lookup can be performed at each LTT boundary switch. The register renaming cache hit pattern is read out, and the number of fuops coming out of the IDQ 114 is counted and matched / aligned with the fuops to be read out from the register renaming cache 122. On a lookup miss, the fuop is sent through the conventional renaming block 116, which can represent, for example, a ten-fuop-wide pipeline. The renaming information of the learned miss trace (e.g., as described below) is learned and, if there is available space, filled into the register renaming cache 122. On a hit, up to two packets (e.g., up to 32 fuops) can rename registers per cycle based on the information from the register renaming cache 122 to match the front-end bandwidth with two branches per cycle (where each branch can correspond to a 16-fuop-long packet, or less if the trace is interrupted by the end of a branch). In cases where the branch density is very high, the FE 112 can be limited to two taken branches and can be less than sixteen fuops. For this reason, to match the taken-branch bandwidth of the FE 112, the renaming using the register renaming cache 122 should also be able to support fuops worth two taken branches. For example, assume the FE 112 predicts two taken branches per cycle, and the first taken-taken trace contains only four fuops, and the next taken-taken trace contains eight fuops. All of these can be renamed through the register renaming cache 122, but the renaming information is stored separately for the first trace (four uops) and the second trace (eight uops). Therefore, the register renaming cache 122 uses two renaming packets (124 and 126).

[0031] In an embodiment, the register renaming cache (e.g., register renaming cache 122) can store the following information for each packet. · Reservation: For a packet, all unique register renamings (e.g., RAT) are read. For example, as Figure 1B indicated, there can be eight to ten unique sources per packet, corresponding to eight to ten read ports per packet (but in various embodiments, the number can depend on the performance, power, and area (PPA) configuration of the simulation). The remaining sources are forwarded. ·Live-In_ID / Fwd_Distance: For each source of each fuop, store where the source was renamed. If it was forwarded from the destination of one of the previous uops in the package, store that forward distance. If it came from a live-in, store the live-in identifier. · Logout: For a package, store the unique logical destination to which it is written. The last writer to the logical destination in the package is the owner of the logout. For example, as Figure 1B indicated, there can be five to seven logouts per package, corresponding to five to seven write ports per package (however, in various embodiments, the number can depend on the performance, power, and area (PPA) configuration of the simulation).

[0032] Figure 2 FIG. illustrates a register renaming circuit 200 according to an embodiment, which includes a RAT 210 and multiplexers 212, 214, and 216. The register renaming circuit 200 can be implemented in a processor, a processor core, an execution core, etc., which can be any type of processor / core, including general microprocessors / cores, such as a processor family or processors / cores in other processor families of a company or another company, dedicated processors or microcontrollers, or any other device or component in an information processing system in which the embodiment can be implemented. For example, the register renaming circuit 200 can be in any one of processors 470, 480, or 415 as described below respectively, Figure 4 in one of processors 500 or cores 502(A) to 502(N) in Figure 5 and / or Figure 6B in core 690 in

[0033] As Figure 2 shown, compared with existing methods, source renaming does not have complex control logic for determining the inline dependencies and selection lines of source renaming multiplexers (such as 212) and logout multiplexers (such as 214). The selection lines directly come from the memory reads of the register renaming cache. Similarly, no inline destination override logic is required because the selection lines for the live-in multiplexer (e.g., 216) directly come from the register renaming cache reads.

[0034] Embodiments may include software assistance for improving the hit rate in the register renaming cache (e.g., for large code with space server workloads). For example, application developers / compilers can reorganize the code to reduce the number of branches taken in the code and better organize the code to reduce the number of traces left and logged out, to improve the utilization of the register renaming cache ports and stores. Embodiments may include profile-guided learning of frequent paths (traces) and re-optimization of the source code and the code generator of the compiler to minimize register renaming constraints. For example, when the number of register renaming stays in sequential code increases, the compiler can optimize the generated code (e.g., similar to the way current compilers reduce register spills and fills).

[0035] In an embodiment, information is learned from the register renaming operations that occur between the last taken target and the next taken branch and stored in the register renaming cache in the enclosure for that trace. When replaying register renaming from the register renaming cache, the number of taken branches that are allowed to be renamed can be limited (e.g., limited to two). In some cases, if the taken branch density is too high, the bandwidth of register renaming will degrade in terms of the number of fuops renamed in a cycle. To address this issue, if some branches have a very high probability of being taken, embodiments can use software and compiler support to fix the polarity of the branches to a static value and not make them look like branches, thereby increasing the bandwidth of register renaming according to the embodiments.

[0036] In embodiments where the number of register renaming cache ports is low, the register renaming cache may be underutilized due to port constraints. Underutilization may result in a poor hit rate in the register renaming cache, leading to performance loss and increased power. In an embodiment, only the unique logical sources (stays) and destinations (logged out) for a trace are stored in one entry of the register renaming cache. When there are benchmarks that produce too many stays and logged outs, the compiler can reorganize the code or a binary editor can be used to minimize the number of stays and logged outs for register renaming traces, thereby improving the hit rate for large footprint workloads common in current servers.

[0037] Figure 3 FIG. illustrates a method 300 for register renaming according to an embodiment. Method 300 may be performed in whole or in part by apparatuses such as Figure 1B and / or Figure 2 shown and / or in conjunction with operations of apparatuses such as Figure 1B and / or Figure 2 shown; thus, Figure 1B and / or Figure 2 or with Figure 1B and / orFigure 2 All or any part of the foregoing description related thereto may apply to method 300. Note that, in Figure 3 , as described elsewhere in this specification and its accompanying drawings, reference may be made to RAT traps, but only as a specific example of a more general register renaming method according to an embodiment.

[0038] In 310 of method 300, for example, by a lookup block (e.g., the lookup block 120 in Figure 1B ), in response to an LTT boundary switch and in parallel with a write to an instruction decode queue (e.g., the IDQ 114 in Figure 1B ), a register renaming cache lookup is performed.

[0039] In 312, it is determined whether the lookup results in a hit in the register renaming cache (e.g., the register renaming cache 122). If not, method 300 continues in 314. If so, method 300 continues in 316.

[0040] In 314, in response to a register renaming cache miss, register renaming may be performed according to a previous method (e.g., using RAT) to provide register allocation performed in 318 and cause method 300 to continue in 320.

[0041] In 316, in response to a register renaming cache hit, register renaming may be performed based on information from the register renaming cache to provide register allocation performed in 318 and cause method 300 to continue in 330.

[0042] In 320, information about the attributes of the renaming performed in 314 (as described above) may be learned such that if it is determined in 322 that the register renaming cache is not full, the information may be stored in the register renaming cache in 324. However, if it is determined in 322 that the register renaming cache is full, then, if it is determined in 326 that there is a high taken branch density, in 328, the compiler or binary editor may reorganize the code to reduce taken branches, thereby providing another re-learning for filling the register renaming cache in 324.

[0043] In 330, it may be determined whether the renaming bandwidth is low due to register renaming cache port constraints, and if so, in 332, the compiler or binary editor may reorganize the code to reduce the retention and logout per trace to provide another re-learning for filling the register renaming cache in 324.

[0044] From 332, method 300 may return to 310.

[0045] Example devices, methods, etc.

[0046] According to some examples, a device (e.g., a processor core, a system, a system on a chip (SoC), etc.) includes a register renaming cache, a front-end circuit, a lookup circuit, and an execution circuit. The register renaming cache is used to store register renaming information associated with an instruction trace. The register renaming information will be learned from a first execution of the instruction trace and will be used to perform register renaming in combination with a second execution of the instruction trace. The front-end circuit is used to provide operations for execution based on the instruction trace. The lookup circuit is used to look up an entry corresponding to the operation in the register renaming cache. The execution circuit is used to execute the instruction trace.

[0047] Any such example may include any one or any combination of the following aspects. The instruction trace starts at a taken branch target and ends at the next taken branch. The register renaming information includes an identifier of each uniquely read logical register. The register renaming information includes an identifier of each uniquely written logical register. In response to a lookup hit in the register renaming cache, execution performs register renaming using the register renaming information in combination with a second execution of the instruction trace. In response to a lookup miss in the register renaming cache, execution learns the register renaming information in combination with a first execution of the instruction trace. The device includes a register alias table that is used to perform register renaming in combination with a first execution of the instruction trace. In response to a lookup miss in the register renaming cache, execution uses the register alias table to perform register renaming in combination with a first execution of the instruction trace.

[0048] According to some examples, a method includes: learning register renaming information from a first execution of an instruction trace; storing the register renaming information in a register renaming cache; and performing register renaming using the register renaming information in combination with a second execution of the instruction trace.

[0049] Any such example may include any one or any combination of the following aspects. An instruction trace begins at a taken branch target and ends at the next taken branch. The register renaming information includes an identifier for each uniquely read logical register. The register renaming information includes an identifier for each uniquely written logical register. The method includes looking up an entry corresponding to the instruction trace in a register renaming cache. In response to looking up an entry corresponding to the instruction trace that results in a hit in the register renaming cache, performing a second execution of register renaming that uses the register renaming information in combination with the instruction trace. In response to looking up an entry corresponding to the instruction trace that results in a miss in the register renaming cache, performing a first execution of learning the register renaming information in combination with the instruction trace. The method includes performing register renaming using a register alias table in combination with a first execution of the instruction trace. In response to looking up an entry corresponding to the instruction trace that results in a miss in the register renaming cache, performing register renaming using the register alias table in combination with a first execution of the instruction trace.

[0050] According to some examples, a method may include: profiling code to provide register renaming information to be stored in a register renaming cache from a first execution of an instruction trace in the code; and reorganizing the code to reduce register renaming constraints for performing register renaming using the register renaming information from the register renaming cache in combination with a second execution of the instruction trace.

[0051] Any such example may include any one or any combination of the following aspects. Reorganizing the code includes reorganizing the code to reduce sequential register retention or logout. Reorganizing the code includes reorganizing the code to reduce taken branches.

[0052] According to some examples, a device may include means for performing any of the functions disclosed herein; a device may include a data storage device that stores code that, when executed by a hardware processor or controller, causes the hardware processor or controller to perform any of the methods or any part of the methods disclosed herein; the means, methods, systems, etc. may be as described in the detailed description; a non-transitory machine-readable medium may store instructions that, when executed by a machine, cause the machine to perform any of the methods or any part of the methods disclosed herein. Embodiments may include any of the details, features, etc. described in this specification or combinations of details, features, etc.

[0053] Example computer architecture

[0054] The following describes an example computer architecture in detail. Other system designs and configurations of laptops, desktops, handheld personal computers (PCs), personal digital assistants, engineering workstations, servers, blade servers, network devices, network hubs, switches, routers, embedded processors, digital signal processors (DSPs), graphics devices, video game devices, set-top boxes, microcontrollers, cellular phones, portable media players, handheld devices, and various other electronic devices known in the art are also suitable. Generally, various systems or electronic devices capable of incorporating a processor and / or other execution logic as disclosed herein are generally suitable.

[0055] Figure 4 An example computing system is illustrated. The multi-processor system 400 is a system provided with interfaces and includes a plurality of processors or cores, the plurality of processors or cores including a first processor 470 and a second processor 480 coupled via an interface 450 such as a point-to-point (P-P) interconnect, a fabric, and / or a bus. In some examples, the first processor 470 and the second processor 480 are homogeneous. In some examples, the first processor 470 and the second processor 480 are heterogeneous. Although the example system 400 is shown as having two processors, the system may have three or more processors, or may be a single-processor system. In some examples, the computing system is a system-on-chip (SoC).

[0056] The processors 470 and 480 are shown as including integrated memory controller (IMC) circuits 472 and 482, respectively. The processor 470 also includes interface circuits 476 and 478; similarly, the second processor 480 includes interface circuits 486 and 488. The processors 470, 480 may exchange information using the interface circuits 478, 488, via the interface 450. The IMCs 472 and 482 couple the processors 470, 480 to respective memories, namely memory 432 and memory 434, which may be portions of the main memories locally attached to the respective processors.

[0057] The processors 470, 480 may each use interface circuits 476, 494, 486, 498 to exchange information with a network interface (NW I / F) 490 via separate interfaces 452, 454. The network interface 490 (e.g., one or more of an interconnect, a bus, and / or a fabric, and in some examples, a chipset) may optionally exchange information with the coprocessor 438 via the interface circuit 492. In some examples, the coprocessor 438 is a dedicated processor, such as, for example, a high throughput processor, a network or communication processor, a compression engine, a graphics processor, a general purpose graphics processing unit (GPGPU), a neural-network processing unit (NPU), an embedded processor, and so on.

[0058] A shared cache (not shown) may be included in either of the processors 470, 480, or external to both processors but connected to these processors via an interface (such as a P-P interconnect) such that if the processors are placed in a low power mode, the local cache information of either or both processors may be stored in the shared cache.

[0059] The network interface 490 may be coupled to a first interface 416 via the interface circuit 496. In some examples, the first interface 416 may be an interface such as a Peripheral Component Interconnect (PCI) interconnect, a PCI Express interconnect, or another I / O interconnect. In some examples, the first interface 416 is coupled to a power control unit (PCU) 417, which may include circuitry, software, and / or firmware for performing power management operations related to the processors 470, 480, and / or the coprocessor 438. The PCU 417 provides control information to a voltage regulator (not shown) to cause the voltage regulator to generate an appropriate regulated voltage. The PCU 417 also provides control information to control the generated operating voltage. In various examples, the PCU 417 may include various power management logic units (circuits) for performing hardware-based power management. Such power management may be controlled entirely by the processors (e.g., by various processor hardware, and it may be triggered by a workload and / or power, thermal constraints, or other processor constraints), and / or power management may be performed in response to an external source (such as a platform or a power management source or system software).

[0060] PCU 417 is shown as a separate logic from processor 470 and / or processor 480. In other cases, PCU 417 can execute on a given one or more cores in the core (not shown) of processor 470 or 480. In some cases, PCU 417 can be implemented as a (specialized or general-purpose) microcontroller or other control logic configured to execute its own dedicated power management code (sometimes referred to as P-code). In still other examples, the power management operations to be performed by PCU 417 can be implemented outside the processor, such as, by means of a separate power management integrated circuit (PMIC) or another component external to the processor. In still other examples, the power management operations to be performed by PCU 417 can be implemented within the BIOS or other system software.

[0061] Various I / O devices 414 can be coupled to the first interface 416 along with bus bridge 418, and bus bridge 418 couples the first interface 416 to the second interface 420. In some examples, one or more additional processors 415 (such as, co-processors, high throughput many integrated core (MIC) processors, GPGPUs, accelerators (such as, graphics accelerators or digital signal processing (DSP) units), field programmable gate arrays (FPGAs) or any other processor) are coupled to the first interface 416. In some examples, the second interface 420 can be a low pin count (LPC) interface. Various devices can be coupled to the second interface 420, including, for example, a keyboard and / or mouse 422, a communication device 427, and a storage circuit 428. The storage circuit 428 can be one or more non-transitory machine-readable storage media as described below, such as a disk drive or other mass storage device that can include instructions / code and data 430. Further, audio I / O 424 can be coupled to the second interface 420. Note that other architectures are possible in addition to the point-to-point architecture described above. For example, a system such as multi-processor system 400 can implement a multi-drop interface or other such architecture instead of a point-to-point architecture.

[0062] Example Core Architectures, Processors, and Computer Architectures

[0063] Processor cores can be implemented in different ways, for different purposes, and in different processors. For example, implementations of such cores can include: 1) general-purpose in-order cores intended for general computing; 2) high-performance general-purpose out-of-order cores intended for general computing; 3) specialized cores intended primarily for graphics and / or scientific (throughput) computing. Implementations of different processors can include: 1) a CPU that includes one or more general-purpose in-order cores intended for general computing and / or one or more general-purpose out-of-order cores intended for general computing; and 2) a coprocessor that includes one or more specialized cores intended primarily for graphics and / or scientific (throughput) computing. Such different processors give rise to different computer system architectures, which can include: 1) a coprocessor on a separate chip from the CPU; 2) a coprocessor in the same package as the CPU but on a separate die; 3) a coprocessor on the same die as the CPU (in which case such a coprocessor is sometimes referred to as specialized logic or as a specialized core, such specialized logic being, for example, integrated graphics and / or scientific (throughput) logic); and 4) a system-on-chip (SoC) that can include the described CPU (sometimes referred to as the (one or more) application core or (one or more) application processor), the above-described coprocessor, and additional functionality on the same die. Example core architectures are then described, followed by example processors and computer architectures.

[0064] Figure 5 Block diagram of an example processor and / or SoC 500 that can have one or more cores, and an integrated memory controller. The solid box depicts a processor 500 having a set with a single core 502(A), a system agent unit circuit 510, and a set of one or more interface controller unit circuits 516, while the optional addition of the dashed box depicts an alternative processor 500 having a set with multiple cores 502(A)-502(N), a set of one or more integrated memory controller units 514 in the system agent unit circuit 510, and specialized logic 508 and a set of one or more interface controller unit circuits 516. Note that the processor 500 can be Figure 4 one of processors 470 or 480, or coprocessors 438 or 415.

[0065] Accordingly, different implementations of the processor 500 can include: 1) a CPU, where the dedicated logic 508 is integrated graphics and / or scientific (throughput) logic (which can include one or more cores, not shown), and the cores 502(A)-502(N) are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, or a combination of both); 2) a coprocessor, where the cores 502(A)-502(N) are a large number of dedicated cores designed primarily for graphics and / or scientific (throughput); and 3) a coprocessor, where the cores 502(A)-502(N) are a large number of general-purpose in-order cores. Thus, the processor 500 can be a general-purpose processor, a coprocessor, or a special-purpose processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), an embedded processor, and so on. The processor can be implemented on one or more chips. The processor 500 can be part of one or more substrates and / or implemented on one or more substrates using any of a variety of process technologies, such as, for example, complementary metal-oxide-semiconductor (CMOS), bipolar CMOS (BiCMOS), P-type metal oxide semiconductor (PMOS), or N-type metal oxide semiconductor (NMOS).

[0066] The memory hierarchy includes one or more levels of (one or more) cache unit circuits 505(A)-504(N) within cores 502(A)-502(N), a collection of one or more shared cache unit circuits 506, and an external memory (not shown) coupled to a collection of (one or more) integrated memory controller unit circuits 514. The collection of one or more shared cache unit circuits 506 can include one or more intermediate levels of cache (such as level 2 (L2), level 3 (L3), level 4 (L4)) or other levels of cache (such as last level cache (LLC)) and / or combinations of the foregoing. Although in some examples, interface network circuit 512 (e.g., ring interconnect) provides an interface to dedicated logic 508 (e.g., integrated graphics logic), the collection of (one or more) shared cache unit circuits 506, and system agent unit circuit 510, alternative examples use any number of well-known techniques for providing an interface to such units. In some examples, coherence is maintained between the collection of (one or more) shared cache unit circuits 506 and one or more of cores 502(A)-502(N). In some examples, interface controller unit circuit 516 couples core 502 to one or more other devices 518, such as one or more I / O devices, storage devices, one or more communication devices (e.g., wireless network, wired network, etc.).

[0067] In some examples, one or more of cores 502(A)-502(N) are capable of implementing multithreading. System agent unit circuit 510 includes those components that coordinate and operate cores 502(A)-502(N). System agent unit circuit 510 can include, for example, a power control unit (PCU) circuit and / or a display unit circuit (not shown). The PCU can be the logic and components required to regulate the power states of cores 502(A)-502(N) and / or dedicated logic 508 (e.g., integrated graphics logic), or can include such logic and components. The display unit circuit is used to drive one or more externally connected displays.

[0068] Cores 502(A)-502(N) can be homogeneous in terms of instruction set architecture (ISA). Alternatively, cores 502(A)-502(N) can be heterogeneous in terms of ISA; that is, a subset of cores 502(A)-502(N) can be capable of executing an ISA, while other cores can be capable of executing only a subset of that ISA or another ISA.

[0069] Example Core Architecture - Block Diagrams of In-Order and Out-of-Order Cores

[0070] Figure 6A is a block diagram showing an example in-order pipeline according to the example and both an example register renaming and out-of-order issue / execution pipeline. Figure 6B is a block diagram showing an example in-order architecture core to be included in a processor according to the example and both an example register renaming and out-of-order issue / execution architecture core. Figure 6A - Figure 6B The solid-line boxes in show the in-order pipeline and in-order core, while the optional additions in the dashed-line boxes show the register renaming, out-of-order issue / execution pipeline and core. Considering that the in-order aspect is a subset of the out-of-order aspect, the out-of-order aspect will be described.

[0071] In Figure 6A the processor pipeline 600 includes a fetch stage 602, an optional length decoding stage 604, a decode stage 606, an optional allocation (Alloc) stage 608, an optional rename stage 610, a schedule (also known as dispatch or issue) stage 612, an optional register read / memory read stage 614, an execution stage 616, a write-back / memory write stage 618, an optional exception handling stage 622, and an optional commit stage 624. One or more operations can be performed in each of these processor pipeline stages. For example, during the fetch stage 602, one or more instructions are fetched from the instruction memory, and during the decode stage 606, the one or more fetched instructions can be decoded, an address using the forwarded register ports (e.g., load store unit (LSU) address) can be generated, and branch forwarding (e.g., immediate offset or link register (LR)) can be performed. In one example, the decode stage 606 and the register read / memory read stage 614 can be combined into one pipeline stage. In one example, during the execution stage 616, the decoded instructions can be executed, LSU address / data pipelining to an Advanced Microcontroller Bus (AMB) interface can be performed, multiplication and addition operations can be performed, arithmetic operations with branch results can be performed, and so on.

[0072] As an example, Figure 6BAn example register renaming, out-of-order issue / execution architecture core can implement pipeline 600 as follows: 1) Instruction fetch circuit 638 performs the fetch stage 602 and the length decoding stage 604; 2) Decoding circuit 640 performs the decoding stage 606; 3) Rename / allocator unit circuit 652 performs the allocation stage 608 and the rename stage 610; 4) (One or more) scheduler circuits 656 perform the scheduling stage 612; 5) (One or more) physical register file circuits 658 and memory unit circuit 670 perform the register read / memory read stage 614; (One or more) execution clusters 660 perform the execution stage 616; 6) Memory unit circuit 670 and (one or more) physical register file circuits 658 perform the write-back / memory write stage 618; 7) Various circuits may be involved in the exception handling stage 622; and 8) Retirement unit circuit 654 and (one or more) physical register file circuits 658 perform the commit stage 624.

[0073] Figure 6B Illustrates a processor core 690, which includes a front-end unit circuit 630 coupled to an execution engine unit circuit 650, and both are coupled to a memory unit circuit 670. The core 690 can be a reduced instruction set architecture computing (RISC) core, a complex instruction set architecture computing (CISC) core, a very long instruction word (VLIW) core, or a hybrid or alternative core type. As another option, the core 690 can be a specialized core, such as, for example, a network or communication core, a compression engine, a coprocessor core, a general purpose computing graphics processing unit (GPGPU) core, a graphics core, and so on.

[0074] The front-end unit circuit 630 may include a branch prediction circuit 632 coupled to an instruction cache circuit 634, the instruction cache circuit 634 being coupled to a translation lookaside buffer (TLB) 636, the translation lookaside buffer 636 being coupled to an instruction fetch circuit 638, and the instruction fetch circuit 638 being coupled to a decoding circuit 640. In one example, the instruction cache circuit 634 is included in the memory unit circuit 670 rather than in the front-end circuit 630. The decoding circuit 640 (or decoder) may decode the instructions and generate, as output, one or more micro-operations, micro-code entry points, micro-instructions, other instructions, or other control signals decoded from, otherwise reflecting, or derived from the original instructions. The decoding circuit 640 may further include an address generation unit (AGU, not shown) circuit. In one example, the AGU generates an LSU address using the forwarded register ports and may further perform branch forwarding (e.g., immediate-offset branch forwarding, LR register branch forwarding, etc.). The decoding circuit 640 may be implemented using a variety of different mechanisms. Examples of suitable mechanisms include, but are not limited to, lookup tables, hardware implementations, programmable logic arrays (PLAs), micro-code read only memories (ROMs), etc. In one example, the core 690 includes a micro-code ROM (not shown) or other medium (e.g., in the decoding circuit 640 or otherwise within the front-end circuit 630) storing micro-code for certain macro-instructions. In one example, the decoding circuit 640 includes a micro-operation (micro-op) or operation cache (not shown) to save / cache decoded operations, micro-tags, or micro-operations generated during the decoding phase or other phases of the processor pipeline 600. The decoding circuit 640 may be coupled to a rename / allocator unit circuit 652 in the execution engine circuit 650.

[0075] The execution engine circuit 650 includes a rename / allocator unit circuit 652 that is coupled to a retirement unit circuit 654 and a collection 656 of one or more scheduler circuits. The (one or more) scheduler circuits 656 represent any number of different schedulers, including reservation stations, a central instruction window, and the like. In some examples, the (one or more) scheduler circuits 656 may include an arithmetic logic unit (ALU) scheduler / scheduling circuit, an ALU queue, an address generation unit (AGU) scheduler / scheduling circuit, an AGU queue, and so on. The (one or more) scheduler circuits 656 are coupled to the (one or more) physical register file circuits 658. Each physical register file circuit in the (one or more) physical register file circuits 658 represents one or more physical register files, where different physical register files store one or more different data types, such as scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points, status (e.g., an instruction pointer that is the address of the next instruction to be executed), and so on. In one example, the (one or more) physical register file circuits 658 include a vector register unit circuit, a write mask register unit circuit, and a scalar register unit circuit. These register units can provide architectural vector registers, vector mask registers, general-purpose registers, and so on. The (one or more) physical register file circuits 658 are coupled to the retirement unit circuit 654 (also referred to as a retirement queue (“retire queue” or “retirement queue”)) to illustrate various ways in which register renaming and out-of-order execution can be implemented (e.g., using the (one or more) reorder buffers (ROBs) and the (one or more) retirement register files; using the (one or more) future files, the (one or more) history buffers, and the (one or more) retirement register files; using register mapping and register pools, and so on). The retirement unit circuit 654 and the (one or more) physical register file circuits 658 are coupled to the (one or more) execution clusters 660. The (one or more) execution clusters 660 include a collection of one or more execution unit circuits 662 and a collection of one or more memory access circuits 664. The (one or more) execution unit circuits 662 can perform various arithmetic, logical, floating-point, or other types of operations (e.g., shifts, additions, subtractions, multiplications) and can operate on various data types (e.g., scalar integers, scalar floating points, packed integers, packed floating points, vector integers, vector floating points). Although some examples may include multiple execution units or execution unit circuits dedicated to a particular function or set of functions, other examples may include only one execution unit circuit or multiple execution units / execution unit circuits that all perform all functions.(One or more) scheduler circuits 656, (one or more) physical register file circuits 658, and (one or more) execution clusters 660 are shown as potentially multiple because some examples create separate pipelines for certain types of data / operations (e.g., scalar integer pipeline, scalar floating-point / tight integer / tight floating-point / vector integer / vector floating-point pipeline, and / or memory access pipeline each having its own scheduler circuit, (one or more) physical register file circuits, and / or execution cluster—and in the case of a separate memory access pipeline, some examples where only the execution cluster of that pipeline has (one or more) memory access circuits 664). It should also be understood that in the case of using separate pipelines, one or more of these pipelines can be out-of-order issue / execution, and the remaining pipelines can be in-order issue / execution.

[0076] In some examples, the execution engine unit circuit 650 can perform load store unit (LSU) address / data pipelining to an Advanced Microcontroller Bus (AMB) interface (not shown), as well as address staging and write-back, data staging load, store, and branch.

[0077] A set of memory access circuits 664 is coupled to a memory unit circuit 670, which includes a data TLB circuit 672, which is coupled to a data cache circuit 674, which is coupled to a second-level (L2) cache circuit 676. In one example, the memory access circuit 664 can include a load unit circuit, a store address unit circuit, and a store data unit circuit, each of which is coupled to the data TLB circuit 672 in the memory unit circuit 670. The instruction cache circuit 634 is further coupled to the second-level (L2) cache circuit 676 in the memory unit circuit 670. In one example, the instruction cache 634 and the data cache 674 are combined into a single instruction and data cache (not shown) in the L2 cache circuit 676, a third-level (L3) cache circuit (not shown), and / or main memory. The L2 cache circuit 676 is coupled to one or more other levels of cache and ultimately to main memory.

[0078] The core 690 can support one or more instruction sets (e.g., x86 instruction set architecture (optionally with some extensions added with newer versions); MIPS instruction set architecture; ARM instruction set architecture (optionally with optional additional extensions such as NEON)), including the (one or more) instructions described herein. In one example, the core 690 includes logic for supporting a packed data instruction set architecture extension (e.g., AVX1, AVX2), thereby allowing operations used by many multimedia applications to be performed using packed data.

[0079] Example(s) of execution unit circuitry

[0080] Figure 7 The figure illustrates examples of (one or more) execution unit circuitry, such as Figure 6B the (one or more) execution unit circuitry 662. As shown, the (one or more) execution unit circuitry 662 may include one or more ALU circuits 701, an optional vector / single instruction multiple data (SIMD) circuit 703, a load / store circuit 705, a branch / jump circuit 707, and / or a floating-point unit (FPU) circuit 709. The ALU circuit 701 performs integer arithmetic and / or Boolean operations. The vector / SIMD circuit 703 performs vector / SIMD operations on packed data (such as SIMD / vector registers). The load / store circuit 705 performs load and store instructions to load data from memory into registers or store data from registers to memory. The load / store circuit 705 may also generate addresses. The branch / jump circuit 707 causes a branch or jump to a memory address depending on the instruction. The FPU circuit 709 performs floating-point arithmetic. The width of the (one or more) execution unit circuitry 662 varies depending on the example and may be in the range of, for example, from 16 bits to 1024 bits. In some examples, two or more smaller execution units are logically combined to form a larger execution unit (e.g., two 128-bit execution units are logically combined to form a 256-bit execution unit).

[0081] Program code can be applied to input information to perform the functions described herein and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microprocessor, or any combination thereof.

[0082] The program code can be implemented in a high-level procedural programming language or an object-oriented programming language in order to communicate with the processing system. If desired, the program code can also be implemented in assembly language or machine language. In fact, the mechanisms described herein are not limited to the scope of any particular programming language. In any case, the language can be a compiled language or an interpreted language.

[0083] Examples of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or a combination of such implementations. The examples can be implemented as a computer program or program code executing on a programmable system that includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0084] One or more aspects of at least one example can be implemented by representative instructions stored on a machine-readable medium that represent various logic within a processor, which instructions when read by the machine cause the machine to fabricate logic for performing the techniques described herein. Such representations, referred to as “intellectual property (IP) cores,” can be stored on tangible machine-readable media and supplied to various customers or manufacturing facilities to load into the manufacturing machines that actually make the logic or processor.

[0085] Such machine-readable storage media may include, but are not limited to, non-transitory tangible arrangements of articles manufactured or formed by a machine or device, including storage media such as: hard disks; any other type of disk (including floppy disks, optical disks, compact disk read-only memory (CD-ROM), rewritable compact disks (CD-RW), and magneto-optical disks); semiconductor devices such as read-only memory (ROM), random access memory (RAM) such as dynamic random access memory (DRAM) and static random access memory (SRAM), erasable programmable read-only memory (EPROM), flash memory, electrically erasable programmable read-only memory (EEPROM); phase change memory (PCM); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.

[0086] Accordingly, examples also include non-transitory tangible machine-readable media that contain instructions or contain design data, such as a Hardware Description Language (HDL), that define the structures, circuits, devices, processors, and / or system features described herein. Such examples may also be referred to as program products.

[0087] Emulation (including binary translation, code morphing, etc.)

[0088] In some cases, an instruction converter may be used to convert instructions from a source instruction set architecture to a target instruction set architecture. For example, the instruction converter may translate (e.g., using static binary translation, dynamic binary translation including dynamic compilation), morph, emulate, or otherwise convert the instructions into one or more other instructions to be processed by the core. The instruction converter may be implemented in software, hardware, firmware, or a combination thereof. The instruction converter may be on the processor, off the processor, or partially on the processor and partially off the processor.

[0089] Figure 8FIG. is a block diagram illustrating the use of a software instruction converter to convert binary instructions in a source ISA to binary instructions in a target ISA according to an example. In the illustrated example, the instruction converter is a software instruction converter, but alternatively, the instruction converter can be implemented in software, firmware, hardware, or various combinations thereof. Figure 8 Shows that a first ISA compiler 804 can be used to compile a program in a high-level language 802 to generate first ISA binary code 806 that can be natively executed by a processor 816 having at least one first ISA core. The processor 816 having at least one first ISA core represents any processor that can perform functions substantially the same as those of a processor having at least one first ISA core by compatibly executing or otherwise processing the following: 1) a substantial portion of the first ISA, or 2) a target code version of an application or other software targeted to run on an Intel processor having at least one first ISA core, so as to obtain substantially the same results as a processor having at least one first ISA core. The first ISA compiler 804 represents a compiler operable to generate first ISA binary code 806 (e.g., target code) that can be executed on the processor 816 having at least one first ISA core with or without additional linking processing. Similarly, Shows that a program in a high-level language 802 can be compiled using an alternative ISA compiler 808 to generate alternative ISA binary code 810 that can be natively executed by a processor 814 without a first ISA core. An instruction converter 812 is used to convert the first ISA binary code 806 into code that can be natively executed by the processor 814 without a first ISA core. The converted code need not be the same as the alternative ISA binary code 810; however, the converted code will perform the general operations and consist of instructions from the alternative ISA. Thus, the instruction converter 812 represents software, firmware, hardware, or a combination thereof that allows a processor or other electronic device without a first ISA processor or core to execute the first ISA binary code 806 through emulation, simulation, or any other process. Figure 8 References to "an example", "example", "an embodiment", "embodiment", etc. indicate that the described example or embodiment may include a particular feature, structure, or characteristic, but each example or embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same example or embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example or embodiment, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in connection with other examples or embodiments whether or not explicitly described.

[0090] References to "an example", "example", "an embodiment", "embodiment", etc. indicate that the described example or embodiment may include a particular feature, structure, or characteristic, but each example or embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same example or embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example or embodiment, it is considered within the knowledge of those skilled in the art to affect such feature, structure, or characteristic in connection with other examples or embodiments whether or not explicitly described.

[0091] In addition, in the various examples described above, unless otherwise specifically indicated, disjunctive language such as the phrase "at least one of A, B, or C" or "A, B, and / or C" is intended to be understood to mean A, B, or C, or any combination thereof (i.e., A and B, A and C, B and C, and A, B, and C). As used in this specification and the claims, and unless otherwise specified, the use of the ordinal adjectives "first", "second", "third", etc. to describe elements merely indicates a particular instance of the element being referenced or different instances of similar elements, and is not intended to imply that the elements so described must be in a particular order in time, space, rank, or any other manner. Additionally, as used in the description of the embodiments, the " / " character between items may mean that the described content may include the first item and / or the second item (and / or any other additional item), or may be used, utilized, and / or implemented in accordance with the first item and / or the second item (and / or any other additional item).

[0092] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. However, it will be apparent that various modifications and changes can be made to the present disclosure without departing from the broader spirit and scope of the present disclosure as set forth in the claims.

Claims

1. A device for register renaming cache, comprising: a register renaming cache to store register renaming information associated with an instruction trace, the register renaming information learned from a first execution of the instruction trace and used to perform register renaming in conjunction with a second execution of the instruction trace; a front end circuit for providing operations for execution based on the instruction trace; a search circuit configured to search the register rename cache for an entry corresponding to the operation; as well as An execution circuit is used to execute the instruction trace.

2. The device according to claim 1, wherein: The instruction trace starts at a taken branch target and ends at the next taken branch.

3. The device according to claim 1, wherein: The register renaming information includes an identifier for each unique logical register that is read.

4. The device according to claim 1, wherein: The register renaming information includes an identifier for each unique logical register written.

5. The device according to claim 1, wherein: In response to a lookup hit in the register renaming cache, performing register renaming using the register renaming information in conjunction with the second execution of the instruction trace is performed.

6. The device according to claim 1, wherein: Responsive to a lookup miss in the register renaming cache, performing the first execution in conjunction with the instruction trace to learn the register renaming information. 7 . The apparatus of claim 1 , further comprising a register alias table for performing register renaming in conjunction with the first execution of the instruction trace.

8. The device according to claim 7, wherein: In response to a lookup miss in the register renaming cache, performing register renaming using the register alias table in conjunction with the first execution of the instruction trace is performed.

9. A method for register renaming cache, comprising: learning register renaming information from a first execution of an instruction trace; Storing the register renaming information in a register renaming cache; as well as Register renaming is performed using the register renaming information in conjunction with a second execution of the instruction trace.

10. The method according to claim 9, wherein: The instruction trace starts at a taken branch target and ends at the next taken branch.

11. The method according to claim 9, wherein: The register renaming information includes an identifier for each unique logical register that is read.

12. The method according to claim 9, wherein: The register renaming information includes an identifier for each unique logical register written.

13. The method of claim 9, further comprising looking up an entry in the register rename cache corresponding to the instruction trace.

14. The method according to claim 13, wherein: In response to looking up an entry in the register rename cache corresponding to the instruction trace that resulted in a hit, performing register renaming using the register renaming information in conjunction with a second execution of the instruction trace.

15. The method according to claim 13, wherein: Responsive to looking up an entry in the register rename cache corresponding to the instruction trace that caused the miss, performing the first execution in conjunction with learning the register renaming information in conjunction with the instruction trace.

16. The method of claim 9, further comprising performing register renaming using a register alias table in conjunction with the first execution of the instruction trace.

17. The method according to claim 16, wherein: In response to looking up an entry in the register rename cache corresponding to the instruction trace that caused the miss, performing register renaming using the register alias table in conjunction with the first execution of the instruction trace.

18. A method for register renaming cache, comprising: profiling the code to provide register renaming information learned from a first execution of an instruction trace in the code to be stored in a register renaming cache; as well as The code is reorganized to reduce register renaming constraints on performing register renaming using the register renaming information from a register renaming cache in conjunction with a second execution of the instruction trace.

19. The method according to claim 18, wherein: Reorganizing the code includes reorganizing the code to reduce sequential register saves or checkouts.

20. The method according to claim 18, wherein: Reorganizing the code includes reorganizing the code to reduce taking branches.