Dynamically configurable over-provisioned microprocessor

A dynamically configurable over-provisioned microprocessor optimizes performance and efficiency by dynamically activating and deactivating resources based on operating conditions, addressing thermal and power density challenges in modern microprocessors.

JP7853286B2Active Publication Date: 2026-04-28ADVANCED MICRO DEVICES INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ADVANCED MICRO DEVICES INC
Filing Date
2021-09-28
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

As transistor sizes decrease, modern microprocessors face challenges with increased power density and thermal energy, leading to thermal runaway and inefficiencies in computational workload performance, while specialized microprocessors for specific workloads are not cost-effective.

Method used

A dynamically configurable over-provisioned microprocessor design that includes more physical computing resources and long control spans, allowing dynamic activation and deactivation of resources based on operating conditions to optimize performance, energy consumption, and clock frequency.

Benefits of technology

The design efficiently balances computational performance, energy consumption, and clock frequency by dynamically configuring resources, enabling optimal workload performance across various applications without the need for specialized processors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007853286000001
    Figure 0007853286000001
  • Figure 0007853286000002
    Figure 0007853286000002
  • Figure 0007853286000003
    Figure 0007853286000003
Patent Text Reader

Abstract

A dynamically configurable over-provisioned microprocessor uses a general-purpose microprocessor design to optimally support a variety of different computing application workloads, with the ability to trade-off between computing performance, energy consumption, and clock frequency for each computing application. In some embodiments, the over-provisioned microprocessor comprises physical computing resources and dynamic configuration logic, the dynamic configuration logic configured to detect an activation-guaranteed operating condition, unencrypt the physical computing resources in response to detecting the activation-guaranteed operating condition, detect a configuration-guaranteed operating condition, and dynamically configure the over-provisioned microprocessor to use the unencrypted physical computing resources in response to detecting the configuration-guaranteed operating condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to general-purpose microprocessors such as central processing units (CPUs) in consumer-class personal computing devices and enterprise-class server computers. More particularly, some embodiments relate to dynamically configurable overprovisioned microprocessors.

Background Art

[0002] Until recently, scaling of complementary metal-oxide-semiconductor (CMOS) technology has been continuously advancing. During this period, metal-oxide-semiconductor field-effect transistors (MOSFETs) have become smaller, and transistor density has increased according to Moore's law. Furthermore, the dynamic switching power consumption per transistor has also decreased according to Dennard's scaling law. As a result, designers and manufacturers of single-core microprocessor chips have been able to increase the clock frequency from one microprocessor generation to the next without significantly increasing the overall power density.

[0003] Recently, the size of transistors has been reduced to the point where it reaches the limit of Dennard's scaling law for single-core microprocessors. In particular, at small transistor sizes (e.g., less than 65 nanometers), increased current leakage and increased power density increase the thermal energy within the microprocessor, threatening thermal runaway that can destroy the chip itself. As a result, as transistor size continues to decrease along with the desire to increase computational workload performance, microprocessor chip designers and manufacturers have not focused as much on increasing the clock frequency in single-core microprocessors, but have focused more on multi-core general-purpose microprocessor designs and special chips such as accelerators or application-specific integrated circuits (ASICs).

[0004] Unfortunately, these multicore designs are also approaching the limits of Dennard scaling law. As transistor sizes become smaller and transistor density increases in multicore designs, some of the transistors in a multicore microprocessor may be "dark" at any given time in order to stay within power constraints and avoid thermal runaway. More specifically, the higher power density of modern multicore designs, facilitated by increasingly smaller transistor sizes, hinders the ability to power on all transistors simultaneously at the nominal operating voltage within thermal design power (TDP) constraints. A substantial portion of the microprocessor may be dark (unused) at any given time. This dark portion is sometimes called "dark silicon."

[0005] Cryogenic cooling of a microprocessor (e.g., using liquid nitrogen, liquid helium, or other suitable cryogen) reduces current leakage energy. This allows a larger proportion of all transistors to be powered on simultaneously at the nominal voltage while remaining within TDP constraints. Cryogenic operation has other beneficial properties. In particular, transistors switch faster, enabling the microprocessor to operate at higher clock frequencies, and integrated circuit wires have lower electrical resistance, resulting in less signal delay.

[0006] Many computing application domains require cost-effective reductions in computation time to solve problems. Such domains include, for example, machine learning, games, image and video editing, and graph processing, among others. One possible solution to this requirement is to design and manufacture specialized microprocessors specifically engineered to improve the computational performance of particular computational workloads compared to more general-purpose microprocessors. An example of such a specialized microprocessor is one specifically designed for cryogenic operation (e.g., using cryogens between approximately 100 and 4 degrees A / B). However, due to the high overhead of design and manufacturing, designing and manufacturing specialized microprocessors for different computational workloads is generally not cost-effective.

[0007] The approaches described in this section are feasible approaches, but not necessarily approaches that have been previously conceived or performed. Therefore, unless otherwise indicated, none of the approaches described in this section should be assumed to be qualified as prior art simply by being included in this section.

[0008] Some embodiments are shown in the figures of the attached drawings as examples, not as limitations, and similar reference numerals refer to similar elements. [Brief explanation of the drawing]

[0009] [Figure 1] This is a schematic diagram of an exemplary microprocessor in which the techniques disclosed herein for a dynamically configurable over-provisioned microprocessor may be implemented, according to several embodiments. [Figure 2] Figure 1 is a schematic diagram of the core of an exemplary microprocessor according to several embodiments. [Figure 3] Figure 1 is a state diagram illustrating the exemplary states of computing resources managed by dynamic configuration logic for computing resources within the microprocessor, according to several embodiments. [Figure 4] Figure 1 is a schematic diagram of a dynamically configurable hybrid in-order / out-of-order CPU of a microprocessor core, according to several embodiments. [Figure 5] Figure 1 is a schematic diagram of a dynamically configurable memory-level parallel processing unit within the CPU of a microprocessor core, according to several embodiments. [Figure 6] Figure 1 is a schematic diagram of a dynamically configurable simultaneous multithreading unit within the CPU of a microprocessor core, according to several embodiments. [Modes for carrying out the invention]

[0010] The figures show several embodiments for the purpose of providing clear examples, but some embodiments may omit, add, rearrange, or modify any of the elements shown in the figures.

[0011] The following description includes many specific details to provide a thorough understanding of several embodiments for illustrative purposes. However, it will be apparent that some embodiments can be carried out without these specific details. In other examples, well-known structures and devices are shown in block diagrams to avoid unnecessarily obscuring some embodiments.

[0012] (overview) To provide optimized multi-purpose processing capabilities using a general-purpose microprocessor design, the microprocessor is over-provisioned with physical computing resources such as more transistors, longer control spans, and larger data storage structures. During operation, the over-provisioned microprocessor is dynamically configured to activate (de-darken) its computing resources. These activated computing resources are then used during operation to provide more optimal computing workload performance for a given computing application or a given portion of a computing application.

[0013] When an activated resource is no longer needed, the microprocessor may be dynamically configured to deactivate (dimm) the resource to reduce energy consumption or to increase the clock frequency. Throughout the process of handling a given computational workload, different computational resources may be dynamically activated and deactivated to balance computational performance, energy consumption, and clock frequency.

[0014] In some embodiments, an over-provisioned microprocessor is dynamically configured to activate computing resources in response to detecting activation-guaranteed operating conditions. One non-limiting example of an activation-guaranteed operating condition is cryogenic operation of an over-provisioned microprocessor. In this case, the over-provisioned microprocessor can be configured to dynamically activate computing resources so that alternative computing resources can be activated simultaneously, in which case one set of computing resources is currently in use while an alternative set of computing resources is activated and made available. Simultaneously activating alternative computing resources allows for efficient dynamic configuration of the over-provisioned microprocessor from using one set of computing resources to using an alternative set of computing resources, without having to wait for the alternative set of computing resources to become active (de-darken) after the configuration decision has been made. This activation-guaranteed operating condition and other activation-guaranteed operating conditions are described in more detail below.

[0015] In some embodiments, an over-provisioned microprocessor is dynamically configured to switch the use of computing resources in response to the detection of configuration-guaranteed operating conditions. One non-limiting example of a configuration-guaranteed operating condition is underutilization of reorder buffers. In this case, the over-provisioned microprocessor can be dynamically configured to switch from using computing resources for out-of-order instruction execution to using computing resources for in-order instruction execution. This configuration-guaranteed operating condition and others are described in more detail below.

[0016] Therefore, a technique is provided for the dynamic configuration of over-provisioned microprocessors that uses a general-purpose microprocessor design to optimally support a variety of different computing application workloads and has the ability to trade off computing performance, energy consumption, and clock frequency for each computing application.

[0017] (Example microprocessor) Figure 1 is a schematic diagram of exemplary microprocessors in which the techniques disclosed herein for a dynamically configurable over-provisioned microprocessor 100 may be implemented according to several embodiments. As used herein, the term “dynamic” in dynamically configurable means that the over-provisioned microprocessor is configured in operation while executing one or more computational tasks (e.g., processes or threads) without requiring the tasks to be restarted. Alternatively, the over-provisioned microprocessor can continue to execute tasks after configuration.

[0018] The microprocessor 100 has two or more separate cores 102-1, ..., 102-N that support parallel processing and multitasking. Each core 102 has its own central processing unit 104, its own set of registers 106, and its own cache 108. The cores 102-1, ..., 102-N are physically coupled to a bus 110 via one or more intermediate components, which may not be shown, for sending and receiving data and commands between the cores 102-1, ..., 102-N and a memory device 112 and an input / output device 114.

[0019] Figure 2 is a schematic diagram of core 102-1 according to several embodiments. Other cores 102 of the microprocessor 100 may have the same or equivalent components. However, heterogeneous cores 102-1, ..., 102-N are also possible. Core 102-1 may be a multithreaded central processing unit (CPU), or it may be a single-threaded core of the microprocessor 100 being multithreaded. Core 102-1 may utilize general-purpose processor design techniques, including but not limited to superscalar architecture, simultaneous multithreading, fine-grained multithreading, speculative execution, branch prediction, out-of-order execution, and / or register renaming. Core 100 may include physical computing resources for executing instructions according to a predefined instruction set architecture. For example, a given instruction set architecture may be X86, ARM, POWERPC, MIPS, SPARC, RISC, or any other complex or reduced instruction set architecture. A non-exclusive set of physical computing resources that may be included in core 100 may include an instruction fetch unit 110, an instruction cache 115, a decode unit 120, a register renaming unit 125, an instruction queue 130, an execution unit 135, a load / store unit 140, a data cache 140, and other circuitry in core 100. Other computing resources that may be included in core 102-1 (not shown) may include, but are not limited to, a prefetch buffer, branch prediction logic, global / bimodal logic, loop logic, indirect jump logic, loop stream decoder, microinstruction sequencer, retirement register file, register allocation table, reorder buffer, reservation station, arithmetic logic unit, or memory ordering buffer.

[0020] The microprocessor described above is presented for the purpose of illustrating an example of a basic microprocessor in which several embodiments can be implemented. However, it should be understood that other microprocessors, including those with more, fewer, or different computing resources than those described above, may be used in implementations. Furthermore, for illustrative purposes, the following description presents an example of a dynamically configurable over-provisioned microprocessor in a multi-core microprocessor context. However, some embodiments are not limited to any particular microprocessor configuration. In particular, the multi-core microprocessor is used to provide a framework for explanation, although it is not required for all implementations. Instead, some embodiments may be implemented in any type of microprocessor or other integrated circuit capable of supporting the methods of the embodiments presented in detail below.

[0021] (Over-provisioned microprocessors) As illustrated by the examples described below, in some embodiments, the microprocessor 100 is over-provisioned with physical computing resources. Here, “over-provisioned” includes the microprocessor 100 having more physical computing resources than can be powered on simultaneously at the nominal operating voltage within the target thermal design power (TDP) constraint. The target TDP constraint may be based on (assumed) non-cryogenic operation of the microprocessor 100. For example, the target TDP may be based on “room temperature” operation where the use of a cryogen (e.g., liquid nitrogen, liquid helium, etc.) between 100 and 4 degrees absolute temperature is not used.

[0022] Additionally or alternatively, "over-provisioned" includes the microprocessor 100 having a long physical control span (long physical signaling wire) between computational resources. The physical length of the control span may be too long to meet the signaling timing constraints at the target clock frequency of the microprocessor 100 in room temperature operation. For example, it may be too long to meet the signaling timing constraints at the above "on-the-box" clock frequency of the microprocessor 100 in non-cryogenic operation. In non-cryogenic operation (e.g., room temperature operation), the electrical resistance is greater in the control span than in cryogenic operation. As a result, to use these long control spans and still meet the signaling timing constraints, the clock frequency of the microprocessor 100 may need to be reduced (underclocked), or cryogenic operation of the microprocessor 100 may be required.

[0023] According to some embodiments, the microprocessor 100 can be dynamically configured to use a long control span when it detects that the microprocessor 100 is in cryogenic operation or when it detects that operation at the target clock frequency is not required (e.g., due to large-scale main memory I / O), and thus can temporarily reduce the operating clock frequency to enable the use of a long control span. A long control span can be used to connect computational resources that are not normally connected to each other in this way for timing constraints. By using a long control span between computational resources that are not typically connected in this way, dynamic configuration of the microprocessor 100 that utilizes the long control span is possible. Examples of such dynamic configurations are described in more detail below.

[0024] The overprovisioning of the microprocessor 100 can take various forms. According to some embodiments, at least three different forms, namely, (1) alternative computing resources, (2) long control spans, and (3) extended data storage structure headroom, are conceivable. Examples of each of these forms are provided in more detail below.

[0025] (Alternative computing resources) Generally, alternative computing resources are overprovisioned computing resources of the microprocessor 100 that can be used as an alternative during operation. An example of alternative computing resources is an in-order execution unit versus an out-of-order execution unit. Using the techniques disclosed herein, the microprocessor 100 can be overprovisioned with computing resources for both in-order and out-of-order execution, and the overprovisioned microprocessor 100 can be dynamically configured to use one or the other, for example, in response to detecting a configuration-guaranteed operating condition such as excessive stalls when using in-order execution computing resources. In this example, the overprovisioned microprocessor 100 can be configured to use out-of-order execution computing resources. This example and other examples of utilizing alternative computing resources in the overprovisioned microprocessor 100 are described in more detail below.

[0026] Another example of a computing resource that can be used as an alternative is a simple load-store unit for in-order main memory access, relative to the associated load-store unit for parallel main memory access. Using the techniques disclosed herein, the microprocessor 100 can be over-provisioned with computing resources for both in-order and out-of-order memory access, and the over-provisioned microprocessor 100 can be dynamically configured to use one or the other in response to detection of configuration-guaranteed operating conditions, such as the execution of a computationally limiting application when using out-of-order memory access computing resources. In this example, the over-provisioned microprocessor 100 can be configured to use in-order memory access computing resources. This and other examples of utilizing alternative computing resources in the over-provisioned microprocessor 100 are described in more detail below.

[0027] (Long control span) According to some embodiments, the microprocessor 100 may be over-provisioned with a long control span, which can be used to enable communication between computing resources that are not normally connected to each other. As shown by the following examples, a long control span can be used to implement dynamic configuration of the over-provisioned microprocessor 100 between in-order and out-of-order execution, and dynamic configuration of the over-provisioned microprocessor 100 between single-threaded processing mode and concurrent multi-threaded processing mode.

[0028] (Extended data storage structure headroom) The microprocessor 100 may include many data storage structures, such as register files, rename tables, reorder buffers, load / store units, instruction queues, and other physical data storage structures having a fixed number of entries for storing data items. The fixed number (which may vary between different structures) is typically determined during microprocessor design based on timing constraints at the target clock frequency and target TDP in non-cryogenic operation. In some embodiments, data storage structures are designed in the over-provisioned microprocessor 100 to have an extended number of entries to increase the data storage headroom of the structures. For example, the over-provisioned microprocessor 100 may be dynamically configured to use an extended register file. As another example, the over-provisioned microprocessor 100 may be dynamically configured to increase the instruction window size using an extended data storage structure. These and other examples of utilizing extended data storage structures in the over-provisioned microprocessor 100 are described in more detail below.

[0029] (Dynamic configuration logic) According to some embodiments, the over-provisioned microprocessor 100 is configured using one or more dynamic configuration logics for dynamically configuring the physical computing resources of the microprocessor 100. The dynamic configuration logic can be implemented using firmware, finite state machine logic, or other suitable logic. Different dynamic configuration logics may dynamically configure different computing resources, or a single dynamic configuration logic may serve to dynamically configure multiple computing resources.

[0030] According to some embodiments, the physical computing resources of a microprocessor 100, which can be dynamically configured by dynamic configuration logic, may be power-gated. Power gating refers to a technique in a microprocessor to reduce leakage power consumption by computing resources when they are not in use. Power gating can be implemented in the microprocessor 100, for example, using P-type metal-oxide-semiconductor (PMOS) or N-type metal-oxide-semiconductor (NMOS) sleep transistors. Alternatively, other circuit techniques can be used to place the physical computing resources into a dormant, sleep, or other low-power state. For example, voltage scaling techniques for reducing static power consumption by computing resources, as described in the following paper, can be applied to alternate computing resources between an active and a dormant state. K. Flautner, Nam Sung Kim, S. Martin, D. Blaauw and T. Mudge, "Drowsy caches: simple techniques for reducing leakage power," Proceedings of the 29th Annual International Symposium on Computer Architecture, Anchorage, AK, USA, 2002, pp. 148-157. A potential advantage of using this voltage scaling technique is that it requires fewer clock cycles to transition computing resources between active and sleep states compared to power gating.

[0031] (Power status) Figure 3 is a state diagram illustrating the exemplary power state of a computing resource managed by dynamic configuration logic for computing resources, according to several embodiments. Initially, the computing resource may be in a dark state 332. In the dark state 332, the computing resource may be power-gated and not being used for computing tasks.

[0032] When the dynamic configuration logic detects the activation guarantee operating conditions, it can transition the computing resource from the dark state 332 to an active but low-power standby mode (active-standby state 334). Alternatively, the dynamic configuration logic may directly transition the computing resource to the non-standby active state 336. In either case, de-darkening the computing resource from the dark state 332 may include the dynamic configuration logic removing the power gate® on the computing resource.

[0033] A variety of different activation guarantee operating conditions are possible, and no specific activation guarantee operating conditions are required. Some examples of activation guarantee operating conditions include dynamic configuration logic that detects cryogenic operation of the microprocessor 100, dynamic configuration logic that receives or obtains commands to activate computing resources (e.g., via instruction set architecture (ISA) commands or via memory-mapped I / O), or dynamic configuration logic that detects configuration guarantee operating conditions that guarantee the use of computing resources.

[0034] The dynamic configuration logic detects configuration-guaranteed operating conditions that guarantee the use of computing resources, and if the computing resources are in a dark state 332, the dynamic configuration logic can treat the configuration-guaranteed operating conditions as an activation-guaranteed operating system in order to directly transition the computing resources from the dark state 332 to a non-standby active state 336, or to first transition the computing resources from the dark state 332 to an active-standby state 334, and then to a non-standby active state 336.

[0035] On the other hand, if the configuration guarantee operating conditions are detected by the dynamic configuration logic and the computing resource is already in the active-standby state 334, the dynamic configuration logic can transition the computing resource from the active-standby state 334 to the non-standby-active state 336. This transition can be achieved by the dynamic configuration logic removing the clock gate on the computing resource, or by the dynamic configuration logic using voltage scaling techniques to transition the computing resource from a dormant state to an active state.

[0036] The dynamic configuration logic can, upon detecting configuration guarantee operating conditions, configure the microprocessor 100 so that it no longer uses the computing resources. In this case, the computing resources can transition back to the active-standby state 334. For this transition from the non-standby-active state 336 to the active-standby state 334, the dynamic configuration logic can clock-gate the computing resources or use voltage scaling techniques to transition them from the active state to the sleep state.

[0037] Alternatively, if the dynamic configuration logic detects a deactivation guarantee condition for a computing resource that is in a non-standby active state 336, the dynamic configuration logic may power gate the computing resource to directly transition it to a dark state 332.

[0038] When a computing resource is in the active-standby state 334, the dynamic configuration logic may clock-gate the computing resource to conserve power. Clock gating refers to a technique in a microprocessor to reduce power consumption by a computing resource when it is not in use. Clock gating can be implemented in the microprocessor 100, for example, by removing the clock signal from the computing resource when it is not in use. When a physical computing resource is clock-gated, power is supplied to the circuit, but the clock pulses that drive the circuit are blocked. By doing so, energy consumed due to circuit switching is reduced or eliminated, but leakage power is still consumed. In contrast to clock gating, power gating cuts off the power signal, and therefore current, to the circuit. Power gating typically requires a transition between power states, which is usually a physically longer process (a longer latency process) than enabling and disabling clock pulses to the circuit (clock gating). Therefore, clock gating can be used to more efficiently transition the physical computing resource between the active-standby state 334 and the non-standby state 336.

[0039] Clock gating may be used to transfer physical computing resources between the active-standby state 334 and the non-standby-active state 336. However, other techniques, such as the voltage scaling techniques referenced above for transferring physical computing resources between the active state and the sleep state, may be used to transfer physical computing resources between these states. In this case, the sleep voltage scaling state corresponds to the active-standby state 334, and the active voltage scaling state corresponds to the non-standby-active state 336.

[0040] According to some embodiments, the dynamic configuration logic is a programmable epoch-based system that periodically checks activation guarantee conditions, deactivation guarantee conditions, or configuration guarantee conditions for computing resources. For example, the dynamic configuration logic may check one or more of these conditions every few clock cycles or every few nanoseconds. The periodicity of these checks may also change over time depending on the dynamic configuration logic that detects conditions that guarantee an increase or decrease in the frequency of these checks. The check frequency may also be controlled from high-level logic, for example, by high-level language program instructions or high-level language compiler append instructions to the instruction set executed by the microprocessor 100.

[0041] (Hybrid in-order / out-of-order CPU design) According to some embodiments, the CPU (e.g., 104-1) of the core (e.g., 102-1) of the over-provisioned microprocessor 100 encompasses a hybrid in-order / out-of-order CPU design. In particular, the microprocessor 100 is over-provisioned in both in-order and out-of-order computational resources, and the CPU's dynamic configuration logic dynamically configures the CPU to be either an in-order instruction execution machine or an out-of-order instruction execution machine.

[0042] Generally, when the CPU is in in-order instruction execution mode, computational application instructions are fetched, executed, and committed in compiler-generated order. If an instruction stalls (for example, while waiting for data from main memory), all subsequent instructions will also stall. Instructions are statistically scheduled by the CPU in compiler-generated order. The advantages of in-order instruction execution include simpler implementations, faster clock cycles, fewer computational resources, and lower design, development, and debugging costs.

[0043] On the other hand, when the CPU is in out-of-order instruction execution mode, computational application instructions can still be fetched in compiler-generated order. However, instruction completion can be in-order or out-of-order. Instructions are dynamically scheduled by the CPU. The CPU determines in which order instructions can be executed, and instructions following a stalled instruction can be passed in execution order if they do not depend on the stalled instruction. The advantages of out-of-order execution include higher performance for certain computational workloads with a high level of instruction-level parallelism and low instruction dependency. Other advantages potentially include latency hiding, fewer processor stalls, and higher utilization of execution (function) units.

[0044] When using in-order instruction execution computing resources, out-of-order instruction execution computing resources may be dormant. Alternatively, when using out-of-order instruction execution computing resources, in-order instruction execution computing resources may be dormant. If operating conditions allow for being within the target TDP at cryogenic operation or lower (underclocked) clock frequencies, both in-order and out-of-order instruction execution computing resources may remain active while one of them is in use. In this case, the dynamic configuration between using in-order instruction execution computing resources and out-of-order instruction execution computing resources does not incur the overhead of transitioning computing resources from dormant to active. For example, power gating overhead is avoided.

[0045] The dynamic configuration of a CPU between an in-order instruction execution machine and an out-of-order instruction execution machine may include the dynamic configuration of control and data paths, as well as data storage structures such as instruction queues, rename tables, and reorder buffers. When the CPU is running a computational application with high inherent instruction-level parallelism and relatively low data dependency, this may be a configuration-guaranteed operating condition that triggers the dynamic configuration logic to configure the CPU as an out-of-order instruction execution machine in order to take advantage of the speculation and dynamism provided by out-of-order instruction execution operation. However, when this speculation and dynamism is no longer required by the computational application, this may be a configuration-guaranteed operating condition that triggers the dynamic configuration logic to dynamically configure the CPU as an in-order instruction execution machine in order to avoid the overhead of out-of-order instruction execution mode.

[0046] As described above, when the CPU is in either in-order instruction execution mode or out-of-order instruction execution mode, power can be saved by de-darkening alternative computing resources such as certain data and control paths and data storage structures not used in the current mode. However, if leakage current is reduced, for example, in cryogenic operation or at underclocked clock frequencies, clock gating or other low-power conditions may be used with minimal power overhead to keep currently unused computing resources active. In this way, if dynamic configuration logic decides to dynamically configure the CPU to switch from in-order instruction execution mode to out-of-order instruction execution mode or vice versa, this can be done quickly without the need to de-darken computing resources.

[0047] Figure 4 is a schematic diagram of a hybrid in-order / out-of-order CPU 104-1 of the core 102-1 of the microprocessor 100, according to several embodiments. The CPU 104-1 has an instruction fetch unit 438 that fetches the next instruction from a memory address stored in the program counter and stores the fetched instruction in an instruction register. The CPU 104-1 also has an instruction decode unit 440 for interpreting the fetched instruction. The configurable issue unit 442 can issue the decoded instruction to the in-order instruction execution unit 446 or the out-of-order instruction execution unit 448, depending on the current configuration by the dynamic configuration logic 444.

[0048] When the dynamic configuration logic 444 detects a configuration guarantee operating condition, it can dynamically configure the configurable issue unit 442 to issue instructions to the in-order instruction execution unit 448 for in-order instruction execution, or to the out-of-order instruction execution unit 446 for out-of-order instruction execution. For example, the dynamic configuration logic 444 may track the number of instruction execution stalls during an epoch of several clock cycles or several nanoseconds. If the number of stalls during an epoch exceeds a threshold and the configurable issue unit 442 is currently in in-order instruction execution mode, the dynamic configuration logic 444 can dynamically configure the configurable issue unit 442 to issue instructions to the 448 for out-of-order instruction execution until the dynamic configuration logic 444 detects a configuration guarantee operating condition that guarantees switching to in-order instruction execution mode. For example, the dynamic configuration logic 444 may detect an instruction passed from the instruction decode unit 440 requesting in-order instruction execution. Such instructions may be inserted into a computation application or added to the set of computation application instructions by a compiler or runtime instruction profiler based on the expectation that the subsequent instructions of the computation application being executed do not have a high degree of instruction-level parallelism, and therefore the overhead of out-of-order instruction execution is not guaranteed for these subsequent instructions.

[0049] For example, a high-level programming language compiler may, during an optimization or profiling pass, determine that the instruction schedule generated by the compiler for a given instruction window has little to no register dependencies. In this case, the compiler can insert instructions or otherwise configure the compiled instructions to select in-order instruction execution to execute the instruction window. On the other hand, if the compiler generates a complex instruction schedule with many register dependencies (e.g., a number of register dependencies exceeding a threshold), the compiler can insert instructions or otherwise configure the compiled instructions to select out-of-order instruction execution to execute the instruction window.

[0050] The CPU 104-1 may include other computing resources associated with the in-order instruction execution unit 446 and the out-of-order instruction execution unit 448. Such other computing resources may include an instruction queue 450, a reorder buffer 452, a rename table 454, a function unit 456, and a register file 458.

[0051] In CPU 104-1, the computing resources supported by in-order instruction execution 446 and out-of-order instruction execution 448 share front-end logic such as the instruction fetch unit 438, the decode instruction unit 440, the configurable issue unit 442, and the dynamic configuration logic 444. The in-order instruction execution unit 446 and the out-of-order instruction execution unit 448 may each have their own private computing resources. For example, only the out-of-order instruction execution unit 448 is likely to perform register renaming using the rename table 454 and the reorder buffer 452. Alternatively, the in-order instruction execution unit 446 and the out-of-order instruction execution unit 448 may use their own respective register files 458 (for example, to increase port width), or these units may share the register file 458.

[0052] The long control span 460 is an example of a long control span that may only be usable under specific operating conditions, such as cryogenic operation (lower electrical resistance) or lower (underclocked) clock frequencies. The long control span 460 allows for the transfer of register state information between the rename table 454 and the in-order instruction execution unit 446 when switching from out-of-order instruction execution to in-order instruction execution. The register state information may include, for example, logical register identifiers in the rename table 454 that determine whether the in-order instruction execution unit 446 uses physical register identifiers for in-order instruction execution.

[0053] (Configurable memory-level parallelism) According to some embodiments, the load-store (LS) unit of the CPU (e.g., 104-1) of the core (e.g., 102-1) of the microprocessor 100 for holding memory instructions may be over-provisioned with additional queue entries that can be activated (de-darkened) under certain operating conditions, such as cryogenic operation (lower electrical resistance) or lower (underclocked) clock frequencies. The extra capacity / width in the queue for storing memory requests (e.g., load and store) allows the microprocessor 100 to provide extended memory-level parallelism.

[0054] Additionally or alternatively, the CPU may over-provision (1) simple load-store (LS) units and (2) related load-store (LS) units. Simple LS units can be used for computational applications or parts of computational applications that require in-order memory access, have limited memory-level parallelism due to computational limits, have data dependencies, are difficult to predict branches, or have unpredictable memory access patterns. On the other hand, related LS units with more extended memory scheduling and arbitration capabilities can be used for computational applications or parts of computational applications that require a high degree of memory-level parallelism.

[0055] If the CPU is using a simple LS unit computing resource, the associated LS unit computing resource may be dimmed. Alternatively, if the CPU is using an associated LS unit computing resource, the simple LS unit computing resource may be dimmed. If operating conditions allow for operation within the target TDP, such as at cryogenic temperatures or at lower (underclocked) clock frequencies, both the simple LS unit computing resource and the associated LS unit computing resource may remain active while one of them is in use. In this case, configuring between the simple LS unit computing resource and the associated LS unit computing resource does not incur the overhead of transitioning the computing resource from dimmed to active. For example, power gating overhead is avoided.

[0056] The CPU may include dynamic configuration logic to dynamically configure the CPU between the use of simple LS units and the use of related LS units. If the CPU is running a computational application with high inherent memory-level parallelism and relatively low data dependencies, this may be a configuration-guaranteed operating condition that triggers the dynamic configuration logic to configure the CPU to switch from the use of simple LS units to the use of related LS units. However, if the scheduling and arbitration of related LS units are no longer required by the computational application, this may also be a configuration-guaranteed operating condition that triggers the dynamic configuration logic to dynamically configure the CPU to use simple LS units.

[0057] As mentioned above, when the CPU is using either a simple LS unit or an associated LS unit, power can be saved by de-darkening alternative computing resources such as certain data and control paths and data storage structures that are not being used by the current mode. However, if leakage current is reduced, for example in cryogenic operation, clock gating or other low-power conditions may be used with minimal power overhead to keep currently unused computing resources active. In this way, if the dynamic configuration logic decides to dynamically configure the CPU to switch from using a simple LS unit to using an associated LS unit or vice versa, this can be done quickly without the need to de-darken computing resources.

[0058] Figure 5 is a schematic diagram of a configurable load / store (LS) unit in the CPU 104-1 of the core 102-1 of a microprocessor 100, according to several embodiments. The CPU 104-1 has a configurable memory request issuing unit 562 that receives requests for memory access (e.g., load and store). The configurable issue unit 562 is dynamically configured by dynamic configuration logic 562 to issue memory requests to a simple LS unit 566 or an associated LS unit 568, depending on the current configuration. The associated LS unit 568 may include extended memory scheduling or arbitration logic. Such logic can be configured to accommodate multiple consistency models or various memory sequences, including store / load transfers.

[0059] For example, when the dynamic configuration logic 562 detects configuration guarantee operating conditions, it can dynamically configure the configurable issuer 562 to issue memory access requests to either a simple LS unit 566 or an associated LS unit 568. For example, the dynamic configuration logic 562 can track the utilization of a functional unit over time. If a functional unit is not currently being fully utilized and the configurable issuer 562 is currently in associated LS mode, it is likely that there is insufficient memory-level parallelism in the instruction being executed. Therefore, the dynamic configuration logic 564 can dynamically configure the configurable issuer 562 to switch to using a simple LS unit 566, thereby avoiding the energy and computational overhead of the associated LS unit 568.

[0060] Other possible configuration-guaranteed operating conditions for switching between the simple LS unit 566 and the associated LS unit 568 include compiler instructions, directives, or hints inserted into instructions executed by the high-level programming language compiler. For example, if a compiled instruction has a high proportion of memory instructions (e.g., with respect to memory operations per total operation within the window of compiled instructions), the compiler may insert a directive or hint into the compiled code to indicate that a larger load-store queue will be used for the instruction window. Conversely, if the instruction window has a low rate of memory instructions, the directive or hint may indicate that a load-store queue of normal size or smaller will be used for the instruction window.

[0061] If the compiler can successfully resolve the memory instruction addresses and determines that there are many independent addresses within the instruction window, the compiler may instruct or imply that associated LS unit 568 be used for the window to increase the memory-level parallelism of the window. If the compiler has difficulty resolving memory addresses for the instruction window at compile time, when the window is executed, it may choose a simple LS unit 568 mode to avoid inefficient use of the sorting or transfer logic of associated LS unit 566.

[0062] Address prediction techniques or runtime address profiling that use performance counters to track observed addresses can be utilized by dynamic configuration logic. If there is a phase of high memory-level parallelism detected during the execution of an application program, this can be detected by dynamic configuration logic, and a transition from using the simple LS unit 566 to the associated LS unit 568 can be made. Conversely, if the microprocessor 100 is performing insufficient work utilizing the complex logic in the associated LS unit 568, a transition to using the simple LS unit 566 can be made by dynamic configuration logic. Furthermore, when the dynamic configuration logic detects configuration assurance operating conditions, it can dynamically configure the queues or other data storage structures of the simple LS unit 566 and the associated LS unit 568 to activate or use extra over-provisioned entries. For example, when the dynamic configuration logic detects that the microprocessor 100 is operating at cryogenic temperatures or is underclocked, it can activate and dynamically configure the simple LS566 or associated LS unit 568 to use extra over-provisioned entries to provide a wider width (greater possible memory-level parallelism) for memory access requests. When the dynamic configuration logic detects that the microprocessor 100 is no longer operating at cryogenic temperatures or is no longer underclocked, it can dynamically configure the simple LS556 or associated LS unit 568 to no longer use the additional over-provisioned entries.

[0063] (Configurable simultaneous multithreading processor) Computational applications can be programmed to run in both single-threaded and multi-threaded modes. During execution, the computational application may switch back and forth between single-threaded and multi-threaded execution. According to some embodiments, a single CPU design can be over-provisioned with pipelined computing resources to optimize both single-threaded and multi-threaded execution.

[0064] Figure 6 is a schematic diagram of a configurable simultaneous multithreading CPU 104-1 of a microprocessor 100 core 102-1 in several embodiments. The CPU 104-1 has multiple hardware instruction pipelines to support multiple concurrently executing threads. Each pipeline has its own instruction fetch unit 670, its own instruction decode unit 672, its own configurable instruction issue unit 674, its own instruction queue 676, and its own set of function units 678. Furthermore, the CPU 104-1 is overprovisioned from each configurable instruction issue unit 674 to one or more other hardware pipelines by long control spans. In the example in Figure 6, each configurable instruction issue unit 674 is connected to each other's hardware pipelines by long control spans. However, a configurable instruction unit 674 may be connected to fewer pipelines than all other pipelines by long control spans.

[0065] As a result, when CPU 104-1 is executing a single-threaded computation application or a single-threaded portion of a computation application in one of its pipelines, the front-end computation resources of the other pipelines may be put into a dark state or an active but low-power state (e.g., clock-gated) when the front-end computation resources are not being used in order to conserve power. The front-end computation resources include an instruction fetch unit 670, an instruction decode unit 672, and a configurable instruction issue unit 674.

[0066] Furthermore, in this situation, the configurable instruction issuing unit 674 of the pipeline used to execute a single thread can issue instructions to the backend of an unused pipeline. This can be done so that the functional unit 678 of another pipeline can execute the single-threaded instructions. Instructions can be issued to the backend of the other pipeline over a long control span, which may be available at cryogenic operation or underclocked clock speeds.

[0067] For example, when pipeline 0 is executing a single thread, the front-end units 670-0, 672-0, 674-0 and back-end units 676-0, 678-0 of pipeline 0 can be used to fetch, decode, issue, and execute single-threaded instructions. During this time, the front-end units 670-1, 672-01, 674-1 of pipeline 1 and the front-end units 670-2, 670-2, 674-2 of pipeline 2 may be implicitly deactivated or in a low-power state as they are not being used to conserve power. However, the back-end units 676-1, 678-1, 676-2, 678-2 may be used during this time. In particular, when configuration assurance operating conditions are detected indicating that the functional unit 678-0 of pipeline 0 is about to be fully utilized or is fully utilized, the configurable instruction issuing unit 674-0 may be dynamically configured to begin issuing instructions to the back-ends of other pipelines (e.g., pipeline 1 or pipeline 2) over a long control span. In this way, when pipeline 0 executes a single-threaded application or a part thereof, it can utilize the functional units 678-1 or 678-2 of other pipelines in addition to its own set of functional units 678-0. This increases the overall instruction throughput of the single-threaded application.

[0068] If a computational application causes thread generation for multiple execution threads, the front-end of another pipeline (e.g., pipeline 1 or pipeline 2) can be used so that multiple threads can run concurrently on separate pipelines. In this case, the configurable instruction issuing unit 674-0 can be dynamically configured so that it no longer issues instructions to the back-ends of the other pipelines (e.g., pipeline 1 or pipeline 2) over a long control span, as those back-ends are used by threads running on those pipelines.

[0069] Note that the configurable simultaneous multithreading CPU 104-1 in Figure 6 can be combined with the hybrid in-order / out-of-order CPU 104-1 in Figure 4. In this configuration, each pipeline can be over-provisioned in terms of both in-order and out-of-order instruction execution units and associated computing resources. The configurable instruction issuer 674 can switch between in-order and out-of-order instruction issuing within and across their respective pipelines when in single-thread mode. For example, the instruction execution throughput of instructions in a single-threaded application with a low level of instruction-level parallelism can be improved by using in-order instruction execution units from other pipelines in addition to the in-order instruction execution units in the pipeline where the single thread is running.

[0070] (Configurable register file size) According to some embodiments, the register file of the CPU (e.g., 104-1) of the core (e.g., 102-1) of the microprocessor 100 is over-provisioned with additional entries that can be activated (de-darkened) under specific operating conditions such as cryogenic operation (lower electrical resistance) or lower (underclocked) clock frequencies. The extra capacity / width in the register file allows the microprocessor 100 to increase its utilization of other computing resources. For example, in addition to the over-provisioned register file, the microprocessor 100 may be over-provisioned with more function units, function unit schedulers, instruction queue entries, branch predictor table entries, prefetcher table entries, reorder buffer entries, on-chip cache entries, load-store queue entries, cacheway predictor entries, opcode cache entries, or fetch buffer entries. The size (data capacity) of latches, flip-flips, SRAM structures, CAM structures, etc. can also be increased.

[0071] (Configurable instruction window size) According to some embodiments, a configurable instruction window size is provided. Here, the data storage structure of the CPU (e.g., 104-1) or core (102-1) of the microprocessor 100 implementing the instruction window size may be over-provisioned with extra entries that can be activated (de-darkened) under certain operating conditions, such as cryogenic operation (lower electrical resistance) or lower (underclocked) clock frequencies. Such data store structures may include register files, rename tables, reorder buffers, etc. The instruction window size can be increased by using these extra entries. If an increased instruction window size is not required, these extra over-provisioned entries can be de-darkened or otherwise put into a low-power state (e.g., clock-gated). Whether to use extra entries to increase the instruction window size may depend on the level of instruction-level parallelism in the instructions of the computational application or the portion of the computational application being executed by the CPU. If there is a high level of instruction-level parallelism in the instructions, extra entries can be used. If there is only a low or intermediate level of instruction-level parallelism in the instructions, extra entries can be de-darkened or not used.

[0072] (Programmed logic support) The dynamic configuration logic may be programmed or configured to periodically check activation guarantees, deactivation guarantees, and configuration guarantee operating conditions according to a programmable or configurable epoch. The epoch may be programmed or configured, for example, in terms of the number of clock cycles or nanoseconds. Furthermore, the dynamic configuration logic may maintain counters and other data to determine whether or not operating conditions are met, and when they will be met. Such counters may include the number of entries in the data storage structure being used, the memory instruction rate, the cache miss rate, the bandwidth usage, etc. Based on these counters and data, the dynamic configuration logic may determine that operating conditions have been met and, accordingly, transition computing resources between a dark state, an active-standby state, and a non-standby-active state.

[0073] Compiler-based configuration can also be supported. For example, when compiling program instructions for a computational application programmed in a high-level programming language such as C or C++, the compiler can add instruction hints to the compiled instructions to use extra over-provisioned entries in the register file, based on the compiler's knowledge of register file usage by the program instructions obtained during the compilation of the program instructions into compiled instructions (machine code). At runtime, the instruction hints can be detected by dynamic configuration logic, and the extra over-provisioned entries in the register file are used when the compiled instructions are executed. The hints may apply to all compiled instructions in the computational application, or to only some of them.

[0074] Operating system configurations may also be supported. For example, an operating system running a single-threaded computation application can detect that the computation application has triggered multiple threads, for example by issuing a system call to the operating system, and then issue a system call or other low-level call to the CPU. The dynamic configuration logic can detect the system call or low-level call from the operating system and dynamically configure the CPU pipeline from single-threaded mode to concurrent multi-threaded mode, for example, as described above with respect to Figure 6.

[0075] (Conclusion) Thus, a dynamically configurable over-provisioned microprocessor is disclosed. The microprocessor may be over-provisioned with computing resources such as alternative computing resources, long control spans, or data storage structures with extra headroom. Dynamic configuration logic can detect operating conditions that ensure the transition of computing resources between dark, low-power, and active states to balance computing throughput and energy usage. In some embodiments, certain over-provisioned computing resources, such as long control spans, are used only under specific operating conditions, such as cryogenic operation or underclocking, while other over-provisioned computing resources may still be used at the target clock frequency in room temperature operation by darkening currently unused computing resources through power gating, or by bringing them into a low-power state through clock gating. For example, extra over-provisioned headroom for a register file may be used at the target clock frequency in room temperature operation by darkening or clock gating other unused computing resources to reduce power density and remain within the target TDP.

[0076] (Other aspects of this disclosure) Unless otherwise clearly stated in the context, the term “or” is used in its inclusive sense (rather than its exclusive sense) within the above specification and the attached claims, and therefore, for example, when used to connect a list of elements, the term “or” means one, some, or all of the elements in the list.

[0077] Unless otherwise explicitly stated in the context, terms such as “comprising,” “including,” “having,” “based on,” and “encompassing” are used in an open-ended manner in the above specification and the attached claims and do not exclude additional elements, features, actions, or operations.

[0078] Unless otherwise explicitly stated in the context, conjunctions such as "at least one of X, Y, and Z" should be understood to indicate that an item, term, etc., may be X, Y, Z, or a combination thereof. Therefore, such conjunctions are not intended to imply that a particular embodiment requires the presence of at least one of X, at least one of Y, and at least one of Z, respectively.

[0079] Unless otherwise explicitly stated in the context, the above detailed explanation is intended to include the plural forms of the singular "a," "an," and "the."

[0080] Unless otherwise explicitly stated in the context, the terms "first," "second," etc., are used herein to describe various elements, as may be, but these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first computing device may be called a second computing device, and similarly, a second computing device may be called a first computing device. Both the first and second computing devices are computing devices, but they are not the same computing device.

[0081] In the above specification, several embodiments are described with reference to numerous specific details that may vary from implementation to implementation. Therefore, this specification and the drawings should be considered illustrative rather than restrictive. The sole exclusive indicator of the scope of the invention, and what the applicant intends to be the scope of the invention, is the literal equivalent scope of the set of claims issued from this application in any particular form in which such claims are issued, including any subsequent modifications.

Claims

1. An over-provisioned microprocessor, The first physical computing resource, Equipped with dynamic configuration logic, The aforementioned dynamic configuration logic is To detect the activation guarantee operating conditions, In response to the detection of the activation guarantee operating conditions, the first physical computing resource is de-darkened, To detect the configuration guarantee operating conditions, In response to detecting the aforementioned configuration assurance operating conditions, the over-provisioned microprocessor is dynamically configured to use the de-darkened first physical computing resources, It is configured to do, An over-provisioned microprocessor.

2. The aforementioned dynamic configuration logic is To detect the cryogenic operation of the over-provisioned microprocessor, In response to detecting cryogenic operation of the over-provisioned microprocessor, the first physical computing resource is de-darkened, It is configured to do, An over-provisioned microprocessor according to claim 1.

3. The second physical computing resource, The system further comprises a control span connecting the first physical computing resource to the second physical computing resource, An over-provisioned microprocessor according to claim 1.

4. An in-order instruction execution pipeline comprising a first set of physical computing resources, Further comprising an out-of-order instruction execution pipeline having a second set of computing resources, The first physical computing resource is one of the first set of physical computing resources, The dynamic configuration logic is configured to dynamically configure the overprovisioned microprocessor to use the in-order instruction execution pipeline to execute instructions for the computation application, and no longer use the out-of-order instruction execution pipeline to execute instructions for the computation application, in response to detecting the configuration guarantee operating conditions. An over-provisioned microprocessor according to claim 1.

5. An in-order instruction execution pipeline comprising a first set of physical computing resources, Further comprising an out-of-order instruction execution pipeline having a second set of physical computing resources, The first physical computing resource is one of the second set of physical computing resources, The dynamic configuration logic is configured to dynamically configure the overprovisioned microprocessor to use the out-of-order instruction execution pipeline to execute instructions for the computation application, and no longer use the in-order instruction execution pipeline to execute instructions for the computation application, in response to the detection of the configuration guarantee operating conditions. An over-provisioned microprocessor according to claim 1.

6. A first load-store unit for in-order memory access, comprising a first set of physical computing resources, It further comprises a second load-store unit for parallel memory access, which has a second set of computing resources, The first physical computing resource is one of the first set of physical computing resources, The dynamic configuration logic is configured to dynamically configure the over-provisioned microprocessor to use the first load-store unit to execute memory access requests and no longer use the second load-store unit to execute memory access requests, in response to the detection of the configuration assurance operating conditions. An over-provisioned microprocessor according to claim 1.

7. A first load-store unit for in-order memory access, comprising a first set of physical computing resources, It further comprises a second load-store unit for parallel memory access, which has a second set of physical computing resources, The first physical computing resource is one of the second set of physical computing resources, The dynamic configuration logic is configured to dynamically configure the over-provisioned microprocessor to use the second load-store unit to execute memory access requests and no longer use the first load-store unit to execute memory access requests, in response to the detection of the configuration guarantee operating conditions. An over-provisioned microprocessor according to claim 1.

8. A plurality of pipelines, each of which has a set of front-end computing resources and a set of back-end computing resources, each set of front-end computing resources comprises an instruction fetch unit, and each set of back-end computing resources comprises a set of functional units, The system further comprises at least one control span that connects the computing resources of each set of front-end computing resources of a first pipeline among the plurality of pipelines to the computing resources of each set of back-end computing resources of a second pipeline among the plurality of pipelines. An over-provisioned microprocessor according to claim 1.

9. The system further comprises a configurable instruction issuing unit for each set of front-end computing resources of the first pipeline, configured to send instructions for a single-threaded computing application running on the first pipeline to the computing resources of each set of back-end computing resources of the second pipeline via the control span. An over-provisioned microprocessor according to claim 8.

10. The first physical computing resource is one or more over-provisioned entries in a data storage structure. An over-provisioned microprocessor according to claim 1.

11. The aforementioned data storage structure is a queue, register file, rename table, or reorder buffer. An over-provisioned microprocessor according to claim 10.

12. An over-provisioned microprocessor core, The first physical computing resource, Equipped with dynamic configuration logic, The aforementioned dynamic configuration logic is To detect the activation guarantee operating conditions, In response to the detection of the activation guarantee operating conditions, the first physical computing resource is de-darkened, To detect the configuration guarantee operating conditions, In response to the detection of the aforementioned configuration guarantee operating conditions, the first physical computing resource is dynamically transitioned from an active standby power state to a non-standby power state. It is configured to do, Over-provisioned microprocessor cores.

13. The aforementioned dynamic configuration logic is To detect the cryogenic operation of the over-provisioned microprocessor core, In response to detecting cryogenic operation of the over-provisioned microprocessor core, the first physical computing resource is de-darkened, It is configured to do, An over-provisioned microprocessor core according to claim 12.

Citation Information

Patent Citations

  • Cross bar switch

    JP1990001086A

  • Over-provisioned multicore processor

    US20090094438A1

  • Directed Resource Folding for Power Management

    US20120096293A1

  • Apparatus and method for adaptive guard-band reduction

    WO2015094373A1