Platform Efficiency Tracker

A system management circuit dynamically tracks power efficiency and reallocates resources in computing systems, addressing inefficiencies from conservative power loss assumptions, enabling improved performance by optimizing power usage.

JP2025524450APending Publication Date: 2025-07-30ADVANCED MICRO DEVICES INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024575324
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-06-29
Filing Date
2023-06-01
Publication Date
2025-07-30

AI Technical Summary

Technical Problem

Existing computing systems inefficiently allocate power due to conservative assumptions about power loss, leading to performance limitations despite having more power available than assumed.

Method used

Implement a system management circuit that dynamically tracks power efficiency and reallocates power within the system based on real-time conditions, using a model to estimate power loss and adjust voltage and frequency to improve performance without exceeding power consumption limits.

Benefits of technology

Enhances system performance by utilizing available power more efficiently, allowing components to operate at higher performance states when conditions permit, thereby optimizing power usage and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025524450000001_ABST
    Figure 2025524450000001_ABST
Patent Text Reader

Abstract

A system, apparatus, and method for dynamically estimating power losses in a computing system are disclosed. A system management circuit tracks the state of the computing system and dynamically estimates power losses in the computing system based in part on that state. Based on the estimated power losses, the power consumption of the computing system is estimated. The system management circuit is configured to increase the power performance state of one or more circuits of the computing system while staying within the power allocation limit of the computing system in response to detecting a reduction in power losses in at least a portion of the computing system.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] (Description of Related Art) During the design of a computer or other processor-based system, many design factors must be considered. To succeed in the design, various trade-offs may be required between power consumption, performance, heat output, etc. For example, in the design of a computer system that emphasizes high performance, it may allow for greater power consumption and heat output. Conversely, in the design of a portable computer system that is sometimes powered by a battery, reducing power consumption may be emphasized even at the expense of some performance. Whatever the specific design goal, a computing system typically has a predetermined amount of power available during operation. This power must be allocated among various components within the system, such as partly to the central processing circuit, another part to the memory subsystem, and a part to the graphics processing circuit, etc. Also, how the power is allocated among system components may change during operation.

[0002] Computers and other complex electronic systems are typically designed with a thermal budget and a power budget. Power consumption and heat output must be maintained within certain limits. However, the maximum values of power consumption and heat output are parameters that must be considered in the context of system performance requirements. Placing too much emphasis on power requirements and thermal requirements in the design of a system may prevent the performance goals from being achieved. Conversely, placing too much emphasis on performance may result in exceeding the power and thermal goals. Since variations in the required processing load can lead to wide variations in power consumption and heat output, many processors have the ability to adjust the operating voltage and operating clock frequency. This can enable control over power consumption and heat output and may enable these parameters to meet the design requirements.

[0003] When a computing system is designed, the total amount of power required and available is determined. For example, a central processing circuit is determined to have a certain range of power requirements, a memory subsystem is determined to have a certain power requirement, and so on. Then, the computing system power requirements are determined based on the requirements of all the components of the system. Generally speaking, the components within a computing system have inefficiencies that result in power losses. For example, voltage regulators, circuit board wiring, fans, and other components within the system are not perfectly efficient with respect to power consumption. To account for such inefficiencies, assumptions are made during design regarding how much power loss exists and how much power is actually available. For example, during design, a voltage regulator within the system may be determined to be 80 - 95% efficient when operating under various conditions. Since the system designer must consider all conditions, a conservative assumption is made that the voltage regulator is 80% efficient, and the total available circuit board power is determined based on this assumption. By making conservative assumptions, it is ensured that adequate power is always available. However, during actual system operation, the efficiency is not always 80%. Sometimes, the efficiency is higher, and actually more power may be available than assumed. Therefore, power that could otherwise be used to improve the performance of the system is not used, and the performance is unnecessarily limited.

[0004] The advantages of the methods and mechanisms described herein can be better understood by reference to the following description in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0006] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the methods and mechanisms presented herein. However, one of ordinary skill in the art should recognize that various embodiments may be practiced without these specific details. In some instances, well-known structures, components, signals, computer program instructions, and techniques have not been shown in detail in order to avoid obscuring the approach described herein. For the sake of brevity and clarity, it should be understood that the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements.

[0007] A system, apparatus, and method are disclosed for tracking power efficiency in a computing system and changing a power performance state. In this context, power efficiency refers to the ratio of the total power drawn from a power source that is actually usable by the SoC. The computing system includes a system management circuit that estimates the power efficiency of one or more components of the system based on various system conditions. In one embodiment, the system management circuit allocates power to components within the system based on determined system requirements. In one embodiment, a maximum available power budget is allocated to a given component, and that given component is required to operate within that range. During operation of the computing system, various conditions are monitored. In response to detecting a first condition, it is determined that a given component is operating with improved power efficiency, and the power performance state of the given component is improved. In various embodiments, the power performance of a given component is improved without increasing the estimated power consumption. In this way, improved performance is obtained while maintaining a given estimated power consumption level.

[0008] In one embodiment, the estimated power consumption by a given component is determined based in part on a pre-determined characterization and current operating conditions. Such characterization can be performed either pre-silicon or post-silicon. Such operating conditions can include one or more of operating temperature, operating frequency, current being drawn, and other conditions. In one embodiment, the power efficiency tracking circuit is configured to generate an estimate of power consumption based on dynamic calculations using the operating conditions and / or other parameters described above. In some embodiments, the dynamic calculations are performed based on equations implemented in hardware (i.e., circuitry). In other embodiments, a combination of hardware and software is used to perform the calculations.

[0009] Referring to FIG. 1, a block diagram of one embodiment of a computing system 100 is shown. In this example, a power supply 104 coupled to a substrate 102 including components of the system is shown. In this embodiment, the power supply 104 represents the total amount of power available to the substrate 102 and the components of the system components within the system. In this embodiment, the illustrated computing system 100 includes a system on chip (SoC) 105 coupled to a memory 160. However, embodiments are possible and contemplated in which one or more of the illustrated components of the SoC 105 are not integrated on a single chip. In some embodiments, the SoC 105 includes a plurality of processor cores 110A-110N and a GPU 140. In the illustrated embodiment, the SoC 105, the memory 160, and other components (not shown) are part of a system board 102 (e.g., a mother board), and one or more of the peripheral devices 150A-150N and the GPU 140 are separate entities (e.g., daughter boards, etc.) coupled to the system board 102. In other embodiments, one or more of the GPU 140 and / or the peripheral devices 150 may be permanently mounted on the substrate 102 or otherwise integrated into the SoC 105. Note that the processor cores 110A-110N can also be referred to as processing circuits or processors. The processor cores 110A-110N and the GPU 140 are configured to execute instructions of one or more instruction set architectures (ISAs) that can include operating system instructions and user application instructions. These instructions include memory access instructions that can be converted and / or decoded into memory access requests or memory access operations targeting the memory 160.

[0010] In another embodiment, SoC 105 includes a single processor core 110. In a multi-core embodiment, the processor cores 110 may be identical to each other (i.e., symmetric multi-core) or one or more cores may be different from the others (i.e., asymmetric multi-core). Each processor core 110 includes one or more execution circuits, cache memory, a scheduler, a branch prediction circuit, etc. Further, each of the processor cores 110 is configured to assert an access request to the memory 160 that functions as the main memory of the computing system 100. Such requests include read requests and / or write requests and are first received from the respective processor cores 110 by the bridge 120. Also, each processor core 110 can include a queue or buffer that holds inflight instructions that have not yet completed execution. This queue can be referred to herein as an "instruction queue". Some of the instructions within the processor core 110 may still be waiting for their operands to become available, while other instructions may be waiting for an available arithmetic logic unit (ALU). Instructions waiting on an available ALU can be referred to as pending ready instructions. In one embodiment, each processor core 110 is configured to track the number of pending ready instructions.

[0011] The input / output memory management circuit (IOMMU) 135 is coupled to the bridge 120 in the illustrated implementation. In one embodiment, in the computing system 100, the bridge 120 functions as a north bridge device and the IOMMU 135 functions as a south bridge device. In other embodiments, the bridge 120 can be a fabric, a switch, a bridge, any combination of those components, or another component. Several different types of peripheral buses (e.g., peripheral component interconnect (PCI) bus, PCI-Extended (PCI-X), PCIE (PCI Express), gigabit Ethernet (GBE) bus, universal serial bus (USB)) may be coupled to the IOMMU 135. Various types of peripheral devices 150A - 150N can be coupled to some or all of the peripheral buses. Such peripheral devices 150A - 150N include, but are not limited to, keyboards, mice, printers, scanners, joysticks, other types of game controllers, media recording devices, external storage devices, network interface cards, etc. At least some of the peripheral devices 150A - 150N coupled to the IOMMU 135 via the corresponding peripheral buses can assert memory access requests using direct memory access (DMA). These requests (which can include read requests and write requests) are transmitted to the bridge 120 via the IOMMU 135.

[0012] In some embodiments, SoC 105 includes a graphics processing circuit (GPU) 140 configured to be coupled to a display 145 (not shown) of computing system 100. In some embodiments, GPU 140 is a separate integrated circuit separate from SoC 105. GPU 140 performs various video processing functions and provides the processed information to display 145 for output as visual information. Also, GPU 140 can be configured to perform other types of tasks scheduled for GPU 140 by an application scheduler. GPU 140 includes "N" computing circuits for executing tasks of various applications or processes, where "N" is a positive integer. Also, the "N" computing circuits of GPU 140 may be referred to as "processing circuits". Each computing circuit of GPU 140 is configured to assert a request for access to memory 160.

[0013] In one embodiment, memory controller 130 is integrated with bridge 120. In other embodiments, memory controller 130 is separate from bridge 120. Memory controller 130 receives memory requests transmitted from bridge 120. Data accessed from memory 160 in response to a read request is transmitted by memory controller 130 to the requesting agent via bridge 120. In response to a write request, memory controller 130 receives both the request and the data to be written from the requesting agent via bridge 120. If multiple memory access requests are pending at a given time, memory controller 130 arbitrates between those requests. For example, if the power budget allocated to memory controller 130 limits the total number of requests that can be made to memory 160, memory controller 130 can give priority to critical requests while delaying non-critical requests.

[0014] In some embodiments, memory 160 includes a plurality of memory modules. Each of the memory modules includes one or more memory devices (e.g., memory chips) mounted thereon. In some embodiments, memory 160 includes one or more memory devices mounted on a motherboard or other carrier on which SoC 105 is also mounted. In some embodiments, at least a portion of memory 160 is implemented on the die of SoC 105 itself. Embodiments having combinations of the above-described embodiments are also possible and contemplated. In one embodiment, memory 160 is used to implement a random access memory (RAM) for use with SoC 105 during operation. The implemented RAM may be static RAM (SRAM) or dynamic RAM (DRAM). Types of DRAM used to implement memory 160 include, but are not limited to, double data rate (DDR) DRAM, DDR2 DRAM, DDR3 DRAM, etc.

[0015] Although not explicitly shown in FIG. 1, SoC 105 can also include one or more cache memories inside processor core 110. For example, each of processor cores 110 can include an L1 data cache and an L1 instruction cache. In some embodiments, SoC 105 includes a shared cache 115 shared by processor cores 110. In some embodiments, shared cache 115 is a level two (L2) cache. In some embodiments, an L2 cache is implemented within each of processor cores 110, and thus shared cache 115 is a level three (L3) cache. Cache 115 may be part of a cache subsystem that includes a cache controller.

[0016] In one embodiment, the system management circuit 125 is integrated into the bridge 120. In other embodiments, the system management circuit 125 may be separated from the bridge 120 and / or the system management circuit 125 may be implemented as multiple individual components in multiple locations of the SoC 105. The system management circuit 125 is configured to manage the power states of the various processing circuits of the SoC 105. The system management circuit 125 may also be referred to as a power management circuit. In one embodiment, the system management circuit 125 uses dynamic voltage and frequency scaling (DVFS) to change the frequency and / or voltage of the processing circuits in order to limit the power consumption of the processing circuits to a selected power allocation.

[0017] The SoC 105 includes a plurality of temperature sensors 170A-170N representing any number of temperature sensors. The sensors 170A-170N are shown on the left side of the block diagram of the SoC 105, but it should be understood that the sensors 170A-170N can be distributed throughout the SoC 105 and / or located adjacent to the major components of the SoC 105 in the actual implementation of the SoC 105. In one embodiment, there are sensors 170A-170N for each core 110A-110N, the computing circuits of the GPU 140, and other major components. In this embodiment, each sensor 170A-170N tracks the temperature of the corresponding component. In another embodiment, there are sensors 170A-170N for different geographical regions of the SoC 105. In this embodiment, the sensors 170A-170N are distributed throughout the SoC 105 to track the temperature within various areas of the SoC 105 and monitor whether there are hot spots within the SoC 105. In other embodiments, other ways of arranging the sensors 170A-170N within the SoC 105 are possible and contemplated.

[0018] Also, the SoC 105 includes a plurality of performance counters 175A - 175N that represent any number and type of performance counters. Although the performance counters 175A - 175N are shown on the left side of the block diagram of the SoC 105, it should be understood that the performance counters 175A - 175N can be distributed throughout the SoC 105 and / or can be located within the major components of the SoC 105 in the actual implementation of the SoC 105. For example, in one embodiment, each of the cores 110A - 110N includes one or more of the performance counters 175A - 175N, the memory controller 130 includes one or more of the performance counters 175A - 175N, the GPU 140 includes one or more of the performance counters 175A - 175N, and the other performance counters 175A - 175N are utilized to monitor the performance of other components. The performance counters 175A - 175N can track various different performance metrics, including the instruction execution rate, the consumed memory bandwidth, the row buffer hit rate, the cache hit rate of various caches (e.g., instruction cache, data cache), and / or other metrics, of the cores 110A - 110N and the GPU 140.

[0019] In one embodiment, SoC 105 includes a phase-locked loop (PLL) circuit 155 coupled to receive a system clock signal. The PLL circuit 155 includes a number of PLLs configured to generate corresponding clock signals and distribute them to each of the processor cores 110 and other components of the SoC 105. In one embodiment, the clock signals received by each of the processor cores 110 are independent of each other. Further, the PLL circuit 155 in this embodiment is configured to individually control and change the frequency of each of the clock signals provided to each of the processor cores 110 independently of each other. The frequency of the clock signal received by any given processor core among the processor cores 110 can increase or decrease according to the power state allocated by the system management circuit 125. The various frequencies at which the clock signals are output from the PLL circuit 155 correspond to different operating points of each of the processor cores 110. Therefore, changing the operating point of a particular processor core among the processor cores 110 is implemented by changing the frequency of the clock signal received by each of them.

[0020] For the purposes of the present disclosure, the operating point can be defined as the clock frequency and can also include the operating voltage (e.g., the supply voltage provided to the functional circuit). Raising the operating point of a given functional circuit can be defined as raising the frequency of the clock signal provided to that circuit and can also include increasing its operating voltage. Similarly, lowering the operating point of a given functional circuit can be defined as lowering the clock frequency and can also include lowering the operating voltage. Limiting the operating point can be defined as limiting the clock frequency and / or the operating voltage to the maximum value specified for a particular set of conditions (however, not necessarily the maximum for all conditions). Therefore, if the operating point is limited for a particular processing circuit, that processing circuit can operate at the clock frequency and operating voltage up to the value specified for the current set of conditions, but can also operate at clock frequencies and operating voltage values lower than the specified value.

[0021] If changing the operating point of each of one or more processor cores 110 includes changing one or more respective clock frequencies, the system management circuit 125 changes the state of the digital signals provided to the PLL circuit 155. In response to these signal changes, the PLL circuit 155 changes the clock frequency of the affected processor core 110. Additionally, the system management circuit 125 can also cause the PLL circuit 155 to inhibit each clock signal from being provided to the corresponding processor core among the processor cores 110.

[0022] In the illustrated embodiment, SoC 105 includes a plurality of voltage regulators (VRs) 165A - 165M included on substrate 102. Each of them is coupled to one or more components within the system and provides a predetermined voltage. In other embodiments, voltage regulator 165 may be implemented separately from SoC 105. In various embodiments, power supply 104 represents a power supply that establishes the maximum amount of power available to substrate / platform 102. A portion of the power supplied by power supply 104 is actually available for use by SoC 105 as usable power, but a portion is lost. Power losses occur in various ways within the system. For example, power is lost in the transmission of power from power supply 104 to voltage regulator 165. For example, losses occur in the signal wiring of substrate 102 when power is transmitted from one location to another. Similarly, power losses occur within voltage regulator 165. As is known to those skilled in the art, voltage regulator 165 is not perfectly efficient and does not convert power completely efficiently. Each of the components of SoC 105 is likewise not perfectly efficient in its power usage, and some power loss occurs during operation. More generally, a portion of the maximum amount of power that power supply 104 can supply is consumed by the SoC and other components of system 100, and the remainder of the power is consumed in the form of platform / power supply losses. As a result, a portion of the power supplied by power supply 104 is lost. Voltage regulators 165 provide supply voltages to each of processor cores 110 and other components of SoC 105. In some embodiments, voltage regulator 165 provides a supply voltage that can be changed according to a specific operating point. In some embodiments, each of processor cores 110 shares a voltage plane. Thus, each processor core 110 in such an embodiment operates at the same voltage as the other processor cores 110. In another embodiment, the voltage planes are not shared, and thus the supply voltage received by each processor core 110 is set and adjusted independently of the respective supply voltages received by the other cores of processor core 110.Accordingly, the adjustment of the operating point, including the adjustment of the supply voltage, can be selectively applied to each processor core 110 independently of other cores in an implementation having non-shared voltage planes. When changing the operating point includes changing the operating voltage of one or more processor cores 110, the system management circuit 125 can change the state of the digital signal provided to the voltage regulator 165. In response to the change in the signal, the voltage regulator 165 adjusts the supply voltage provided to the affected cores among the processor cores 110. When power is removed from a processor core among the processor cores 110 (i.e., a closed processor core), the system management circuit 125 sets the state of the corresponding signal among the signals so as not to supply power to the processor core 110 affected by the voltage regulator 165.

[0023] In various embodiments, the computing system 100 may be any of a computer, laptop, mobile device, server, web server, cloud computing server, memory system, or various other types of computing systems or devices. Note that the number of components of the computing system 100 and / or the SoC 105 may vary from embodiment to embodiment. There may be more or fewer each component / sub-component than shown in FIG. 1. Also note that the computing system 100 and / or the SoC 105 may include other components not shown in FIG. 1. Additionally, in other embodiments, the computing system 100 and the SoC 105 may be structured in other ways than shown in FIG. 1.

[0024] Next, referring to FIG. 2, a block diagram of a portion of the substrate 102 of FIG. 1 coupled to the power supply 104 of FIG. 1 is shown. In the illustrated example, an embodiment of the system management circuit 210 is shown. The system management circuit 210 is coupled to the computing circuits 205A-205N, the memory controller 225, the phase-locked loop (PLL) circuit 230, and the voltage regulator 235A. As shown, the power supply 104 is coupled to supply power to a plurality of voltage regulators 235A-235C and other components 240 (not shown) on the substrate. In this example, the power supply 104 is shown to supply power to the voltage regulator 235A coupled to the computing circuit 205, the voltage regulator 235B coupled to the system management circuit 210, and the voltage regulator 224 coupled to the memory controller 225. As can be understood, many other voltage regulators and system components are coupled to receive power supplied by the power supply 104, and the diagram of FIG. 2 is merely illustrative. Also, the system management circuit 210 may be coupled to one or more other components not shown in FIG. 2. The computing circuits 205A-205N represent any number and type of computing circuits, and the computing circuits 205A-205N may also be referred to as processors or processing circuits. For example, in one embodiment, at least one computing circuit is a CPU and another computing circuit is a GPU.

[0025] The system management circuit 210 includes an efficiency tracking circuit 202, a power allocation circuit 215, and a power management circuit 220. The efficiency tracking circuit 202 is configured to dynamically track and estimate the power efficiency of various components within the system. By dynamically tracking the power efficiency (or power loss), the tracking circuit 202 can dynamically estimate the power consumption. It should be noted that the total power consumption is composed of the power consumed by all components of the platform, including the SoC, other components, and other elements of the platform (such as the power distribution network, etc.). In this context, the power efficiency or power loss being tracked generally corresponds to the entire platform and is not generally tracked as part of the power estimation / tracking of the SoC components. Based on the estimated value of the power consumed by the SOC and the estimated value of the power loss in the board / platform, the total power drawn from the power supply can be estimated. The power allocation circuit 215 is configured to allocate power budgets to each of the calculation circuits 205A - 205N, the memory subsystem including the memory controller 225, and / or one or more other components. The total amount of power available to the power allocation circuit 215, which will be distributed among the components, can have an upper limit for the host system or device. The power allocation circuit 215 receives various inputs from the calculation circuits 205A - 205N, including the status of the miss status holding register (MSHR) of the calculation circuits 205A - 205N, the instruction execution rate of the calculation circuits 205A - 205N, the number of execution-ready instructions waiting in the calculation circuits 205A - 205N, the instruction and data cache hit rates of the calculation circuits 205A - 205N, the consumed memory bandwidth, and / or one or more other input signals. The power allocation circuit 215 can use these inputs to determine whether the calculation circuits 205A - 205N have tasks to execute, and then the power allocation circuit 215 can adjust the power budgets allocated to the calculation circuits 205A - 205N according to those determinations.Also, the power allocation circuit 215 can receive inputs from the memory controller 225, and those inputs include the consumed memory bandwidth, the total number of requests in the pending request queue, the number of critical requests in the pending request queue, the number of non-critical requests in the pending request queue, and / or one or more other input signals. The power allocation circuit 215 can utilize the status of these inputs to determine the power budget allocated to the memory subsystem.

[0026] The PLL circuit 230 includes any number of PLLs configured to receive a system clock signal, generate corresponding clock signals, and distribute them to each of the arithmetic circuits 205A to 205N and other components. The power management circuit 220 is configured to transmit a control signal to the PLL circuit 230 to control the clock frequencies supplied to the arithmetic circuits 205A to 205N and other components. The voltage regulator 235 provides supply voltages to each of the arithmetic circuits 205A to 205N and other components. The power management circuit 220 is configured to transmit a control signal to the voltage regulator 235 to control the voltages supplied to the arithmetic circuits 205A to 205N and other components.

[0027] Memory controller 225 is configured to control the memory (not shown) of the host computing system or device. For example, memory controller 225 issues read, write, erase, refresh, and various other commands to the memory. In one embodiment, memory controller 225 includes the components of memory controller 225 (of FIG. 2). When memory controller 225 receives a power budget from system management circuit 210, memory controller 225 converts the power budget into the number of memory requests per second that memory controller 225 is permitted to make to the memory. The number of memory requests per second is enforced by memory controller 225 to ensure that memory controller 225 stays within the power budget allocated to the memory subsystem by system management circuit 210. Also, the number of memory requests per second can take into account the status of the DRAM so that memory controller 225 can issue pending critical and non-critical requests to currently open DRAM rows as long as certain memory power constraints are met. Memory controller 225 prioritizes processing critical requests without exceeding the number of requests per second that memory controller 225 is permitted to make. If all critical requests are processed and memory controller 225 has not reached the specified per-second request limit, memory controller 225 processes non-critical requests.

[0028] Referring to FIG. 3, a sample chart is presented showing how the power efficiency of the platform can vary according to various operating conditions. In the example shown, the y-axis lists the estimated power range loss, loss (W) (e.g., 0 to 400 watts), and the x-axis shows the range of current, I out (e.g., 0 to 800 amperes) drawn by the platform. As shown, during operation, the power loss increases as the current draw increases. In various embodiments, the loss can vary based on various conditions, including for example temperature. For example, I outWhen I = 100 A, the power loss is approximately 10 W. On the other hand, when I out = 300 W, the power loss is approximately 70 W. Furthermore, it can be seen that this relationship is non-linear. In the illustrated example, the relationship between current and power loss is a quadratic curve. Therefore, the proportion of power loss in the platform increases as the current increases.

[0029] As described above, when a computing system is designed, the designer must consider the inefficiencies that result in power loss when determining how much power is allocated to the various components within the system. Designers and vendors of the various components (e.g., processors, memory controllers, memory, I / O circuits, etc.) used in such a computing system must determine the power requirements of the components under design. For example, when determining the power requirements of a GPU, a GPU designer makes assumptions about how much power is lost in the power supply network during operation. As an example, a GPU designer may determine that all components of the GPU require a total of up to 400 watts of power to operate. Furthermore, it may be determined that 7 - 15% of the power is lost in the GPU's power supply network during operation. Taking this estimated power loss into account, the GPU designer determines the actual power budget required to supply 400 watts of power to the components of the GPU and reports to the platform / substrate designer that the GPU requires an allocation of power in excess of 400 watts to account for these losses. For example, a GPU vendor may report that, assuming an efficiency of only 85%, approximately 400 / 0.85 = 471 watts are required by the GPU.

[0030] In the above example, the GPU vendor uses the upper part of the power loss range (15% loss) to ensure proper operation of the GPU for all operating conditions. In other words, the GPU vendor assumes the worst-case power loss, which is always statically assumed during operation. However, the actual power loss varies during operation, and the platform may actually operate more efficiently (with respect to power) than suggested by the 15% power loss assumption under various times and conditions. Therefore, if the GPU is consuming the maximum power based on an estimate assuming 85% efficiency, the GPU is restricted from further improving performance even if it is not actually consuming the maximum power.

[0031] To enable a component (the GPU used in this example) to better track power efficiency, in one embodiment, at least one power efficiency tracking circuit is incorporated into the system. FIG. 4 shows a system management circuit 410 that includes an efficiency tracking circuit 402, a power allocation circuit 415, and a power performance management circuit 440. Also shown is that the system management circuit 410 is configured to receive any number of various system parameters shown as 420A - 420Z corresponding to the conditions, operations, or states of the system. In the illustrated example, the parameters are shown to include the operating temperature 420A of a given circuit, the current draw 420B by a given circuit, and the operating frequency of a given circuit. Other parameters are possible and contemplated. In various embodiments, one or more of the parameters 420 are reported from other circuits or parts of the system (e.g., based on sensors, performance counters, other event / activity detections, etc.). In some embodiments, one or more parameters are tracked within the system management circuit 410. For example, the system management circuit 410 can track the current power performance state of components within the system, the duration of the power performance state, previously reported parameters, etc.

[0032] In the example of FIG. 4, the efficiency tracking circuit 402 includes a model 430. The model 430 is used to estimate the power efficiency of a circuit based on parameters 420. In one embodiment, the model 430 includes a circuit configured to perform calculations representing the relationship between power loss and current based on the parameters 420. In other embodiments, the model 430 may include a combination of hardware and software (e.g., firmware) for calculating an estimate. In various embodiments, the model 430 is developed based at least in part on an assessment of the characteristics of the operation of the circuit being tracked. For example, taking a GPU as an example, a designer may perform a number of tests to assess the power loss of a GPU in operation under a wide range of conditions. Such conditions include assessing the power loss according to the operating frequency, voltage, current, type of workload (e.g., high computational load vs. high memory load), circuits in operation, temperature, etc. Based on these assessments, a model representing the power loss based on these various conditions is developed. In one embodiment, an equation / function representing such a model (as described above) is generated to represent the power efficiency of the circuit, and a circuit implementing that function is designed. For example, in some embodiments, the power loss is represented using a simplified model represented by the equation P loss =c2×I out 2 +c1×I out +c0. In this example, c is a fixed coefficient that can be determined experimentally, through simulation, or by other means, and I out is the current. In other embodiments, the equation P loss =f2(v,temp)×I out 2 +f1(v,temp)×I outA more advanced model represented by +f0(v, temp) may be used. In this model, the coefficients are replaced with functions that depend on voltage and temperature. In other embodiments, a look-up table or other structure may be used to estimate the current power efficiency. During system operation, the system management circuit 410 monitors the estimated power loss based on the estimation of the efficiency tracking circuit 402. Based on the estimated power loss, the system management circuit 410 can change the power performance state of one or more circuits of the system. For example, in one embodiment, a circuit (e.g., a computing circuit) is configured to operate in multiple power performance states. Given a sufficient power budget, the computing circuit can operate in a higher power performance state and complete work at a faster speed. However, with a reduced power budget, the computing circuit may be limited to a lower power performance state, and as a result, the work is completed at a slower speed.

[0033] Referring to FIG. 5, one embodiment is shown of how the dynamically determined power loss can be used to improve system performance. Two processes, process 500 and process 520 are shown in FIG. 5. In one embodiment, processes 500 and 520 correspond to functions performed by a system management circuit (e.g., system management circuit 410) to track power efficiency and change the power performance state based on the dynamically determined power consumption estimate. For example, in one embodiment, the system management circuit 410 includes a circuit configured to perform the functions represented by processes 500 and 520. In some embodiments, processes 500 and 520 operate simultaneously, but this is not necessary.

[0034] The process 500 of FIG. 5 corresponds to the function of an efficiency tracking circuit that monitors system conditions (e.g., the parameters 420 of FIG. 4, etc.) and dynamically calculates an estimated power consumption. As described above, a model used to estimate power losses under various conditions can be incorporated into the circuit. As described above, conventionally, static assumptions are made regarding the always assumed power loss (e.g., 15%). In the illustrated embodiment, the process 500 is shown to include an initial power consumption estimate (PCE) 502. However, in other embodiments, the initial estimate may be calculated based on block 512 described below. The process 500 is configured to monitor and detect various conditions (e.g., conditions 504 and 506) and calculate a new power consumption estimate based on the dynamic detection and estimation of power losses (or power efficiency). Note that only two conditions are shown for simplicity of the figure. In an embodiment, in fact, any number of conditions can be monitored.

[0035] In the illustrated example, a determination is made as to whether a first condition has been detected (504). This condition takes into account the state of the system represented by parameters received by the system management circuit (e.g., circuit 410) or made available in some other way. Based on the state of the system, a determination is made as to whether a change in the estimated power loss has been detected (decision block 504). If a reduction in the estimated power loss (i.e., an increase in the estimated efficiency) is detected, the current estimated power consumption is decreased (508), and a new power consumption estimate (PCE) is generated or otherwise calculated. Instead, if no estimated decrease is detected (decision block 504), a determination is made as to whether an increase in the estimated power loss has been detected (decision block 506). If so, the estimated power consumption is increased (510), and based on that increase, a new estimated power consumption is generated or otherwise calculated (512). If neither condition is detected (504 or 506), monitoring continues until such a condition is detected or some event (e.g., reboot, power-off, power reduction state of the system management circuit, override signal to prevent monitoring, etc.) occurs.

[0036] As described above, processes 500 and 520 operate simultaneously in various embodiments. Regardless of whether both are operating simultaneously all the time or at various times, process 520 is configured to modify the power-performance state (PPS) of components within the system based on the estimated power consumption generated by process 500. In various embodiments, the estimated power consumption is generated by the efficiency tracking circuit 402 and made available to the power performance management circuit 440 configured to perform the functions indicated by process 520.

[0037] In one embodiment, process 520 is configured to compare the current estimated power consumption with the maximum power allocated to a given circuit. For example, in one embodiment, this given circuit is a GPU. However, in other embodiments, the circuit is considered only a part of the GPU (or other system component). As can be understood, the methods and mechanisms described herein are applicable to computing circuits at any of various granularities. For example, the power consumption of the entire GPU can be estimated and operated based thereon. Alternatively, specific computing circuits of the GPU can be tracked for power efficiency. All such embodiments are possible and contemplated.

[0038] As shown in FIG. 5, the current power consumption estimate (PCE) is compared with the maximum power allocated to the circuit. Using the entire GPU as an example, the platform of which the GPU is a part allocates 471 watts of power to the GPU as described above. Depending on the operating conditions, process 500 may generate a current estimated power consumption that is less than 471 watts due to the estimation of power loss reduction. If the estimate is lower than the maximum value (condition block 522), it may be possible to increase the PPS of the GPU, or a part of the GPU, while staying within the total power limit of the power supply. As shown, a determination is made as to whether a change in the power performance condition (PPS) is indicated (condition block 524). For example, if the current task benefits from a higher operating frequency of the computing circuit (e.g., a video game where a higher frame rate is desired), the power management circuit increases the PPS of the computing circuit. This may involve or otherwise cause a higher frequency and voltage to be applied to the computing circuit, which may consume more power. Also, changing the power performance state may include allocating more or less power budget to the circuit. After increasing the PPS of the computing circuit, a new estimate is generated by process 500.

[0039] Conversely, if it is determined that the PCE exceeds the maximum power assigned (decision block 528), the PPS of the circuit can be changed if such a change is indicated (conditional block 530). In some scenarios, the PCE may be permitted to exceed the maximum value for a limited period. Otherwise, a PPS change is indicated and the PPS can be decreased (block 532). Note that the circuit whose PPS is changed by process 520 need not be the circuit directly tracked for efficiency. For example, if one or more first circuits (e.g., circuit A) are operating more power-efficiently and a reduction in power loss is dynamically detected within the system, the power saved by circuit A can be allocated for use by one or more different circuits (e.g., circuit B). The system management circuit 410 can be configured to detect scenarios in which it is possible to make such an allocation. For example, models developed during characterization can identify such possibilities and various other combinations of system operations that can improve performance in one area when efficiency improves in other areas. Next, referring to FIG. 6, an example of a method 600 for changing the power performance state (PPS) of a circuit based on the estimated power consumption in a computing system is shown. This example shows how power reallocation can be changed based on dynamically tracking power efficiency in a computing system. Generally speaking, reallocation of the power budget in a system can reduce a portion X of the assigned power budget from one circuit and allocate that portion X to another circuit. However, this example shows how the above-described dynamic power consumption estimates can change this reallocation.

[0040] In this example, it is determined that the computational circuit is operating with reduced power loss, and an increase in the PPS of the memory subsystem is initiated (condition 605). The memory subsystem includes a memory controller and one or more memory devices. In one embodiment, the determination to increase the PPS of the memory subsystem is based in part on detecting the execution of a task that requires an increase in memory bandwidth (e.g., due to the type of workload), a critical memory access request pending, or other means. Depending on the embodiment, the system management circuit may utilize one or more of the number of tasks that one or more processors must execute, the current operating point of one or more processors, the consumed memory bandwidth, the number of critical and non-critical pending requests in the memory controller, the temperature of one or more components and / or the overall temperature of the system, and / or one or more other metrics for determining how much power should be allocated to the memory subsystem.

[0041] When condition (condition 605) is detected, the power budget allocated to the computing circuit is reduced by an amount X (block 615). This amount X may represent an amount that does not take into account the reduced power loss of the computing circuit. In other words, an algorithm for reallocating the power budget within a system that does not consider the efficiency tracking described above can be established. For example, the reallocation of the power budget can be performed within the system regardless of the reduction in power loss (e.g., due to changes in the workload, etc.). Based on this algorithm, it is determined that X is reallocated to the memory subsystem. Considering that the system management circuit 410 is configured to dynamically track the power loss and the corresponding increase in efficiency, the system management circuit 410 calculates the amount of the power budget to be transferred to the memory subsystem when the current power loss is considered. In the illustrated example, due to the reduced power loss, more than the amount X of the power budget reduced from the computing circuit can be allocated to the memory subsystem. Therefore, the amount of power X + Y is allocated to the memory subsystem, increasing the PPS and power consumption of the memory subsystem, and leveraging this newly allocated power (block 620). Note that the reverse can also occur. When a relatively high (higher) power loss is detected and a power budget reallocation condition is detected, the amount of the power budget to be reallocated can be reduced to account for the reduction in efficiency. A number of such scenarios are possible and contemplated.

[0042] In various embodiments, program instructions of a software application are used to implement the methods and / or mechanisms described previously. Those program instructions describe the behavior of hardware in a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog is used. Those program instructions are stored on a non-transitory computer-readable storage medium. A number of types of storage media are available. The storage medium is accessible by a computing system during use, and as a result, provides the program instructions and accompanying data to the computing system for program execution. The computing system includes at least one or more memories and one or more processors configured to execute the program instructions.

[0043] It should be emphasized that the above-described embodiments are merely non-limiting examples of embodiments. Once the above disclosure is fully understood, numerous variations and modifications will become apparent to those skilled in the art. The following claims are intended to be construed as encompassing all such variations and modifications.

Claims

1. A system comprising: a system management circuit, wherein the system management circuit is configured to: dynamically estimate the power consumption of a computing system; and change the power performance state of a first circuit among a plurality of circuits of the computing system in response to dynamically detecting either an increase or a decrease in power loss in the computing system. The system.

2. The system according to claim 1, wherein the system management circuit is configured to dynamically estimate the power consumption of each of the plurality of circuits of the computing system.

3. The system according to claim 2, wherein the system management circuit is configured to increase the power performance state of a second circuit among the plurality of circuits in response to detecting a decrease in power loss in the first circuit among the plurality of circuits.

4. The computing system is a graphics processing circuit, and the system management circuit is configured to increase the power performance state of a computing circuit in response to detecting a decrease in power loss in a memory subsystem. The system according to claim 2.

5. The system management circuit is configured to dynamically estimate the power consumption of the computing system based on the state of the computing system, wherein the state is partially based on one or more of a plurality of parameters including the amount of current drawn, the operating frequency, and the operating temperature. The system according to claim 2.

6. The system management circuit is configured to dynamically estimate the power consumption of the plurality of circuits based in part on calculations representing the relationship between power loss and one or more of current, voltage, and temperature. The system according to claim 1.

7. The system management circuit is configured to: determine that the computing system is estimated to be consuming the maximum amount of power allocated for use by the computing system; and increase the power consumption used by at least a part of the system without reducing the power consumed elsewhere in the system and without increasing the estimated power consumption in response to detecting a condition. The system according to claim 1.

8. ​ ​ ​ ​ The condition is to dynamically estimate a reduction in power loss in at least a part of the computing system. The system of claim 7. **Claim 9** A method comprising: dynamically estimating power consumption of a computing system; and changing a power performance state of a first circuit among a plurality of circuits of the computing system in response to dynamically detecting either an increase or a decrease in power loss in the computing system. The method. **Claim 10** including dynamically estimating power consumption of each of a plurality of circuits of the computing system. The method of claim 9. **Claim 11** including increasing a power performance state of a second circuit among the plurality of circuits in response to detecting a reduction in power loss in a first circuit among the plurality of circuits. The method of claim 9. **Claim 12** The computing system is a graphics processing circuit, and the method includes increasing a power performance state of a computing circuit in response to detecting a reduction in power loss in a memory subsystem. The method of claim 10. **Claim 13** including dynamically estimating power consumption of the computing system based on a state of the computing system, wherein the state is partially based on one or more of a plurality of parameters including an amount of current drawn, an operating frequency, and an operating temperature. The method of claim 10. **Claim 14** including dynamically estimating power consumption of the plurality of circuits based in part on calculations representing relationships between power loss, current, and one or more of voltage and temperature. The method of claim 9. **Claim 15** determining that the computing system is estimated to be consuming a maximum amount of power allocated for use by the computing system; and increasing power consumption used by at least a part of the computing system without reducing power consumed elsewhere in the computing system and without increasing the estimated power consumption in response to detecting a condition. The method of claim 9. **Claim 16** The condition is to dynamically estimate a reduction in power loss in at least a part of the computing system. The method of claim 15. **Claim 17** A system comprising: A system comprising a plurality of circuits including a central processing circuit, a graphics processing circuit, a memory subsystem, and a system management circuit. The system management circuit: Dynamically estimates the power consumption of the plurality of circuits; Changes the power performance state of a first circuit among the plurality of circuits in response to dynamically detecting either an increase or a decrease in power loss in the system; Is configured to perform the above. System.

18. The system management circuit is configured to increase the power performance state of a second circuit among the plurality of circuits in response to detecting a decrease in power loss in a first circuit among the plurality of circuits. The system of Claim 17.

19. The system management circuit is configured to dynamically estimate the power consumption of the plurality of circuits based at least in part on calculations representing the relationship between power loss and current. The system of Claim 17.

20. The system management circuit: Determines that the computing system is estimated to be consuming the maximum amount of power allocated for use by the computing system; Increases the power consumption used by at least a part of the system without reducing the power consumed elsewhere in the system and without increasing the estimated power consumption in response to detecting a condition. Is configured to perform the above. The system of Claim 17.