Power management of processing units

By introducing a distributed power-aware management algorithm into the processing platform, allowing each processing unit to independently perform power management, solving the problem of frequent suppression performance and insufficient energy efficiency in the prior art based on utilization rates, achieving more refined and efficient power management.

CN120010646APending Publication Date: 2025-05-16TAHOE RES LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510150064.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2017-12-15
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing processing platforms reactively suppress processor performance levels to prevent system failures when detection power limits are exceeded, but this approach may frequently affect processing performance and allocation frequency based on utilization alone may not be energy efficient enough.

Method used

By introducing distributed power-aware management algorithms into the processing platform, each processing unit (such as CPU, GPU, IPU) allows each processing unit (such as CPU, GPU, IPU) to perform power management independently, selecting frequency based on the unit's power limiting and scalability, ensuring optimized power efficiency and performance under different workload situations.

Benefits of technology

A more refined and efficient power management is achieved, reducing the impact of processing performance due to frequent suppression performance, and improving the reliability and energy efficiency of the system when facing power limits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010646A_ABST
    Figure CN120010646A_ABST
Patent Text Reader

Abstract

Power management circuitry is provided to control performance levels of processing units of a processing platform. The power management circuit comprises: a measurement circuit for measuring a current utilization rate of the processing unit at a current operating frequency and determining any change in utilization rate or power, and a frequency control circuit is provided for updating the current operating frequency to a new operating frequency by determining a new target quantized power loss to be applied in a subsequent processing cycle according to a determination of any change in the utilization or power. A new operating frequency is selected to meet the new target quantized power based on a scalability function that specifies a change in a given utilization or power value with the operating frequency. The processing platform and machine readable instructions are provided to set a new quantization target power for the processing unit.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application with a filing date of December 15, 2017, application number 201780095402.6, and invention name “Power Management of Processing Unit”. Technical Field

[0002] Embodiments described herein relate generally to the field of power management, and more particularly, to power management of processing units for controlling performance levels of a processing platform. Background Art

[0003] The processor has dynamically changing power requirements depending on the processing application (workload) requirements. For example, by selecting the execution frequency f and the corresponding processor core voltage V, many different performance states of the processor can be achieved. The processor power consumption P can be approximated as P=f*V 2 *C + leakage, where C is capacitance. Leakage is approximated as a constant corresponding to the power wasted as a result of applying voltage to the transistor. Thus, the processor frequency and voltage can be increased when the processing workload is high to run faster and this results in increased power consumption, while the processor frequency and voltage can be reduced to reduce power consumption when the processor has a low workload or is idle. The processor performance level can be set based on both the processing workload and the maximum capacity of the available power supply.

[0004] On current processing platforms, active power management is performed by dynamically scaling at least one of voltage and frequency, a technique known as dynamic voltage and frequency scaling (DVFS). DVFS can be performed when the processor requires a higher (or lower) performance state, and DVFS can be based on changes in processor utilization. A higher performance state (higher frequency state) is often granted by the DVFS controller unless there is some other constraint or limitation that is mitigated for the higher frequency selection, such as a thermal violation or peak current violation detected during processing.

[0005] As processing platforms evolve, the form factors of integrated circuits such as systems on chips (SOCs) are shrinking into more power-constrained and thermally constrained designs. Current platforms tend to detect that power limits are exceeded or approached, and respond by reactively suppressing processor performance levels to bring the platform back to the desired operating state. Performing this suppression may adversely affect processing performance if it is performed too frequently. In some cases, the reactive response to the power limit being breached may not provide enough warning to enable the processing platform to reliably prevent unexpected system failures. In addition, allocating frequencies to processors based solely on utilization may not be energy efficient for all processing tasks, for example, when the processing speed is reduced due to the waiting time for accessing data in memory. There are some situations when it may be appropriate to be more tolerant when allocating higher frequencies as utilization levels require, and there are other situations when it may be appropriate to be more conservative when allocating frequencies to be more energy efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The embodiments described herein are illustrated by way of example and not by way of limitation in the figures of the accompanying drawings in which like reference numerals refer to similar elements:

[0007] Figure 1 schematically illustrates a processing platform having a platform level power management circuit and a SOC level power management circuit;

[0008] Figure 2A schematically illustrates a software stack suitable for power management distributed among multiple processing units;

[0009] Figure 2B schematically illustrates example architectural counters for tracking central processing unit (CPU) performance;

[0010] Figure 3 schematically illustrates processing units of a SOC and their corresponding architectural counters;

[0011] Figure 4 schematically illustrates a functional flow for a power-aware frequency selection algorithm for a processing unit;

[0012] Figure 5 is a flow chart schematically illustrating a power-aware management algorithm for a processing unit;

[0013] Fig. 6A is a three-dimensional graph schematically illustrating an example of a power profile of a CPU, the three-dimensional graph depicting the relationship between frequency, utilization, and power consumption for a specific example workload;

[0014] Figure 6Bis a two-dimensional graph of frequency versus utilization of a processing unit, the two-dimensional graph depicting equidistant power lines and scalability lines for the processing units; and

[0015] Figure 7 is a table listing a sequence of processing elements used to implement power management for a processing unit. DETAILED DESCRIPTION

[0016] Illustrative embodiments of the present disclosure include, but are not limited to, methods, systems, and apparatus, and machine-readable instructions for peak power determination in an integrated circuit device.

[0017] Figure 1 A processing platform 100 is schematically illustrated having a platform-level power management circuit 110 and an SOC-level power management circuit 122 capable of controlling power to one or more individual processing units of the SOC 120. The processing platform 100 includes a power selector 102, a charger 104, a battery 106, a fuel gauge 108, the platform-level power management circuit 110, and one or more system power sensors 112. A system bus 114 provides a communication path between the platform-level power balancer circuit 110 and the SOC 120 and a rest of the platform (ROP) 130 representing a set of system components other than the SOC 120 to which the platform-level power balancer 110 controls power from available power sources.

[0018] In this example, ROP 130 includes memory 132, display 134, modem 136, and camera 138. SOC 120 includes a set of SOC power management circuits 122, one or more SOC-based sensors 123, SOC bus 124, a central processing unit (CPU) 125 including one or more physical processing cores (not shown), an image processing unit (IPU) 127, and a graphics processing unit (GPU) 129.

[0019] The IPU 127 may be an imaging subsystem for processing camera images from a camera that is integrated into an electronic device. The IPU 127 may include a sensor and camera control chipset and an image signal processor. The IPU 127 may support image capture, video capture, face detection, and image processing capabilities. The GPU 129 is a dedicated highly parallel processing unit that uses frame buffers and memory manipulation to process blocks of data to accelerate computer graphics and image processing. The GPU 129 may include multiple execution units (EUs), each of which includes at least one floating point unit (FPU).

[0020] The power selector 102 may select the AC adapter 96, which performs alternating current (AC) to direct current (DC) conversion on power from the mains AC power source 96 to supply power to the processing platform 100. However, when the mains power is not available, the power selector 102 may select the battery 106 of the processing platform 100 to supply power to the SOC 120 and the ROP 130. If the battery 106 is not fully charged, the AC adapter may be controlled by the power selector 102 to supply power to increase the charge level of the battery 106 and to supply power to the SOC 130 and the ROP 130.

[0021] The fuel gauge 108 may be used to determine at least the battery charge and the battery current, for example, by using a coulomb counter or a sense resistor to monitor the total amount of charge supplied to the battery during a charging cycle or received from the battery during a discharging cycle. The fuel gauge may provide an indication of at least one of the battery charge level and the full battery capacity in units of measure such as coulombs or ampere hours. The full battery capacity may decrease during the life of the battery due to the effects of multiple charge-discharge cycles. The fuel gauge may therefore provide an indication of the peak capacity of the battery 106 at a given time, which may depend on the calibration of the battery and the battery charge level at a given time.

[0022] The platform power management circuitry 110 may be coupled to one or more sensors 112 to receive information indicative of a status of the one or more sensors 112. The sensors 112 may be disposed proximate to a component of the system, such as proximate to the battery 106 or the power selector 102 or proximate to an interconnect such as the bus 114 or any other component associated with the processing platform 100. The SOC sensor 123 may be disposed proximate to one or more of the CPU 125, the IPU 127, the GPU 129, and the bus 124. The sensors 112, 123 may provide, for example, measurements of a battery charge level, a battery current value, an adapter current value, a temperature, an operating voltage, an operating current, an operating power, inter-core communication activity, an operating frequency, or any other parameter related to power or thermal management of the processing platform 100.

[0023] Figure 1 The SOC 120 may receive a system power signal P from the platform level power balancing circuit 110 via the system bus 114. SYS 111 and can use P SYS Signal 111 is used to monitor when the system power meets one or more threshold conditions including a peak power threshold condition. The threshold condition can be related to the average power limit or the instantaneous power limit of the processing platform 100. The threshold condition can be related to V SYS The minimum value of P SYS The maximum value (or the corresponding current I SYS or voltage VSYS ) is related to, below V SYS The minimum value of P is likely to cause system failure. SYS The maximum value of the system failure is likely to occur. The threshold voltage V for SOC TH and the maximum permissible system power P MAX_SYS and the maximum allowable system current I MAX or voltage V MAX The power limit parameters may be set by the SOC power management circuit 122, for example, by software in the SOC based on information from the platform level power management circuit 110. In other examples, these power limit parameters for the SOC 120 may be set by an embedded controller (not shown) on the processing platform. TH can depend on the minimum system voltage V set once by the user MIN Therefore, V TH Probably higher than V MIN The triggering of the assertion of the Suppress signal may depend on the value of V TH ,I MAX and P SYS_MAX One or more of .

[0024] Additional power control signals may be sent to the SOC 120 via the system bus 114 based on feedback from at least one of the system sensor 112 and the SOC sensor 123, for example, according to thermal limits of the processing unit. For example, the SOC sensor 123 may feed back temperature and power consumption measurements from the processing units 125, 127, 129 to the platform power management circuit 110 via the bus 114. Power control to activate a power reduction feature of the SOC 250 may be performed by the SOC power management circuit 122 based on counting the number of times the system signal has satisfied a corresponding threshold condition.

[0025] The SOC 120 may respond to assertion of a suppression signal (not shown) to activate a power reduction feature of the SOC. For example, power consumption may be reduced in response to a suppression signal from one of the platform power management circuit 110 and the SOC power management circuit 122, the suppression signal resulting in a reduction in processor frequency within a predictable time window Δt1 in which the suppression signal is asserted. The power reduction feature is implemented reactively by the power management circuits 110, 122 in some examples to respond to a threshold being crossed, or preemptively by the power management circuits 110, 122 to prevent a threshold from being crossed. Some power thresholds may be related to average power indicators, whereas other power thresholds may be related to instantaneous power characteristics associated with "spikes" in processing activity.

[0026] The platform power management circuitry 110 may handle thresholds that apply to the platform as a whole, including the SOC 122 and the ROP 130. For example, the platform power management circuitry 110 may ensure that the total power drawn by the ROP 130 and the SOC 120 does not exceed the maximum system power P currently available from the selected power supply. SYS The available maximum power may change, for example, as the charge level of the battery 106 is depleted. The platform power management circuitry 110 may receive an input voltage 103 from the power selector, an indication of the battery charge level via a signal 109 from the fuel gauge 108, an indication of a user preference representing an EPP (Energy Performance Preference) indicating whether a power saving mode or a performance enhancing mode is currently selected, and indications of system events such as docking and undocking of the processing platform 100 from a docking station that connects it to the mains AC power source 98. The EPP may be provided by the operating system or may be a user programmable value.

[0027] The platform power management circuit 110 may also receive one or more weights to provide information about how the available system power P should be allocated between at least a subset of the different components 132, 134, 136, 138 of the different processing units 125, 127, 129 of the SOC and the ROP 130. SYS In some embodiments, processing platform 100 may include only SOC 120 without ROP 130 components. SYS , and the platform power management circuit 110 can also ensure that the processing platform 100 is properly thermally managed to maintain consistency with any power dissipation and junction temperature operating condition limits associated with the processing unit such as CPU 125. Operating the CPU 125 within the thermal limits and any other design limits can prevent accidental damage to the CPU 125 and other components of the platform 100. At least one of the platform power management circuit 110 and the SOC power management circuit 122 can implement one or more power limits, such as a first power limit PL1, to provide a threshold for the average power of the platform 100 that can be maintained indefinitely. The value of PL1 can be set to be below or close to the thermal design limit of the processing platform. The second power limit PL2 can be a higher power limit than PL1, which may last up to a limited duration, such as, for example, 100 seconds.

[0028] Many different performance parameters can be used to monitor the performance of a processing unit. For example, the "utilization" of a processing unit can refer to the proportion of all available processing cycles when the processing unit (which may include multiple physical cores) is in an active state rather than in a dormant state or in a power saving state or in a shut down state. Utilization is sometimes expressed as "load", but this "load" is different from the workload (the processing task including the instructions to be executed). The workload when executed by a given processing unit can result in a corresponding utilization and a corresponding scalability of the processing unit, where the scalability reflects the time taken to complete the execution of the workload, which is likely to vary according to stops, etc. The instantaneous measurement of the utilization and scalability of a given workload can vary during the execution time. In the dormant state, power consumption is reduced by suspending processing but retaining some power, so that "wake-up" is faster than it is from the power supply state. The processing unit can provide many different power states even in the active mode, but the active state refers to the processor state is fully turned on and the system clock is turned on. The active state is the normal state of the physical core when the code is being executed. For multi-threaded operation, if any thread in the processor core is active, the state of the core should be parsed as active. In the "sleep" state the processor state is not fully turned on, but the system clock is turned on. The processor performance level may be controlled via the operating system or via dedicated or special purpose hardware or using a combination of hardware and software. The platform power management circuitry 110 may take into account one or more of the following: processing workload requirements (e.g., the type of program application being executed), thermal limits of the processing hardware, maximum power, voltage, frequency, and current levels, and an active performance window requested by the operating system.

[0029] The "scalability" of a processing unit may refer to how the execution time of a given processing workload of the processing unit may vary with operating frequency. For example, a workload that results in many stalls may be less scalable than a workload that results in few stalls. For example, stalls may occur due to dependencies on data returned from memory. Thus utilization may provide a measure of when a processing unit is active, whereas scalability may provide a measure of useful (non-stall) work done while the processor is active. It should be appreciated that increasing the processing frequency when scalability is low is likely to result in less rate increase for the workload than when scalability is high. This is because stalls such as memory-related stalls are not ameliorated by increasing the frequency of the processing unit, since stall time is not an explicit function of the processing unit execution clock rate.

[0030] In previously known systems, the selection of a performance level for a processing unit, such as an operating frequency selection, is based on a system power threshold and may have considered processing unit utilization and processing unit scalability in order to allocate a new frequency when selecting a new performance level. However, the power impact of the new frequency selection is not evaluated before setting the new frequency, but the power may have been suppressed to a lower value in response to a suppression signal. In contrast, according to the present technology, the power impact of each frequency selection is evaluated by at least one of the platform power management circuit 110 and the SOC power management circuit 122 before the frequency is allocated to the corresponding processing unit. Therefore, according to an example embodiment, a separate power limit can be assigned to each of at least a subset of ROP components 132, 134, 136, 138 and SOC components 125, 127, 129. The unit-based power limit can be applied to each or at least a subset of the following: memory 132, display 134, modem 136, camera 138, CPU 125, IPU 127, and GPU 129. Some units may have more than one associated operating frequency that affects the performance level of the unit. For example, the IPU 127 may have an input subsystem frequency and a processing subsystem frequency that may be controlled separately. The CPU may include multiple physical cores. A unit-based power limit may be set based on utilization measurements of the processing unit. Each processing unit may have an associated weight for allocating system power between the multiple processing units. The unit-based power limit may be dynamically updated by the processing platform.

[0031] Power saving policies may be implemented by ROP 130 components to comply with any unit power limits received from platform power management circuitry 110. For example, memory 132 may be placed in a self-refresh state, display 134 may reduce memory reads when performing a display refresh, or backlight brightness may be adapted or refresh rate may be changed depending on the type of media being displayed.

[0032] In a multi-core processor, all active processor cores can share the same frequency and voltage by selecting, for example, the highest frequency performance state requested among all active cores as the frequency to be allocated. CPU 125 can have multiple performance operating points with associated frequency and voltage parameters. The operating point can be selected to prioritize power efficiency or performance based on the value of the EPP parameter. The frequency selection can be software controlled by writing to the CPU registers. In the case of CPU 125, the operating voltage can then be selected based on the selected frequency and the number of active physical cores. Due to the low latency of the transitions between performance states, multiple performance level transitions per second are possible for each processing unit.

[0033] The GPU 129 may have a driver for dynamically adjusting between performance states to maintain performance, power, and thermal constraints. The voltage of the GPU 129 may be adjusted downward to put it in a sleep state. The frame rate may be limited to reduce the load on the GPU 129 and allow it to run at a lower speed to save power. Thus, the GPU 129 may be controlled by the SOC power management circuit 122 to operate within the power limit assigned by the platform power management circuit 110 based on the allocation of system power PSYS among the processing units of the processing platform 100.

[0034] In some previously known systems, the allocation of operating frequencies to the CPU may have been controlled according to a DVFS algorithm that was arranged to monitor processor core utilization and scalability at regular intervals (e.g., approximately every millisecond) and apply an average to the actual measurements. Any other platform components that support DVFS, such as the IPU 127 and GPU 129, may have sent their frequency requests to the CPU power management algorithm. In previously known systems, although frequency requests could be made, processing units other than the CPU did not perform power-aware control for the corresponding processing units. This is because there has been a strong focus on management focused on performance without proper consideration of the power impact of performance-based tuning. The operating frequency for the CPU may have been selected based on observed utilization and scalability changes. CPU power and system power P may have been used. SYS The performance state is determined by considering the power from all platform components other than the CPU as static SOC power. A high frequency performance level or increased performance level called "boost frequency" may have been assigned based on utilization thresholds being exceeded. The system may have responded to higher than desired power consumption corresponding to, for example, a thermal warning being triggered or a peak current violation by reactively reducing power consumption by reducing the operating frequency of the processing unit to reduce power consumption to the desired level.

[0035] In such previously known systems, the power manager algorithm for the CPU may have been a centralized arbiter and controller for granting frequencies to other processing units such as IPU 127 and GPU 129. In contrast, according to the present technology, each processing unit supporting DVFS may have a unit-specific algorithm for performing power management to enable power-aware frequency allocation of the corresponding unit. In some example embodiments, only a subset of one or more of the CPU and other processing units comprising platform 100 may have separate power-aware management algorithms. Distributing power-aware management between two or more processing units is more scalable to different platform architectures than previously known CPU-centric power management methods. The distributed power-aware management of the example embodiments allows performance level preferences that distinguish between power efficiency and performance enhancement to be input to each processing unit 125, 127, 129, rather than just to the CPU 125. In addition, thermal limits and peak current, voltage, and power consumption limits may be provided to the control logic for each processing unit. This allows for more effective and efficient performance level selection and power efficiency.

[0036] In some examples, processing platform 100 may represent a suitable computing device, such as a computing tablet, mobile phone or smartphone, laptop computer, desktop computer, Internet of Things (IoT) device, server, set-top box, wireless-enabled e-reader, or the like.

[0037] Figure 2A A software stack suitable for power management distributed among multiple processing units is schematically illustrated. The software stack includes an application software layer 210, an operating system layer 220, and a platform component layer 230. The application software layer includes a first workload 212, a second workload 214, and a third workload 216. The first workload 212 may be a video editing application, the second workload 214 may be a deoxyribonucleic acid (DNA) sequencing application 214, and the third workload 216 may be a gaming application. Different processing workloads 212, 214, 216 may place different processing demands on the processing units of the platform component level 230, which include CPU hardware 242, GPU hardware 252, and IPU hardware 262. For example, the gaming application 216 is likely to cause the GPU hardware 252 to consume proportionally more power than the CPU hardware 242, whereas the DNA sequencing application 214 is likely to cause the CPU hardware 242 to consume more power than the GPU hardware 252.

[0038] According to the present technique, the power impact of candidate target frequency selections for each set of processing hardware 242, 252, 262 may be considered before allocating those frequencies. Different processing workloads 212, 214, 216 may result in different utilization levels and different scalability levels, and those levels may also vary for each set of hardware 242, 252, 262.

[0039] According to the present technology, the operating system level 220 of the software stack can have a platform-level power control algorithm 222, which can assign a per-component power limit to each of at least a subset of the CPU hardware 242, the GPU hardware 252, and the IPU hardware 262 via the bus 225. The per-component power limit can be determined by the platform-level power control algorithm 222 based on factors such as the available system power P. SYS The platform level power control algorithm 222 may be set by one or more constraints and may take into account one or more of the following: threshold temperature, threshold current, and threshold voltage of the platform hardware. The platform level power control algorithm 222 may also supply mode selection parameters that apply globally to platform components or to a subset of components or to individual components to select between optimization (or at least improvement) of processing performance (throughput) or power efficiency.

[0040] At the platform component level 230 of the software stack, the CPU hardware 242 has a corresponding CPU performance level selection algorithm 244, which has an interface 246 to the platform level power control algorithm 222. The CPU performance level selection algorithm 244 takes as input a predicted CPU power profile 248, which it uses to make predictions about power usage for different candidate frequencies before allocating frequencies to the CPU hardware 242. Similarly, the GPU hardware 252 has a corresponding GPU performance level selection algorithm 254, which has an operating system (OS) GPU interface 256 to the platform level power control algorithm 222. The GPU performance level selection algorithm 254 takes as input a predicted GPU power profile 248, which it uses to make predictions about power usage for different candidate frequencies before allocating frequencies to the GPU hardware 252. Likewise, the IPU hardware 262 has a corresponding IPU performance level selection algorithm 254, which has an operating system (OS) IPU interface 266 to the platform level power control algorithm 222. The IPU performance level selection algorithm 264 takes as input the predicted IPU power profile 268 , which it uses to make predictions about power usage for each of the input subsystem frequency 265a and the processing subsystem frequency 265b before allocating frequencies to the IPU hardware 262 .

[0041] The OS CPU interface 246, OS-GPU interface 256, and OS IPU interface 266 allow performance level selection algorithms and circuits in the individual processing units to receive system parameters to be fed into the processing unit frequency selection and allow the individual processing units to feed back power consumption information associated with the frequency selection (e.g., new target quantized power consumption) to the platform level power control algorithm 222. The duplication of the common performance level selection algorithms 244, 254, 264 in multiple processing units and the ability to perform power-aware frequency allocation in the individual processing units enables distributed control efficiencies and the ability to easily determine performance per watt for each performance level selection decision.

[0042] Each of the CPU performance level selection algorithm 244, the GPU performance level selection algorithm 254, and the IPU performance level selection algorithm 264 receives a corresponding predicted unit-specific power profile 248, 258, 268 from the platform-level power control algorithm 222. Each of the three performance level selection algorithms may also receive unit-specific scalability values ​​and unit-specific utilization values ​​from hardware counters as outlined in Tables 1 and 2 below. Each unit-specific power profile 248, 258, 268 may provide a "priori" relationship between utilization, frequency, and power for a given unit (e.g., determined based on a processor model before runtime or even before production). The power profile for each of the CPU, GPU, or IPU may be based on a pre-silicon model or on post-silicon measured data or on a synthetic workload. A pre-silicon model may be a model based on a prefabricated or processor design simulation. Some power models may assume power viruses, which means 100% utilization. Other power models may assume specific workloads with corresponding processor dynamic capacitance (Cdyn). Equation P = Cdyn*V 2 *f can be used to determine Cdyn, where P is the power drawn, V is the voltage and f is the operating frequency of a given processing unit. The value of Cdyn is workload dependent, so it can vary based on processor utilization and scalability. The power profile for a given processing unit can be generated in any of a number of different ways, but no matter how it is generated, the predicted power profile is used to generate the following processing unit management equation:

[0043] 1) Power as a function of frequency and utilization,

[0044] 2) frequency as a function of utilization and power, and

[0045] 3) Utilization as a function of power and frequency

[0046] The three metrics described above can be used without also involving management equations for scalability. Scalability is an inherent property of the workload; more specifically, how the workload affects a processing unit or, for example, a CPU execution pipeline. Since analytical modeling of a large number of different workloads may be impractical, it is useful to base management algorithms on utilization, frequency, and power rather than scalability. While the equations may not be 100% accurate across different workloads for all possible Cdyn for a processing unit (platform component), they are still accurate enough to determine general trends in the power consumption of a given processing unit to enable higher and effective performance management.

[0047] Each of the CPU performance level selection algorithm 244, the GPU performance level selection algorithm 254, and the IPU performance level selection algorithm 264 can have three different categories of potential inputs. The three categories are: (i) system-level inputs; (ii) utilization inputs specific to each processing unit; and (iii) scalability inputs specific to each processing unit. System-level inputs can include power limits, thermal limits, and energy performance preference information such as a first power limit "PL1" and a second power limit "PL2". These system-level inputs provide centralized guidance from the platform (system) level, allowing each of the processing units such as the CPU hardware 242, the GPU hardware 252, and the IPU hardware 262 to operate autonomously to a certain extent. The energy performance preference information can be, for example, platform-wide or unit-specific or component-specific or SOC-specific. Platform-wide energy performance preferences can be used to guide each of the individual performance level selection algorithms 244, 254, 264. Utilization inputs can be different between processing units. For example, the CPU hardware 242 and the GPU hardware 252 can each have their own metrics for measuring current utilization. According to the present technique, each processing unit may expose the current utilization value in an architecturally consistent manner.

[0048] For example, CPU utilization can be measured using a number of performance monitor counters whose values ​​can be stored in registers. The counters include:

[0049] • Reference counter ("MPERF"): An active running counter that counts at a fixed timestamp counter (TSC) clock rate and counts only during the CPU's active state (when at least one physical core is active in a multi-core system). The TSC clock rate typically corresponds to the guaranteed base frequency of the active CPU or some other baseline counter that increments (or decrements in other examples) at constant intervals.

[0050] • Execution Counter (“APERF”): An active run counter that counts at the actual execution clock rate at that moment in time. This actual clock rate may vary over time based on performance level management and / or other algorithms. This register counts only during the active state.

[0051] • Useful Work Counter ("PPERF"): A counter for active work that counts at the actual clock rate, similar to APERF, except that it does not count when the activity is "stopped" due to some dependency. An example of such a dependency is when the CPU is gated on the clock domain of another platform component such as memory.

[0052] Using the above CPU performance counters, the CPU utilization and CPU scaling factor can be defined as follows:

[0053] Utilization: U = (ΔAPERF / ΔMPERF)*ΔTSC Equation 1.1

[0054] Scaling factor: S = ΔPPERF / ΔAPERF Equation 2.1

[0055] Where the symbol "Δ" represents the corresponding count value change in a given count sampling interval Tz. The value TSC is the time interval between counter increments (or decrements) of the baseline counter. The utilization actually represents the "work" done, because APERF is a free-running counter running at the current frequency, thereby representing the degree of CPU activity, and MPERF is a free-running counter at a fixed frequency. Note that if the execution of program instructions by the CPU is not stopped, PPERF = APERF, and therefore the scaling factor is equal to one. In such a situation, the time taken to complete each processing activity in a given time window is just the inverse of the actual frequency in that window. However, in practice the scaling factor for real workloads can be less than 1. Typical values ​​can be in the range of 0.75 to 0.95, but values ​​outside this range are not uncommon.

[0056] Note that although the utilization equation 1.1 does not involve the non-stop activity count ΔPPERF, the utilization does take into account scalability and the impact or stalls in processing. For example, this can be understood by considering utilization as "work" done, meaning that the CPU is in an active state doing some "work", rather than in an idle or sleep state or another low-power state. This may be "pure work", for example, pure CPU calculations. However, this work may also include time when the CPU is busy not doing useful work (but the counter APERF is still running) due to being stopped, waiting for memory, waiting for input / output, etc.

[0057] So, if at a first frequency f1, the CPU experienced a particular utilization (i.e., 80% utilization at 800 MHz), it would be useful to know what the corresponding utilization would be for running the same processing workload (e.g., a program application) at a different frequency f2, such as 1600 MHz. In a purely "scalable" workload, 40% utilization might be expected due to doing the same work at double the speed (this workload represents a scalability of 1, or 100% scalability). However, in practice workloads are rarely perfectly scalable. Doubling the frequency to f2 (in this example) may reduce utilization by a different amount due to inherent stalls or other artifacts, as stalls may not inherently scale—waiting for memory will still be waiting for memory even if the CPU is running at a higher frequency.

[0058] Figure 2B Schematically illustrates example architectural counters for tracking CPU performance. A first count block 270 within a sampling period Tz includes a number of counts 272 of ΔPPERF corresponding to useful work performed by the CPU, which are incremented at three CPU execution frequencies. The execution frequency may vary within the sampling time interval according to a frequency assigned by at least one of the SOC power management circuit 122 or the CPU performance level selection algorithm 244. Also in the first block 270 where all counts are made at the CPU execution rate is a number of counts ΔAPERF 274 corresponding to cumulative counts when the CPU is active (including active but still stopped). The last portion 274 of the first block 270 of duration Tz includes a duration 276 when the CPU is not active, so the counter based on the execution frequency does not increment or decrement counts.

[0059] exist Figure 2B In the case of , a continuous block of active but not stopped counts and active but stopped counts are illustrated in Tz. However, within a given sampling interval, the active but not stopped counter (PPERF) may be triggered intermittently, with groups of discontinuous periods of stopped activity occurring simultaneously in the sampling window.

[0060] Figure 2B The second counting block 280 in FIG. 2 schematically illustrates the operation of counting at a given constant frequency F TSC The counter counts, and the sampling period T z The total number of counts ΔTSC 282 in ΔTSC 282 includes a baseline count value ΔMPERF 284 of the counts when the CPU is in an active state. By comparing the first block 270 and the second block 280, it can be seen that the duration of the active count t act is the same for both blocks, but the count values ​​are different because the counts are driven by different clock signals. The duration t of the count ΔPPERF 272act 286 corresponds to the number of cycles of CPU activity that are not stopped. Duration t act +t stall 288 corresponds to the CPU activity period including when the CPU is stopped. Duration T z =t act +t stall +t off 289 corresponds to the sampling window duration including CPU activity, CPU stop, and CPU inactivity durations.

[0061] GPU utilization can be calculated using different counters than those used for the CPU. A given GPU can have more than one EU, so a count can be maintained for each EU when the EU is active (not idle) and the count will track the number of graphics processing cycles when each EU is stalled. A EU can be considered stalled when at least one GPU thread is active but not executing.

[0062] Table 1 below specifies some example CPU-related performance monitoring parameters, whereas Table 2 specifies some example GPU-related performance monitoring parameters.

[0063] Table 1

[0064] CPU Terminology describe APERF Architecture counter that increments at actual / current frequency MPERF An architectural counter that increments at a base / guaranteed frequency, typically High Frequency Mode (HFM) TSC Timestamp counter frequency (ΔAPERF / ΔMPERF)*HFM Utilization (ΔMPERF / MPERF) *ΔTSC Scalability Factor ΔPPERF / ΔAPERF P Pe is the most efficient operating point, below which power is wasted Palpha Palpha reflects the maximum power the operating system is willing to pay to obtain a certain required performance.

[0065] Table 2

[0066] GPU Terminology describe GPU tick count The number of cycles that the execution units in the GPU are busy EUNotIdlePerSubslice EU utilization when EU is not idle (can include time when EU is stopped) EUStallPerSubslice The number of slice clocks when the EU has threads but is not running any FPU or EM instructions (that is, when the EU is stalling) NumActiveEUs Number of active EUs Scalability for GPUs ((ΔEuNotIdlePerSubSlice – ΔEuStallPerSubslice)*4 * 8) / (ΔGPUTicks * NumActiveEUs) Note that this metric of graphics scalability is calculated in this example because there is no ready counter in hardware like the CPU, but a counter is available

[0067] In Table 2, consider that a GPU can be conceptually thought of as being built from "slices," each of which can contain multiple (such as three) "sub-slices." These sub-slices can each include multiple (e.g., 8, 16, or 32) "execution units" (EUs) to perform large blocks of processing. Each sub-slice can also include a texture sampler for retrieving data from memory and passing it to the EU, and may include one or more other components. The GPU can also include what may be referred to as "unsliced" components, which are GPU components that process fixed geometry and some media functions outside of the slice.

[0068] In some example GPUs, unsharded may have a separate power and clock domain that may be independent of sharded. Thus, if the hardware encoding and decoding capabilities of unsharded are currently being used, all shards may be powered down or powered off to reduce energy consumption. Additionally, unsharded may run at a higher or lower rate than sharded, thereby providing the ability to improve performance or power usage depending on the specific processing tasks being run. Note that the scalability equation for a GPU in Table 2 is only an example, and the factor 4*8 in the numerator is GPU architecture specific and is non-limiting.

[0069] Calculation of workload scalability

[0070] The scalability of workload with frequency can be calculated differently for each processing unit. For example, for a CPU, a scalability equation can be derived for the CPU by making many approximations. Here, for the current frequency f c To the new target frequency f n The new utilization u associated with the target frequency can be evaluated based on the following factors: n :

[0071] - From existing f c Frequency change to the new target frequency fn

[0072] - Current utilization u c

[0073] - Current scalability factor S as the ratio of ΔPPERF / ΔAPERF c

[0074] In particular, the following equation shows how the current frequency and the current scalability S can be determined from the CPU architecture counters as indicated in Table 1 above: c Calculate the new target frequency f n The predicted new utilization rate.

[0075] Equation 3

[0076] This allows the impact of changes in target frequency on utilization to be evaluated before frequencies are assigned to CPUs.

[0077] The derivation of Equation 3 involves many simplifying assumptions. Different embodiments may adopt different equations for scalability based on the approximation made to derive the functional relationship between frequency, scalability, and utilization. The following equation defines the scalability s in a time window Z z and utilization rate U z In this example, t act It is in T z The duration corresponding to the CPU being active; t stall It is in T z The duration of time when the CPU is active but stopped; and t off It is in T z The duration of time when the CPU is inactive (e.g., turned off or in sleep state). In any time window T, Scalability and utilization can be defined in terms of duration as follows:

[0078] CPU scalability, Equation 4

[0079] CPU utilization, Equation 5

[0080] Figure 2B Schematically illustrates how the execution counter (APERF) and the useful work counter (PPERF) for the CPU can be visualized in time and in frequency.

[0081] A further simplifying assumption made in the derivation of Equation 3 above is that the stopping time t stall is not an explicit function of the local DFVS (execution clock) and is therefore invariant to changes in execution frequency. It is equivalent to separately for all current frequencies f c and the target frequency f n For example, c ΔPPERF = at f n Thus the work range associated with the processing task as counted by the useful work (no-stop) counter ΔPPERF remains the same for different execution frequencies, but its corresponding activity duration t act It does vary with frequency.

[0082] Calculation of scaled power at different target frequencies

[0083] Based on pre-silicon (i.e., pre-fabricated) or other power models for each PU (CPU, GPU, etc.), appropriate equations can be derived for scaled power. The power model can typically be based on a power virus for the processing unit, but may be as accurate as the model allows. The specific equations used to calculate the scaled power at different frequencies can also be specific to each processing unit. Typically, the cumulative "scaled CPU power" can be expressed as a function of the power of the individual logical CPUs (hyperthreads). Such a mathematical relationship can be derived using appropriate curve fitting tools (algorithms) such as power distribution diagrams suitable for a given processing platform. Similar scaled power equations can be derived for other processing units (PUs) such as GPUs and IPUs. Scaled power can characterize how the power consumption of a PU changes from a given initial value of frequency to any final value due to frequency changes.

[0084] Similarly one can derive equations for:

[0085] - Frequency as a function of utilization and power

[0086] - Utilization as a function of power and frequency

[0087] The equation for power as a function of frequency and utilization may be part of a pre-silicon model. This pre-silicon model may be, for example, a spreadsheet giving values ​​for power for different frequencies and utilizations in addition to other parameters such as temperature, implemented process technology, etc.

[0088] Using the above assumptions and inputs from the scaling power and system utilization equations, one can use e.g. Figure 2A An example of a PU-specific performance level selection algorithm implemented by the CPU performance level selection algorithm 244 is as follows:

[0089] 1. Check PU utilization (architecture counters); apply averaging (tau / alpha for short and long sampling windows Tz). Averaging allows observation of coarse and fine granularity of workloads to distinguish, for example, more or less "bursty" workloads. This can be done by (a) observing and averaging utilization over a relatively long window (i.e. 32 ms). This is denoted as alpha. Then (b) the same workload can be observed over a relatively short window, i.e. 4 ms tau). The alpha window can provide trends in utilization, whereas the tau window allows observation of spikes in utilization.

[0090] 2. Calculate system utilization (across all PU cores), which can be done via a priori mathematical derivation based on individual CPU utilization, just like scaling power. Architectural performance counters such as APERF and MPERF can be used to measure the utilization of each thread. The utilization can be calculated using the sum of the utilization of multiple cores, such as by calculating a weighted average across multiple cores.

[0091] 3. Check the current scalability Sc

[0092] 4. Check the polarity (util+ / util-) of any change in utilization of a given PU during the recent time period

[0093] 5. Using the above scaled power and system utilization, estimate the scaled utilization and resulting power

[0094] 6. Reward Scalability

[0095] a. If the current scalability Sc is above a certain (programmable) threshold, a higher frequency option is more likely to be chosen from the range of available target frequencies

[0096] b. If the current scalability Sc is below a threshold, the algorithm may be more conservative in allowing higher frequency selections, since the additional power consumption associated with higher frequencies is less likely to result in increased useful work due to a higher proportion of stops characterizing the workload.

[0097] 7. Choose the best (or at least most efficient) frequency for a given scaling power and system utilization

[0098] 8. Check if the scaled power is within the limits specified by the EPP, as the EPP can guide the instantaneous power boundaries

[0099] 9. Resolve frequency allocation for a specific processing unit (CPU, GPU, IPU or other PU) and set the clock to the resolved frequency as part of DVFS.

[0100] With respect to item 8 above, the EPP can guide the instantaneous power boundary by providing a scale that indicates the highest frequency selection at one end and the lowest frequency selection at the other end. That is, if, for example, EPP=0 is set by the user to indicate that they want the highest performance, the choice of frequency can be to select the highest frequency. If EPP=100 (lowest power), then the lowest frequency can be selected.

[0101] If the EPP is somewhere between 0 and 100, and there are, for example, ten different frequencies that can be selected to meet common utilization, then the EPP can be used to guide which frequency of the range is selected. For example, if EPP = 50, then the midpoint of the available frequencies can be selected.

[0102] Figure 3 The processing units of the SOC and their corresponding architectural counters are schematically illustrated. The system includes a CPU 310, a GPU 330, an IPU 340, and a system agent 360. The CPU 310 has a local CPU power management circuit 324 to monitor the counters in the CPU and communicate with Figure 1 The SOC power management circuit 122 coordinates the CPU execution frequency allocation. The local CPU power management circuit 324 can be implemented entirely in hardware or entirely in software or a combination of hardware and software. The CPU 310 has a source clock generator 312 that controls the execution of program instructions by one or more physical cores 314. The source clock generator 312 operates at a clock frequency that can vary depending on the desired processing performance level. The CPU includes a stop-aware clock activity counter 316 for counting clock cycles of the source clock 312 excluding when a stop is occurring (e.g., performing a ΔPPERF count), and further includes an activity counter 318 that counts at a rate determined by the source clock but counts all activity cycles regardless of whether the process has stopped (e.g., performing a ΔAPERF count). The CPU 310 also has a fixed clock frequency f TSC The fixed clock generator 320 has a fixed clock frequency f TSCThe count rate of the fixed timestamp counter 322 (eg, the ΔMPERF count) is controlled. The CPU power management circuit 324 receives one or more power limits from the SOC power management circuit 122 to constrain the selection of an execution frequency to be selected from a range of target frequencies.

[0103] The GPU 330 includes: a GPU management circuit 332, which is used to set an appropriate GPU performance level; a GPU clock generator 333; a first execution unit 334; a second execution unit 336; and a GT busy counter 338. In this example, the GPU has its own GPU clock generator 333, but in other examples with different GPU architectures, it may not have its own GPU clock generator 333. The GPU 330 can have more than two EUs, and similar to a multi-core CPU, at any one time, all EUs or only a subset of the EUs can be active when graphics processing is being performed. When at least one of these EUs is active, the GT busy counter 338 can count the number of clock cycles Δgt of the GPU clock generator 333 in the sampling interval. busy The GPU 330 of this example has only a single counter 334 and as indicated in Table 2, scalability for the GPU is calculated using the count ΔGpuTicks taken from the GT busy counter 338 and by software tracking the number of active EUs and when each EU is active and when each EU is stopped (e.g., not running any floating point unit instructions). The GPU 330 may receive one or more GPU power limits 335 from the SOC power management circuit 122.

[0104] The GPU management circuit 332 may be an on-demand manager that operates as one or more power control algorithms in a graphics microcontroller (not shown). When no frames per second (FPS) target is specified, the management circuit 332 may use active→idle and idle→active interrupts to identify GPU active time. For example, a minimum of 30 FPS may be desired for good picture quality in order to play a video, and some workloads such as games may typically run at a relatively high FPS, such as 90 FPS. The FPS target may be used as a performance tuning parameter. The ratio of GPU active time to total time within a given sampling interval may give GPU utilization, which in turn may be used to drive GPU frequency selection.

[0105] The utilization and scalability equations may be different for different processing units.Example equations for GPU utilization and GPU scalability may be as specified below.

[0106] Equation 1.2

[0107] Equation 2.2

[0108] Where ∆(AggregateCounter) is the cumulative count of active cycles of all EUs 336, 338 in a given sampling interval; NumActiveEUs is the number of active EUs in a given sampling interval; and ∆(GpuTicks) may be equal to ∆gt busy .

[0109] The IPU 340 has an IPU power management circuit 342 that can be implemented entirely in hardware, entirely in software, or via a combination thereof. The IPU 340 has an input subsystem clock 344 (IS_clk) and a processing subsystem clock 346 (PS_clk), which makes it possible to control the input subsystem and processing subsystem frequencies separately. The IPU 340 can have an architectural counter (not shown) similar to the CPU to count active IPU cycles and non-stop active CPU cycles at the relevant IPU frequency. The IPU 340 can also have a fixed timestamp counter (not shown) to count active IPU cycles. A separate set of active and non-stop active architectural counters can be provided to count at the input subsystem clock frequency and the processing subsystem clock frequency. The IPU can receive one or more IPU power limits 345 from the SOC power management circuit 122, which can constrain the frequency selection implemented in the IPU 340. In an alternative embodiment, the IPU 340 can use one or more IPU drivers to monitor utilization and make power-aware frequency selection based on the IPU power profile model.

[0110] At a high level, the IPU 340 may have an input subsystem (IS) and a processing subsystem (PS) and each may be individually controlled with respect to their respective operating frequencies. The IPU driver 348 may be an internal heuristic to request the respective frequencies based on utilization and usage. For example, during image capture, only the IS component may be controlled to respond at the desired frequency. Subsequently, after the image capture (or in an interleaved manner if necessary), the processing subsystem may be controlled to run at the desired frequency to meet the use case (e.g., a specific data capture rate). In one example arrangement, the IPU 340 may have the following DVFS support:

[0111] • IS_clk 344 is derived from SA_PLL, which is a system agent phase-locked loop that can operate in 1600MHz or in 2000MHz (can be fixed per platform). The driver can request an IS_clk 344 frequency in the range 100-400MHz. A set of "PUNIT" firmware 350 can select the actual divisor setting based on the SA_PLL phase-locked loop frequency. Note that the system agent 360 (see Figure 3 ) is the component that couples the CPU, IPU, and GPU to the display, memory, and I / O controllers.

[0112] • PS_clk 346 can operate in 25 MHz steps from 200 MHz to 800 MHz, with additional frequencies below 200 MHz used for debug purposes. The lowest possible frequency may be, for example, 50 MHz. Frequencies below 200 MHz may not be efficient for functional use.

[0113] The frequency ranges stated above are non-limiting example ranges. The IPU driver 348 requests the corresponding IS and PS frequencies from the PUNIT firmware 350, and the final granting of the requested frequencies may be under the control of the PUNIT firmware 350. If one or more IPU power limits 345 (e.g., system limits such as thermal limits) do not allow the clock frequency to be increased, the PUNIT hardware 350 may ignore the request from the IPU driver 348 to change the clock frequency.

[0114] The IPU 340 is an example of a complex processing unit that has internal coordination implemented between the IS and PS subsystems, so implementation of power-aware frequency allocation to each subsystem according to the present technology can result in improved power efficiency through intelligent allocation of power between the two subsystems and improved overall performance for a given IPU power budget.

[0115] Figure 4 The functional flow for a power-aware frequency selection algorithm for a processing unit is schematically illustrated. At block 410, the current utilization of the processing unit is determined based on an architectural counter or otherwise. At block 412, filtering of the current utilization value may be performed by averaging over different sampling periods to improve accuracy. Since the utilization may vary within the sampling time window, reducing the sampling window may improve accuracy. The overhead of shorter time sampling intervals may be reduced by using a microcontroller-based or firmware-based implementation. The sampling interval may be different for different processing units. For example, the CPU sampling interval may be shorter than the GPU or IPU sampling interval. At block 414, the current power usage of the processing unit may be determined, for example, using a local power sensor, and the current power Pc Supplied to comparison block 416. At box 420, the current scalability S is determined, for example, according to equation 1.1 for the CPU or according to equation 1.2 for the GPU. c .

[0116] At block 422, a value for the scaled utilization is calculated for a range of target frequencies that could potentially be selected as the new frequency, although in some examples a single target frequency may be evaluated. The value for a given target frequency f may be calculated by applying the following to Equation 3: ti Scaling utilization: (i) S determined at block 420 c The value of; (ii) the current frequency; (iii) the given target frequency f ti ; and (iv) current utilization u c At block 422, for each target frequency, the predicted power profiles 248, 258, 268 (see FIG. 2) may be used to determine a corresponding value set {utilization factor U ti ; Power P ti and scalability}. Once the parameter sets for the multiple target frequencies have been determined at block 422, a new frequency may be selected based on a PU-specific relationship between frequency, power, and utilization at block 424. The frequency selection may be performed such that the new target quantized power is known before the new frequency is implemented by the processing unit. The selected new frequency is implemented at block 432, wherein the performance level of the PU is adjusted to align with the new frequency value.

[0117] Once the selected frequency has been achieved, the selection of a new power at block 426 may be made based on trends observed in the most recently measured utilization and scalability values ​​or based on the difference between the expected power and / or utilization calculated in the previous frequency update round and the corresponding actual (measured) values. If the utilization change is below a threshold, the PU may continue to operate at the current power P. c However, if the utilization increases or decreases by more than a threshold amount, a new power may be allocated at block 426. At block 426, a new power P is assigned. nIt may also or alternatively depend on system parameters such as EPP 428 and minimum power, maximum power, or another power limit. For example, a PU may assign relative priority to performance and power saving based on EPP 426. The new power allocated at box 426 may also depend on the per-unit power limit assigned by the platform power management circuit 110 for the PU. Therefore, even without any significant changes in utilization or power consumption, a new power may be assigned to the process at box 426 in response to an increase in the per-unit power limit. An example of a formula that may be used in the calculation of the new power at box 426 is Pn=(Pc±Err)*K*f(EPP), where K is a constant and Err is the difference between the current power and the expected power calculated in the previous frequency update cycle based on the target power and scalability function, and f(EPP) means a function of EPP. EPP may be a system parameter or may alternatively be processing unit specific.

[0118] The selection of the new power at box 426 feeds into the selection of the new frequency at box 424. The new frequency selected at box 424 may have an associated predicted (i.e., expected) new power and predicted new utilization and predicted new scalability. At box 430, the expected power calculated in the previous frequency update cycle (where the current operating frequency was determined before it was implemented) is set equal to the predicted new power determined at box 424. The expected power is the second input to the comparison box 416. The comparison of the expected power set at block 430 (corresponding to the previous cycle and the currently implemented frequency) and the current power output from box 414 at box 416 allows any deviation between the assigned new power and the current power to be corrected at box 426 (represented as "Err" in the above equation). Therefore, there is a feedback loop to correct errors in the power prediction made using the power profile.

[0119] Figure 5 is a flow chart schematically illustrating a power-aware management algorithm for a processing unit. The process of selecting a processing unit performance level begins at element 510, where the current processing unit utilization U is determined, for example, using equation 1.1 above for a CPU or using equation 1.2 above for a GPU. c Alternatively, or in addition, current power consumption of the processing unit may be measured at block 510. The power consumption may be determined based on values ​​obtained from one or more power sensors, or alternatively may be derived from utilization measurements based on a corresponding power profile.

[0120] Next, at element 520, any utilization change ΔU or change in power consumption ΔP relative to the previous period may be determined. Figure 4The measured values ​​of utilization or power at the current operating frequency at the current time are compared with the expected utilization or power image predicted from the power profile in the previous frequency update cycle as shown in blocks 414, 416 and 430 in FIG. Alternatively, ΔU and ΔP may be determined based on past trends of measured values ​​of utilization and power observed from the current execution of a given processing workload or from stored data corresponding to the same or similar workload being executed on the same or similar processing unit.

[0121] At flow chart element 530, a determination is made as to whether a frequency change from the current operating frequency is appropriate. This determination may depend on at least one of the ΔU and ΔP determined at element 520, but may also or alternatively depend on other data inputs as shown at element 535, where power limits such as the minimum power consumption for the processing unit, the maximum power consumption for the processing unit, and one or more additional power limits such as the increased power limit that can be maintained only until the maximum time may be considered when making the determination. In addition, an EPP specific to the processing unit or applicable to the entire processing platform may be considered at element 530. In one example, if ΔU or ΔP is greater than or equal to the corresponding minimum threshold amplitude, the frequency change may be considered appropriate at element 530. In this example, if ΔU is greater than the corresponding threshold amplitude, the frequency change corresponding to the power consumption change is considered appropriate, but if ΔU is less than the threshold amplitude, the change in the current operating frequency in this frequency update cycle is not considered appropriate. At element 530, if the change in the operating frequency is not considered appropriate, the process returns to element 510 and waits until the next frequency update cycle is initiated. The frequency update cycle may be performed periodically. The period for frequency update may be different for different processing units in these processing units. For example, the frequency update cycle period may be 15 ms for the CPU, 100 ms for the GPU, and 100 ms for the IPU.

[0122] If it is determined at decision element 530 that a frequency change is in fact appropriate, then the process proceeds to element 540 where a new target quantified power dissipation is determined based at least in part on ΔP or ΔU or both. The target new power may be quantified in units of measure for power such as watts or in some other manner that allows the processing platform to know the power dissipation by each processing unit before implementing a frequency change by changing the clock rate. This conveniently enables more processing platform control over performance per unit of power consumed. As will be discussed below in Fig. 6A As described in , the new target quantized power can be identified as corresponding to the equidistant power lines on the power distribution graph.

[0123] Once the target power is identified at process element 540, the process proceeds to element 550 where a new operating frequency (or operating voltage, since the two parameters are related) is selected with an eye toward achieving as close a match as possible with the new target quantized power dissipation. One way to accomplish this is to utilize one or more observable values ​​(such as the current observed utilization and the current scalability determined using the architectural counters) and use a scalability function such as specified by Equation 3 above to determine which new frequency is most likely to result in achieving the desired new target quantized power dissipation given knowledge of the current utilization. The current scalability Sc may be determined for a processing unit, for example, by using Equation 2.1 for a CPU or by using Equation 2.2 for a GPU. In some examples, the scalability value determined from the architectural counter may correspond to a frequency update period other than the current period. For example, previously measured scalability values ​​may be used in conjunction with known isometric trends to estimate an appropriate current scalability value to use at element 550 when selecting a new frequency.

[0124] In the parameter space representing the operating point of a processing unit, the parameters of frequency utilization and power are both related. The new frequency is what is selected and the new power consumption is the target value that drives the selection of a specific new frequency for power-aware management distributed among multiple PUs. Utilization and frequency are inherently linked, but scalability is another factor that can be considered to improve the control of the power consumption resulting from the frequency changes being implemented. The use of power profiles, target power consumption, and scalability measurements and scalability functions allows the power consumption of a processing unit to be more predictable during frequency updates according to the present technology.

[0125] Once a new operating frequency has been selected at element 550, the process proceeds to element 560 where control is exercised to implement the selected new frequency in a subsequent processing cycle, and then the cycle returns to the beginning of the flow at element 510. Note that when a new frequency is selected at element 550, at least one of the new target quantized power and the corresponding expected utilization can be fed back to element 520 for use in determining ΔP or ΔU for a subsequent frequency update cycle.

[0126] The duration of the frequency update period may be different from the duration of the processing period and the processing period duration itself is potentially variable as a result of DVFS.In some examples the frequency update may be performed intermittently rather than periodically and the frequency update periods specified above are non-limiting examples.

[0127] Fig. 6AAn example of a 3D power profile of a CPU generated using a synthetic workload is schematically illustrated. The power profile 600 is a 2-dimensional (2D) surface in a 3-dimensional (3D) parameter space, with CPU frequency in megahertz (MHz) along the x-axis, utilization (or equivalently load) in percentage of active cycles along the y-axis, and CPU power in milliwatts (mW) along the z-axis. The grid lines parallel to the x-axis are equidistant utilization lines on which the utilization has the same value, the grid lines parallel to the y-axis are equidistant frequency lines on which the frequency has the same value, and the curves 610a, 610b, 610c plotted in the 2D surface (which is actually a power surface) are equidistant power lines on which the power dissipation has the same value. Each equidistant power line is parallel to the xy plane, because each point on the line has the same power consumption value (z-axis). Therefore, the equidistant power lines are similar to contour lines on a map. Note that power consumption along equidistant power lines such as line 610a appears visually different due to the 3D nature of graph 600. The same applies to equidistant frequency and equidistant utilization lines on the 2D power surface of the 3D graph.

[0128] In some examples, the power profile 600 (or at least a portion thereof) may be generated based on monitoring the operation of the processing unit. In some other examples, the power profile may be generated at least in part by monitoring the operation of one or more other processing units (e.g., which may be external to the processing platform 100), wherein the one or more other processing units have similar characteristics and are of a similar type as the processing unit 125, 127, 129 of interest. In some other embodiments, the power profile may be generated by computer simulation of a model of the processing unit. In still other examples, a combination of these above-discussed methods may be used to generate the power profile.

[0129] Any point on the power surface can correspond to the operating point O(f, U, P) of the CPU characterized by the frequency value, the utilization value and the power value. Therefore, for example, the power distribution diagram can be used to predict or estimate how the CPU utilization changes with the increase of the operating frequency for a given processing workload. The power penalty for increasing the operating frequency can also be determined based on the power distribution diagram. It can be seen from the 2D power surface rising towards the right rear corner of the 3D diagram that power consumption tends to increase with both the increase of utilization and the increase of frequency. When the frequency is relatively low or the load is relatively low or when both the frequency and the load are relatively low, the power consumption is also relatively low.

[0130] When dynamically changing the processor frequency, certain assumptions about the processing workload (e.g., a particular program application being executed) may be made to allow the use of power profile 600 to predict expected utilization and expected power consumption when the operating frequency is changed from a current frequency fc to a new frequency fn, which may depend on at least one of the utilization observed at fc and the power consumption observed at fc. Different power surfaces corresponding to different processing workloads may be available, such as different program applications, such as a gaming application, a video conversion application, and a DNA sequencing application.

[0131] In conventional systems implementing DVFS, frequency may be used as the primary parameter for determining the operating point of the CPU (processing core). For example, there may be a maximum predetermined operating frequency that cannot be exceeded, so the DVFS circuit may set the operating frequency such that the selected frequency is constrained by the maximum frequency. There may also be a minimum operating frequency. However, as can be seen by Fig. 6A As can be seen from the power surface of FIG. 620 , power consumption can vary greatly along equidistant frequency lines (e.g., line 620a) around 2800 MHz. In particular, along equidistant frequency line 620a, as utilization (load) increases, power consumption also increases in this non-limiting example from about 1000 mW at the lowest utilization to about 14000 mW at 100% load. Even if a maximum power threshold is set, simply selecting an operating point O(fn, Un, Pn) based on a frequency range and utilization without knowing the power consumption changes associated with performance level changes may be potentially problematic. Indeed, while the operating frequency can be changed as in a conventional DVFS system, even if it can be predicted based on a 3D power profile whose operating point is located on a particular equidistant power line, utilization cannot be easily controlled because any utilization changes are the response of the processing unit to the frequency change based on the workload being executed. Therefore, although the power profile can provide some guidance, it may be difficult to predict what effect frequency changes will have on utilization and power consumption.

[0132] However, according to the present technology, energy efficiency can be improved and more flexibility can be achieved in setting appropriate frequency values ​​by building power consumption awareness into the frequency selection process. This power consumption awareness can take into account both target power and scalability, and some examples can use equidistant power lines of a power distribution diagram to help set a new target power consumption and also use scalability values ​​(e.g., scalability values ​​read from an architectural register of a processing unit) to help guide the implementation of power consumption at or close to the new target power consumption. This can provide the processing unit with a reliable sense of the power consumption of the new operating point even before the new frequency has been achieved. This is likely to result in fewer assertions of suppression signals and reduce the possibility of unexpected system failures due to thermal limit violations, peak current violations, voltage drops, etc.

[0133] Furthermore, rather than the CPU's operating system centrally managing power control on a processing platform, power-aware management can be repeated in each processing unit of the platform with the ability to change operating frequency. This allows power consumption to become a common item that is managed in all processing units. The use of power-aware management algorithms repeated in two or more processing units of a processing platform also allows the ability to achieve distributed control efficiency, which can be defined as performance per watt. This becomes possible because power-aware management means that power (e.g., wattage) is well quantified and known at each decision point.

[0134] Figure 6B Schematically illustrates a graph of frequency versus utilization of a processing unit. The graph shows equidistant power lines 650, which can be considered to be Fig. 6A The projection of the equidistant power lines 610a, 610b, and 610c in the 3D diagram onto the load-utilization plane. Figure 6B A scalability line 660 is plotted in the graph of FIG. 660, which represents the utilization versus frequency for a particular processing unit. The scalability line in this example has been generated according to the scalability equation (Equation 3), which allows the new load to be calculated based on the current frequency, the new (target) frequency, and the current scalability. The current frequency is known, a new frequency can be selected for evaluation, and the current scalability can be measured via architectural counters (e.g., APERF and PPERF for the CPU).

[0135] exist Figure 6B In the example of , the first data point 672 and the second data point 674 each correspond to the current frequency. The first data point 672 corresponds to the expected utilization at the current frequency, which is predicted as the utilization of the current frequency in the previous frequency setting cycle before the current frequency is implemented by the processing unit. Data point 672 is similar to Figure 4 The expected power at element 430 in the process flow diagram of . The second data point 674 corresponds to the utilization actually observed at the current frequency, for example, by using the architecture counters APERF and MPERF and equation 1.1.

[0136] For example, a deviation ΔU 676 between the observed utilization value of the second data point 674 (y-coordinate) and the expected utilization value of the first data point 672 (y-coordinate) may occur due to changes in workload characteristics (e.g., scalability or changes in the nature or processing tasks) since the last frequency update cycle. However, there may be multiple factors that affect the frequency selection for the frequency update, so system parameters (e.g., EPP or system level power limit) or processing unit parameters (e.g., from the platform power management circuit 110 (see Figure 1At least one of the unit-specific power limits received by the processor 670 (i) and the unit-specific power limits received by the processor 670 (ii) may trigger a change in the operating frequency setting of the processing unit. Rather than determining a utilization difference ΔU 676 between the measured (observed) operating point 674 and the expected operating point 672, in an alternative example a power difference may be determined between these two points. The power difference will suggest a utilization difference at the given frequency at which the prediction is made.

[0137] The scalability line (or curve) 660 may be generated from Equation 3 using the measured values ​​of scalability and utilization corresponding to the second data point 674, so the second data point 674 is a point on the scalability line 660. In other examples, the scalability line may not pass through the second data point 674 corresponding to data measured in the current cycle, but may be determined based on scalability trends or measured data values ​​from different cycles. The scalability line 660 may correspond to Fig. 6A The scalability line 660 is likely to cut through the different trajectories on the 2D power surface relative to the equidistant power lines 650. Fig. 6A scalability line 660 may characterize how the utilization of a PU may change from a given initial value of frequency to any final value due to a frequency change and may therefore provide a causal relationship between the frequency change and the utilization change. In contrast, the equidistant power lines simply represent points in the 3D parameter space but without such a causal relationship. In a map analogy, the scalability line 660 may be thought of as a sidewalk on the surface of a slope that intersects (at one point) the contour line corresponding to the power consumption target. Figure 6B can be considered as Fig. 6A 2D projection of the 3D graph of FIG. 650 , but note that the scalability line 650 has different power values ​​along its length. The scalability function provides the information needed to predict how utilization varies with frequency so that a new target power, or at least a value close to the target, can be achieved when a new frequency is achieved by the processing unit.

[0138] The processing unit power management circuitry (e.g., 324, 332, 342) in the current frequency setting cycle will set a power consumption target taking into account at least the utilization change corresponding to the deviation ΔU 676 and possibly other factors such as the energy performance preference for the individual processing units or for the platform as a whole. Any changes in the processing unit power limits, which may also change dynamically, may also be taken into account in making decisions about the new target quantized power consumption. In this example, the equidistant power line 650 corresponds to the new target quantized power consumption, and this may have an associated power value in mW. Note that at the current frequency, the second data point 674 is not located on the equidistant power line 650.

[0139] In this example, the new target quantized power consumption is higher than the power consumption corresponding to the second data point 774 (not shown). This is consistent with the observed utilization being higher than the utilization predicted by the previous cycle. The equidistant power lines 650 define a range of frequencies and associated utilizations, but the scalability line 660 can be used to determine what value the new operating frequency can be set to to allow the processing unit to reach or most likely achieve a power close to the new target quantized power consumption. Otherwise, if the target power is difficult to achieve without multiple trial and error implementations of setting the new operating frequency and monitoring the resulting power consumption and utilization changes, the change in utilization with frequency may be difficult to predict. In this example, the intersection of the scalability line 660 and the equidistant power line 650 gives an appropriate operating point 680 from which the new frequency can be obtained. Thus, during the frequency update process, a new frequency is allocated by determining any power change indicated as appropriate by at least the utilization change ΔU 676, setting a new target quantified power loss, and selecting a frequency to meet the new target quantified power loss using the equidistant power lines of the power model and the scalability function.

[0140] Scalability line 660 is not an equidistant power line, so the power consumption may vary for different points along the trajectory of line 660. However, it can be seen from the third data point 678 located on the equidistant power line 650 at the current frequency that the new target quantized power consumption will correspond to a higher utilization at the current frequency, and thus the new target quantized power consumption represents an increase in power consumption relative to the current power consumption. The increase in power consumption associated with achieving a new frequency on the equidistant power line 650 may depend on one or more of ΔU, ΔP, and EPP. In other examples, for example, in response to the observed utilization exceeding a threshold magnitude that is less than the expected utilization, the power consumption for the frequency update may be reduced.

[0141] Figure 7 is a table listing sequences of processing elements for implementing power management of processing units according to the present technology. This table illustrates how similar logical elements of the sequence can be applied to CPUs, GPUs, and IPUs and can be similarly extended to other processing units. The scalability equations of examples according to the present technology can be general scalability equations applicable to different types of processing units (such as CPUs, GPUs, and IPUs) or at least applicable to multiple processing units of the same type.

[0142] In this specification, the phrases "at least one of A or B" and the phrases "at least one of A and B" should be interpreted to mean any one or more of the multiple listed items A, B, etc., taken collectively and separately in any and all permutations.

[0143] In the case where the functional unit has been described as a circuit, the circuit can be a general-purpose processor circuit configured by program code to perform a specified processing function. The circuit can also be configured by modification of processing hardware. The configuration of the circuit for performing a specified function can be used entirely in hardware, entirely in software, or using a combination of hardware modification and software execution. The circuit can alternatively be firmware. Program instructions can be used to configure the logic gates of a general or special-purpose processor circuit to perform a processing function. In some examples, different elements of the circuit can be functionally combined into a single element of the circuit.

[0144] The circuit can be implemented, for example, as a hardware circuit including a processor, a microprocessor, a circuit, a circuit element (e.g., a transistor, a resistor, a capacitor, an inductor, etc.), an integrated circuit, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a digital signal processor (DSP), a field programmable gate array (FPGA), logic gates, registers, a semiconductor device, a chip, a microchip, a chipset, etc.

[0145] The processor may include a general purpose processor, a network processor that processes data transmitted over a computer network, or other types of processors including a reduced instruction set computer (RISC) or a complex instruction set computer (CISC). The processor may have a single core or a multi-core design. A multi-core processor may integrate different processor core types on the same integrated circuit die.

[0146] Machine readable program instructions may be provided on a temporary medium such as a transmission medium or on a non-temporary medium such as a storage medium. Such machine readable instructions (computer program code) may be implemented in a high-level procedural or object-oriented programming language. However, the program may be implemented in an assembly language or a machine language as desired. In any case, the language may be a compiled language or an interpreted language and combined with a hardware implementation. The machine readable instructions may be executed by a processor or an embedded controller.

[0147] Embodiments of the present invention are suitable for use with all types of semiconductor integrated circuit ("IC") chips. Examples of these IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLA), memory chips, network chips, etc. In some embodiments, one or more of the components described herein may be implemented as a system on chip (SOC) device. The SOC may include, for example, one or more central processing unit (CPU) cores, one or more graphics processing unit (GPU) cores, an input / output interface, and a memory controller. In some embodiments, the SOC and its components may be disposed on one or more integrated circuit dies, for example, packaged into a single semiconductor device.

[0148] Example

[0149] The following examples relate to further embodiments.

[0150] 1. A power management circuit for controlling a performance level of a processing unit of a processing platform, the power management circuit comprising:

[0151] measurement circuitry for measuring current utilization of the processing unit at a current operating frequency and determining any changes in utilization or power;

[0152] A frequency control circuit is configured to update the current operating frequency to a new operating frequency by determining a new target quantized power dissipation to be applied in a subsequent processing cycle based on the determination of any change in the utilization or power, and select the new operating frequency to meet the new target quantized power based on a scalability function that specifies a change in a given utilization value or a given power value as the operating frequency changes.

[0153] 2. The power management circuit may be the subject matter of Example 1 or any other example described herein, wherein the given utilization value is the currently measured utilization at the current operating frequency, and the given power value is the currently measured power at the current operating frequency.

[0154] 3. The power management circuit may be the subject matter of Example 1 or any other example described herein, wherein the frequency control circuit is configured to determine the new target quantized power consumption based at least in part on feedback corresponding to a deviation between: the actual power consumption, and a value of the target quantized power consumption predicted from a previous frequency update cycle.

[0155] 4. The power management circuit may be the subject matter of any one of Examples 1 to 3 or any other example described herein, wherein the frequency control circuit is configured to determine the new target quantized power consumption based at least in part on feedback corresponding to a deviation between: the actual utilization, and a value of the expected utilization corresponding to a previous frequency update period calculated by applying the scalability function.

[0156] 5. The power management circuit may be the subject matter of any of Examples 1 to 3 or any other example described herein, wherein the measurement circuit is used to perform a determination of any change in utilization based on a difference between: the measured current utilization or power, and the expected utilization or power fed back from a previous frequency update cycle.

[0157] 6. The power management circuit may be the subject matter of any one of Examples 1 to 5 or any other example described herein, wherein, when the measurement circuit detects a change in the utilization or the power, the frequency control circuit is configured to update the current operating frequency to a new operating frequency based on a comparison of the magnitude of the detected change in utilization or power with a corresponding threshold magnitude.

[0158] 7. The power management circuit may be the subject matter of any one of Examples 1 to 6 or any other example described herein, wherein the frequency control circuit is used to update the current operating frequency to a new operating frequency in response to a change in a system parameter of the processing platform.

[0159] 8. The power management circuit may be the subject matter of any one of Examples 1 to 7 or any other example described herein, wherein the frequency control circuit is configured to update the current operating frequency to a new operating frequency according to a change in a power limit assigned to the processing unit PU by the processing platform, the PU power limit representing a portion of a system power limit.

[0160] 9. The power management circuit may be the subject matter of any one of Examples 1 to 8 or any other example described herein, wherein the target quantized power consumption depends on an energy performance preference, such that a higher target quantized power consumption is set when the energy performance preference indicates that performance is to be optimized over power efficiency, and a relatively lower target quantized power consumption is set when the energy performance preference indicates that power efficiency is to be optimized over performance.

[0161] 10. The power management circuit may be the subject matter of any of Examples 1 to 9 or any other example described herein, wherein the frequency control circuit is used to determine the new operating frequency using a power distribution map for the processing unit, wherein the power distribution map defines an a priori relationship between the frequency, utilization and power consumption of the processing unit.

[0162] 11. The power management circuit may be the subject matter of example 10 or any other example described herein, wherein the new target quantized power loss corresponds to a point on an equidistant power line of the power profile.

[0163] 12. The power management circuit may be the subject matter of Example 11 or any other example described herein, wherein the frequency control circuit is used to select the new operating frequency based on the intersection of the equidistant power lines and a scalability line corresponding to the scalability function applied in the load-frequency plane.

[0164] 13. The power management circuit may be the subject matter of any of Examples 10 to 12 or any other example described herein, wherein the power distribution map is generated prior to runtime of the processing unit by at least one of the following steps: performing a computer simulation of a model of the processing unit; monitoring the operation of the processing unit when executing one or more real processing workloads; and monitoring the operation of a processing unit having similar characteristics to the processing unit when executing one or more real processing workloads.

[0165] 14. The power management circuit may be the subject matter of any of Examples 1 to 13 or any other example described herein, wherein the target quantified power dissipation is quantified in Watts.

[0166] 15. A processing platform, the processing platform comprising:

[0167] Two or more processing units which may be the subject matter of any of Examples 1 to 14 or any other example described herein; and

[0168] a platform power management circuit, the platform power management circuit being used to control the allocation of system power to the plurality of processing units;

[0169] The platform power management circuit is arranged to receive a corresponding new target quantized power consumption from each of the processing units, and control one or more system parameters according to the received new target quantized power consumption.

[0170] 16. The processing platform may be the subject matter of example 15 or any other example described herein, wherein the platform power management circuit is configured to determine a performance per watt for the processing platform based on a plurality of the received new target quantized power consumptions.

[0171] 17. The processing platform may be the subject matter of Example 15 or Example 16 or any other example described herein, wherein the two or more processing units include at least a subset of: a processor core, a multi-core processor, a graphics processing unit, and an image processing unit.

[0172] 18. The processing platform may be the subject matter of any of Examples 15 to 17 or any other example described herein, wherein at least a subset of the two or more processing units are configured to receive an allocation of a portion of system power usable by the processing units from the platform power management circuitry, and wherein frequency control circuitry of the respective processing units is configured to determine the new target quantized power consumption based on the allocated portion of the system power.

[0173] 19. Machine readable instructions provided on at least one tangible or non-tangible machine readable medium, which, when executed by a processing unit of a processing platform, cause processing hardware to:

[0174] measuring current utilization of the processing unit at the current operating frequency and determining any changes in utilization or power;

[0175] updating the current operating frequency to a new operating frequency by determining a new target quantized power dissipation to be applied in a subsequent processing cycle based on the calculated change in utilization or power, and selecting the new operating frequency to meet the new target quantized power based on a scalability function that specifies a change in a given utilization or a given power with the operating frequency; and

[0176] The new operating frequency is assigned to the processing unit.

[0177] 20. The machine-readable instructions may be the subject matter of Example 19 or any other example described herein, comprising an interface module for interfacing with an operating system of the processing platform to receive at least one platform-controlled power limit from the processing platform to constrain the new target quantized power consumption.

[0178] 21. The machine-readable instructions may be the subject matter of Example 19 or Example 20 or any other example described herein, wherein the interface module is used to receive an energy performance preference from the processing platform, and wherein the new target quantized power consumption is determined at least in part based on the energy performance preference.

[0179] 22. The machine readable instructions may be the subject matter of any of Examples 19 to 21 or any other example described herein, wherein the interface module is configured to output the determined new target quantized power consumption to the platform operating system.

[0180] 23. A method for controlling a performance level of a processing unit of a processing platform, the method comprising:

[0181] measuring current utilization of the processing unit at the current operating frequency and determining any changes in utilization or power;

[0182] updating the current operating frequency to a new operating frequency by determining a new target quantized power consumption to be applied in a subsequent processing cycle based on the calculated change in utilization or power; and

[0183] A new operating frequency is selected to meet the new target quantized power based on a scalability function that specifies a change in the operating frequency for a given utilization or a given power.

[0184] 24. The method of Example 23 or any other example described herein, comprising determining the new target quantized power consumption based at least in part on feedback corresponding to a deviation between: the measured current utilization, the value of the new target quantized power consumption determined in a previous frequency update cycle.

[0185] 25. An apparatus for controlling a performance level of a processing unit of a processing platform, the apparatus for controlling comprising:

[0186] means for measuring the current utilization of the processing unit at the current operating frequency and calculating a change in utilization or power;

[0187] means for updating the current operating frequency to a new operating frequency by determining a new target quantized power consumption to be applied in a subsequent processing cycle based on the calculated change in utilization or power; and

[0188] Means for selecting a new operating frequency to meet the new target quantized power based on a scalability function, the scalability function specifying a change in a given utilization or a given power with the operating frequency.

[0189] 26. The control device may be the subject matter of Example 25 or any other example described herein, wherein the device for measuring is used to determine a change in the utilization based on a difference between the measured current utilization and an expected utilization fed back from a previous frequency update cycle, the expected utilization having been determined using the scalability function.

Claims

1. A power management circuit for controlling a performance level of a processing unit of a processing platform, the power management circuit comprising: a measurement circuit for measuring a current utilization or power of the processing unit at a current operating frequency and determining any changes in the utilization or power; A frequency control circuit is configured to update the current operating frequency to a new operating frequency by determining a new target quantized power dissipation to be applied in a subsequent processing cycle based on the determination of any change in the utilization or power, and select the new operating frequency to meet the new target quantized power dissipation based on a scalability function that specifies a change in a given utilization value or a given power value as the operating frequency changes.

2. The power management circuit according to claim 1, wherein: The given utilization value is a currently measured utilization at the current operating frequency, and the given power value is a currently measured power at the current operating frequency.

3. The power management circuit according to claim 1, wherein: The frequency control circuit is configured to determine the new target quantized power consumption based at least in part on feedback corresponding to a deviation between: the actual power consumption and a value of the target quantized power consumption predicted from a previous frequency update period.

4. The power management circuit according to claim 1, wherein: The frequency control circuit is configured to determine the new target quantized power consumption based at least in part on feedback corresponding to a deviation between an actual utilization and a value of an expected utilization corresponding to a previous frequency update period calculated by applying the scalability function.

5. The power management circuit according to claim 1, wherein: The measurement circuit is used to perform a determination of any change in utilization based on a difference between: the measured current utilization or power, and the expected utilization or power fed back from a previous frequency update cycle.

6. The power management circuit according to claim 1, wherein: When the measurement circuit detects a change in the utilization rate or the power, the frequency control circuit is configured to update the current operating frequency to a new operating frequency according to a comparison between a magnitude of the detected change in utilization rate or power and a corresponding threshold magnitude.

7. The power management circuit according to claim 1, wherein: The frequency control circuit is used to update the current operating frequency to a new operating frequency in response to a change in a system parameter of the processing platform.

8. The power management circuit according to claim 1, wherein: The frequency control circuit is used to update the current operating frequency to a new operating frequency according to a change of a power limit assigned to the processing unit by the processing platform, wherein the processing unit power limit represents a portion of a system power limit.

9. The power management circuit according to claim 1, wherein: The target quantized power consumption depends on the energy performance preference, so that a higher target quantized power consumption is set when the energy performance preference indicates that performance should be optimized over power efficiency, and a relatively lower target quantized power consumption is set when the energy performance preference indicates that power efficiency should be optimized over performance.

10. The power management circuit according to claim 1, wherein: The frequency control circuit is configured to determine the new operating frequency using a power profile for the processing unit, wherein the power profile defines an a priori relationship between operating frequency, utilization, and power of the processing unit.