Autonomously managing core cluster frequency using performance statistics in processor device

By generating performance and energy consumption models in the processor device and adjusting the core cluster frequency based on the advantage model, the problem that conventional power management circuits cannot consider QoS is solved, and efficient frequency management and power optimization are achieved.

CN121844276APending Publication Date: 2026-04-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

The power management circuitry of conventional processor devices cannot take into account the Quality of Service (QoS) requirements of the processor core workload when making power management decisions.

Method used

The cluster power management circuit collects AMU statistics data from the processor core, generates a performance model and a per-instruction power consumption model, identifies the target frequency operating point based on the advantage model, and adjusts the frequency of the core cluster through the DVFS circuit.

Benefits of technology

It enables efficient management of processor core cluster frequencies while taking QoS requirements into account, improving power efficiency and responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121844276A_ABST
    Figure CN121844276A_ABST
Patent Text Reader

Abstract

The invention discloses autonomously managing core cluster frequency in a processor device using performance statistics. In some aspects, a cluster power management circuit of a processor device collects activity management unit (AMU) statistics for a plurality of processor cores for each of one or more frequency operating points within a time interval. Based on the AMU statistics, the cluster power management circuitry generates a performance model representing processor performance as a function of frequency, and generates an per instruction energy consumption (EI) model representing per instruction energy consumption as a function of frequency using the performance model and the power consumption measurements. The cluster power management circuit then generates a dominant model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency, and identifies a target frequency operating point based on the dominant model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority application

[0002] This application claims priority to U.S. Patent Application Serial No. 18 / 468,242, filed on September 15, 2023, entitled “AUTONOMOUSLY MANAGING CORECLUSTER FREQUENCIES USING PERFORMANCE STATISTICS IN PROCESSOR DEVICES”, the entire contents of which are incorporated herein by reference. background Technical Field

[0003] The technology disclosed herein relates generally to power and performance management in multi-core processor-based devices, and more specifically to frequency management of clusters of processor cores for processor devices. Background Technology

[0004] Conventional processor devices implement the Advanced Configuration and Power Interface (ACPI) specification, which defines an open industry standard encompassing power management across the processor device's hardware, operating system (OS), and application software. Version 5.0 of the ACPI specification includes a feature known as Cooperative Processor Performance Control (CPPC). CPPC provides an interface through which the OS, running on the processor device, can manage the processor device's performance by sending performance change requests to, for example, the processor device's power management circuitry. The power management circuitry determines whether the requested performance level is supported, and if so, updates the processor device's operating frequency and voltage. In this way, performance cues provided by the OS enable better power efficiency and more responsive frequency scaling for the processor device.

[0005] As part of providing power management functionality, the processor device's power management circuitry uses the processor device's Activity Management Unit (AMU) to monitor processor performance. The AMU consists of multiple counters implemented as system registers to track the occurrence of various processor events. These events may include, for example, architectural events (such as processor cycles, retirement instructions, memory pause cycles, etc.) and ancillary events (such as last-level cache (LLC) request misses, LLC request accesses, bus accesses, etc.).

[0006] The power management circuitry of a processor device can use statistics collected from the AMU to determine whether the performance level requested by the OS is supported at a given time. However, the power management circuitry of a conventional processor device may not be able to consider the Quality of Service (QoS) requirements of the workload when making power management decisions. Summary of the Invention

[0007] The aspects disclosed in the detailed description include autonomously managing core cluster frequencies using performance statistics from a processor device. Related apparatus, methods, and computer-readable media are also disclosed. In this regard, the processor device provides multiple core clusters, each core cluster including multiple processor cores and cluster power management circuitry. Each processor core includes an activity management unit (AMU) configured to collect AMU statistics on various aspects of processor core performance. The cluster power management circuitry is configured to collect AMU statistics from the multiple AMUs for each of one or more frequency operating points within a time interval. The cluster power management circuitry then generates a performance model based on the multiple AMU statistics, expressing processor performance (by way of a non-limiting example, measured as average instructions per clock cycle (IPC)) as a function of frequency. Some aspects may include: generating the performance model includes calculating a performance value for each of the one or more frequency operating points, the performance value comprising the frequency operating point multiplied by the average number of instructions per clock cycle at that frequency operating point.

[0008] The cluster power management circuitry further generates a per-instruction energy consumption (EI) model based on a performance model and power consumption measurements. This per-instruction energy consumption (EI) model represents per-instruction energy consumption as a function of frequency. In some aspects, generating the EI model may include calculating an EI value for each of one or more frequency operating points, the EI value comprising the quotient of power consumption at that frequency operating point divided by the performance value at that frequency operating point.

[0009] The cluster power management circuitry then generates an advantage model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency (e.g., by calculating the quotient of the first rate of change of the performance model as a function of frequency divided by the second rate of change of the EI model as a function of frequency). The cluster power management circuitry identifies a target frequency operating point based on the advantage model and sends the target frequency operating point to the processor device's Dynamic Voltage and Frequency Scaling (DVFS) circuitry. According to some aspects, identifying the target frequency operating point may include determining the maximum frequency operating point corresponding to an advantage value indicated by the advantage model as greater than or equal to a threshold. Some such aspects may include: determining the maximum frequency operating point includes determining the maximum frequency operating point as an interpolation between the energy balance frequency operating point and the performance frequency operating point based on an Energy Performance Preference (EPP) cue.

[0010] Several aspects are available: Once the DVFS circuitry receives the target frequency operating point from the cluster power management circuitry, the DVFS circuitry sets the core cluster frequency based on the target frequency operating point. Depending on whether the target frequency operating point is higher than the current frequency of the core cluster, setting the core cluster frequency may include immediately increasing the core cluster frequency. Depending on whether the target frequency operating point is lower than the current frequency of the core cluster, setting the core cluster frequency may include incrementally decreasing the core cluster frequency.

[0011] On the other hand, a processor device is provided. The processor device includes a core cluster comprising multiple processor cores, each processor core including a corresponding multiple AMUs. The processor device also includes cluster power management circuitry and DVFS circuitry. The cluster power management circuitry is configured to collect multiple AMU statistics from the multiple AMUs for each of one or more frequency operating points within a time interval. The cluster power management circuitry is further configured to generate a performance model expressing processor performance as a function of frequency based on the multiple AMU statistics. The cluster power management circuitry is also configured to generate an EI model expressing per-instruction energy consumption as a function of frequency based on the performance model and power consumption measurements. The cluster power management circuitry is additionally configured to generate a dominance model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency. The cluster power management circuitry is further configured to identify a target frequency operating point based on the dominance model. The cluster power management circuitry is also configured to send the target frequency operating point to the DVFS circuitry.

[0012] In another aspect, a processor device is provided. The processor device includes components for collecting multiple AMU statistics from multiple AMUs of multiple processor cores corresponding to a core cluster of the processor device for each of one or more frequency operating points within a time interval. The processor device further includes components for generating a performance model expressing processor performance as a function of frequency based on the multiple AMU statistics. The processor device also includes components for generating an Energy per Instruction (EI) model expressing energy consumption per instruction as a function of frequency based on the performance model and power consumption measurements. The processor device additionally includes components for generating a dominance model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency. The processor device further includes components for identifying a target frequency operating point based on the dominance model. The processor device also includes components for transmitting the target frequency operating point to the processor device's DVFS circuitry.

[0013] On the other hand, a method is provided for autonomously managing core cluster frequency using performance statistics. The method includes having cluster power management circuitry of a processor device collect multiple AMU statistics from multiple AMUs of multiple processor cores corresponding to a core cluster of the processor device for each of one or more frequency operating points within a time interval. The method further includes having cluster power management circuitry generate a performance model representing processor performance as a function of frequency based on the multiple AMU statistics. The method also includes having cluster power management circuitry generate an Energy per Instruction (EI) model representing power consumption per instruction as a function of frequency based on the performance model and power consumption measurements. The method additionally includes having cluster power management circuitry generate a dominance model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency. The method further includes having cluster power management circuitry identify a target frequency operating point based on the dominance model. The method also includes having cluster power management circuitry send the target frequency operating point to the DVFS circuitry of the processor device.

[0014] In another aspect, a non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium stores computer-executable instructions that, when executed, cause a processor device of a processor-based device to collect multiple AMU statistics from multiple AMUs of multiple processor cores corresponding to a core cluster of the processor device for each of one or more frequency operating points within a time interval. The computer-executable instructions further cause the processor device to generate a performance model expressing processor performance as a function of frequency based on the multiple AMU statistics. The computer-executable instructions also cause the processor device to generate an Energy-in-Instruction (EI) model expressing per-instruction energy consumption as a function of frequency based on the performance model and power consumption measurements. The computer-executable instructions additionally cause the processor device to generate a dominance model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency. The computer-executable instructions further cause the processor device to identify a target frequency operating point based on the dominance model. The computer-executable instructions also cause the processor device to send the target frequency operating point to the processor device's DVFS circuitry. Attached Figure Description

[0015] Figure 1 It is a block diagram of an exemplary processor-based device based on some aspects, including a core cluster, which includes cluster power management circuitry configured to autonomously manage the core cluster frequency using performance statistics.

[0016] Figure 2 This is an example based on some aspects. Figure 1 A diagram illustrating an exemplary performance model generated by the cluster power management circuitry;

[0017] Figure 3 This is an example based on some aspects. Figure 1 The cluster power management circuit is based on Figure 2 A diagram illustrating an exemplary per-instruction energy consumption (EI) model generated from the performance model;

[0018] Figure 4 This is an example based on some aspects. Figure 1 The cluster power management circuit is based on Figure 2 Performance model and Figure 3 A diagram illustrating an exemplary advantage model generated by the EI model;

[0019] Figures 5A to 5C Examples are provided based on some aspects. Figure 1 A flowchart illustrating exemplary operations performed by the processor device to autonomously manage the core cluster frequency using performance statistics; and

[0020] Figure 6 Yes, it can include Figure 1 A block diagram of an exemplary processor-based device. Detailed Implementation

[0021] Several exemplary aspects of this disclosure will now be described with reference to the accompanying drawings. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or superior to other aspects. The terms “first,” “second,” etc., are used herein to distinguish similarly named elements and should not be construed as indicating an order relationship between such elements unless expressly stated herein.

[0022] The aspects disclosed in the detailed description include autonomously managing core cluster frequencies using performance statistics from a processor device. Related apparatus, methods, and computer-readable media are also disclosed. In this regard, the processor device provides multiple core clusters, each core cluster including multiple processor cores and cluster power management circuitry. Each processor core includes an activity management unit (AMU) configured to collect AMU statistics on various aspects of processor core performance. The cluster power management circuitry is configured to collect AMU statistics from the multiple AMUs for each of one or more frequency operating points within a time interval. The cluster power management circuitry then generates a performance model based on the multiple AMU statistics, expressing processor performance (by way of a non-limiting example, measured as average instructions per clock cycle (IPC)) as a function of frequency. Some aspects may include: generating the performance model includes calculating a performance value for each of the one or more frequency operating points, the performance value comprising the frequency operating point multiplied by the average number of instructions per clock cycle at that frequency operating point.

[0023] The cluster power management circuitry further generates a per-instruction energy consumption (EI) model based on a performance model and power consumption measurements. This per-instruction energy consumption (EI) model represents per-instruction energy consumption as a function of frequency. In some aspects, generating the EI model may include calculating an EI value for each of one or more frequency operating points, the EI value comprising the quotient of power consumption at that frequency operating point divided by the performance value at that frequency operating point.

[0024] The cluster power management circuitry then generates an advantage model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency (e.g., by calculating the quotient of the first rate of change of the performance model as a function of frequency divided by the second rate of change of the EI model as a function of frequency). The cluster power management circuitry identifies a target frequency operating point based on the advantage model and sends the target frequency operating point to the processor device's Dynamic Voltage and Frequency Scaling (DVFS) circuitry. According to some aspects, identifying the target frequency operating point may include determining the maximum frequency operating point corresponding to an advantage value indicated by the advantage model as greater than or equal to a threshold. Some such aspects may include: determining the maximum frequency operating point includes determining the maximum frequency operating point as an interpolation between the energy balance frequency operating point and the performance frequency operating point based on an Energy Performance Preference (EPP) cue.

[0025] Several aspects are available: Once the DVFS circuitry receives the target frequency operating point from the cluster power management circuitry, the DVFS circuitry sets the core cluster frequency based on the target frequency operating point. Depending on whether the target frequency operating point is higher than the current frequency of the core cluster, setting the core cluster frequency may include immediately increasing the core cluster frequency. Depending on whether the target frequency operating point is lower than the current frequency of the core cluster, setting the core cluster frequency may include incrementally decreasing the core cluster frequency.

[0026] In this respect, Figure 1 This is a block diagram of an exemplary processor device 100 (also referred to as a “processor” or “CPU”). The processor device 100 may include an ordered or out-of-order processor (OoP), and / or may be one of a plurality of processor devices 100. Examples of the processor device 100 may include, but are not limited to, a digital signal processor (DSP), a general-purpose microprocessor, an application-specific integrated circuit (ASIC), a field-programmable logic array (FPGA), or other equivalent integrated or discrete logic circuits.

[0027] like Figure 1As seen, the processor device 100 includes multiple core clusters 102(0)-102(X), each of which includes multiple processor cores, such as processor cores 104(0)-104(C) of core cluster 102(0). Figure 1 The processor device 100 also includes a graphics processing unit (GPU) 106 for performing graphics operations. As a non-limiting example, the GPU 106 may include a dedicated hardware unit with fixed functionality and programmable components for rendering graphics and executing GPU applications. The GPU 106 may also include a DSP, a general-purpose microprocessor, an ASIC, an FPGA, or other equivalent integrated or discrete logic circuitry; for clarity, in... Figure 1 Not shown in the image.

[0028] Figure 1 The processor device 100 further includes additional exemplary components, including an artificial intelligence (AI) engine 108, mobile device management (MDM) circuitry 110, power management circuitry 112, on-chip network (NoC) 114, and memory device 116. As a non-limiting example, the AI ​​engine 108 of the processor device 100 includes circuitry and logic components for providing AI-based functionality, such as search, speech recognition, text and / or image generation, etc. The MDM circuitry 110 provides functionality for supplying, configuring, updating, and / or protecting mobile devices to which the processor device 100 is integrated. The power management circuitry 112 provides advanced performance and power management functionality for the processor device 100 as a whole, while the NoC 114 is configured to manage communication between different devices including the processor device 100. Finally, the memory device 116 provides storage and access to data used by the processor device 100, and in some aspects, as a non-limiting example, may include a double data rate (DDR) synchronous dynamic random access memory (SDRAM) device.

[0029] Figure 1 Exemplary elements of core cluster 102(0) are illustrated in more detail. The processor cores 104(0)-104(C) of core cluster 102(0) are communicatively coupled to a last-level cache (LLC) 118 that stores frequently accessed data for faster access, and are communicatively coupled to a phase-locked loop (PLL) 120 that provides clock signals to the processor cores 104(0)-104(C) and LLC 118. The frequency and voltage of operation of the processor cores 104(0)-104(C) of core cluster 102(0) are controlled by a DVFS circuit 122, and the performance and power management of core cluster 102(0) are handled by a cluster power management circuit 124. It should be understood that, although Figure 1Only exemplary elements of core cluster 102(0) are shown, but each core cluster in core clusters 102(0)-102(X) includes elements corresponding to the illustrated elements of core cluster 102(0). Because processor cores 104(0)-104(C) and LLC 118 all operate using the same clock signal generated by PLL 120, core cluster 102(0) is considered a “synchronous” core cluster.

[0030] exist Figure 1 In the example, processor device 100 is configured to provide performance and power management functionalities, such as those defined by Cooperative Processor Performance Control (CPPC) of version 5 of the Advanced Configuration and Power Interface (ACPI) specification. To support performance monitoring for such performance and power management functionalities, each processor core in processor cores 104(0)-104(C) of core cluster 102(0) includes a corresponding AMU 126(0)-126(C). It should be understood that, although... Figure 1 Only a single AMU 126(0)-126(C) for each of the processor cores 104(0)-104(C) is shown, but in some respects, each of the processor cores 104(0)-104(C) may include multiple AMUs 126(0)-126(C).

[0031] The AMU 126(0)-126(C) consists of individual system registers (not shown), each corresponding to a different architecture event or auxiliary event, and configured to store a counter incremented by the AMU 126(0)-126(C) when the corresponding event occurs. The counter value stored in the AMU 126(0)-126(C) is... Figure 1 The data are collectively referred to herein as “AMU Statistics 128(0)-128(C)” or “Performance Statistics 128(0)-128(C)”. AMU Statistics 128(0)-128(C) may include any data collected by AMU 126(0)-126(C) regarding the occurrence of architectural events or ancillary events related to various aspects of the performance of processor cores 104(0)-104(C).

[0032] Table 1 below lists exemplary architectures and ancillary events that may include AMU statistics 128(0)-128(C):

[0033] Table 1

[0034]

[0035] Figure 1The processor device 100 may encompass any of known digital logic elements, semiconductor circuits, processing cores, and / or memory structures, as well as other elements or combinations thereof. The aspects described herein are not limited to any particular arrangement of elements, and the disclosed techniques can be readily extended to various structures and layouts on semiconductor dies or packages. It should be understood that some aspects of the processor device 100 and / or processor cores 104(0)-104(C) may include, in addition to… Figure 1 Elements other than those exemplified, and / or may include elements that are more than those listed. Figure 1 More or fewer components are illustrated. For example, processor device 100 may further include cache, controller, communication bus and / or persistent storage devices, which, for clarity, are... Figure 1 The middle part is omitted.

[0036] As mentioned above, the cluster power management circuit 124 may use AMU statistics 128(0)-128(C) collected from AMUs 126(0)-126(C) to determine whether the performance level requested by the operating system (OS) executed by the processor device 100 is supported at a given time. However, conventional cluster power management circuitry may not be able to take into account the Quality of Service (QoS) requirements of the workloads executed by the processor cores 104(0)-104(C) when making power management decisions for the core cluster 102(0).

[0037] In this respect, the cluster power management circuit 124 is configured to use performance statistics (i.e., Figure 1 The AMU statistics 128(0)-128(C) autonomously manage the core cluster frequency. In exemplary operation, the cluster power management circuitry 124 of the processor device 100 collects multiple AMU statistics 128(0)-128(C) from multiple AMUs 126(0)-126(C) for each of one or more frequency operating points (not shown) within a time interval. Each frequency operating point represents the processor frequency at the corresponding voltage at which the core cluster 102(0) operates during the time interval. In some respects, the time interval may be programmed by the OS or application software.

[0038] Based on AMU statistics 128(0)-128(C), cluster power management circuit 124 generates a performance model (not shown) that expresses processor performance as a function of frequency. Cluster power management circuit 124 also generates an EI model (not shown) that expresses per-instruction power consumption as a function of frequency based on the performance model and power consumption measurement 130, the power consumption measurement indicating the power consumed at the frequency and voltage used during the time interval. Finally, cluster power management circuit 124 generates an advantage model (not shown) based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency. In some respects, the generation of the exemplary performance model, exemplary EI model, and exemplary advantage model are discussed below with respect to... Figure 2 , Figure 3 and Figure 4 Let's discuss this in more detail.

[0039] Cluster power management circuitry 124 then identifies a target frequency operating point (not shown) based on a dominance model. This may include, for example, determining the maximum frequency operating point corresponding to a dominance value indicated by the dominance model as greater than or equal to a threshold corresponding to a QoS cues from the OS. As used herein, a QoS cues refer to an indication that a workload executed on a processor core from the OS to the corresponding processor core of processor cores 104(0)-104(C) requires a specific QoS level. After identifying the target frequency operating point, cluster power management circuitry 124 sends the target frequency operating point to DVFS circuitry 122, which can then set the frequency of core cluster 102(0) based on the target frequency operating point. In aspects where the target frequency operating point is higher than the current frequency of core cluster 102(0), DVFS circuitry 122 may be configured to set the frequency of core cluster 102(0) by immediately increasing the frequency of core cluster 102(0). Some aspects of the target frequency operating point being lower than the current frequency of the core cluster 102(0) may provide that the DVFS circuit 122 is configured to set the frequency of the core cluster 102(0) by incrementally reducing the frequency of the core cluster 102(0) (i.e., by gradually reducing the frequency of the core cluster 102(0) over multiple time intervals rather than immediately reducing the frequency of the core cluster 102(0)).

[0040] As mentioned above, when performing autonomous frequency management of the core cluster 102(0), the cluster power management circuit 124 uses AMU statistics 128(0)-128(C) to generate a performance model that expresses processor performance as a function of frequency. In this respect, Figure 2 An exemplary performance model 200 is illustrated according to some aspects. Performance model 200 includes a curve 202 generated to fit a plurality of performance values ​​204(0)-204(F), which are derived from... Figure 1The cluster power management circuit 124 operates for multiple corresponding frequency points during the time interval (in Figure 2 The calculation is performed for each frequency operating point in (FREQ OP PT) 206(0)-206(F). Figure 2 As seen, the frequency operating points 206(0)-206(F) are represented by vertical dashed lines arranged along the horizontal frequency axis, while the performance values ​​204(0)-204(F) are represented as points on each vertical dashed line along the vertical performance axis.

[0041] exist Figure 2 In the example, each performance value in performance values ​​204(0)-204(F) indicates the average instruction per clock cycle (IPC) value at the corresponding frequency operating points 206(0)-206(F). In some aspects, the cluster power management circuitry 124 can use Figure 1 The performance values ​​204(0)-204(F) are calculated using the AMU statistics 128(0)-128(C), which include the architectural events and ancillary events referenced in Table 1 above. For example, the performance values ​​204(0)-204(F) can be calculated using the formula illustrated in Table 2 below, which may involve the architectural events and / or ancillary events described in Table 1 above (e.g., BUS_ACCESS, INSTR_RETIRED, etc.):

[0042] Table 2

[0043]

[0044] After calculating the IPC for each frequency operating point in frequency operating points 206(0)-206(F), the cluster power management circuit 124 can use the formula The performance values ​​204(0)-204(F) are calculated, where f is equal to the frequency at each of the frequency operating points 206(0)-206(F). Therefore, each performance value in the performance values ​​204(0)-204(F) consists of the product of the corresponding frequency operating point 206(0)-206(F) and the average number of instructions per clock cycle at the frequency operating point 206(0)-206(F).

[0045] Figure 3 An exemplary EI model 300 is illustrated, which can be derived from... Figure 1 Cluster power management circuit 124 based Figure 2 Performance Model 200 and Figure 1 Power consumption measurement 130 (wherein power consumption measurement 130 indicates the power consumed at the frequency and voltage used at a given frequency operating point) is generated to express per-command energy as a function of frequency. For example... Figure 3 As seen, the EI model 300 includes a curve 302, which is generated to fit a plurality of EI values ​​304(0)-304(F), which are determined by the cluster power management circuit 124 during time intervals. Figure 2 Frequency operating point (at Figure 3 The frequency operating points 206(0)-206(F) are calculated and specified as “FREQOP PT”. Figure 3 The values ​​are represented by vertical dashed lines arranged along the horizontal frequency axis, while the EI values ​​304(0)-304(F) are represented as points on each vertical dashed line along the vertical per-instruction energy consumption axis.

[0046] Cluster power management circuit 124 uses the formula EI values ​​304(0)-304(F) are calculated to generate EI model 300. Power (V, f) corresponds to the power consumption measurement 130 of voltage V and power consumption at frequency f at each frequency operating point 206(0)-206(F), while Perf(F) corresponds to one of the performance values ​​204(0)-204(F) at frequency f at each frequency operating point 206(0)-206(F). Therefore, each EI value in EI values ​​304(0)-304(F) includes the quotient of the power consumption at the corresponding frequency operating point 206(0)-206(F) divided by the performance value 204(0)-204(F) at the corresponding frequency operating point 206(0)-206(F).

[0047] Figure 4 It shows that it can be made by Figure 1 The cluster power management circuit 124 uses Figure 2 Performance Model 200 and Figure 3 An exemplary advantage model 400 is generated from the EI model 300. The advantage model 400 includes a curve 402 based on a first rate of change of the performance model 200 as a function of frequency and a second rate of change of the EI model 300 as a function of frequency. The advantage model 400 can be used by the cluster power management circuitry 124 using a formula advantage. To calculate. Therefore, in some respects, the advantage model 400 involves the cluster power management circuit 124 calculating the quotient of the first rate of change of the performance model 200 as a function of frequency divided by the second rate of change of the EI model 300 as a function of frequency.

[0048] Using the Advantage Model 400, the cluster power management circuit 124 identifies the target frequency operating point (in...). Figure 4The target frequency operating point 404 is defined as “TARGET FREQ OP PT”. The target frequency operating point 404 represents the frequency at which the corresponding increase in normalized performance per instruction power consumption is acceptable for a given QoS cues provided by the OS. The acceptable level of advantage for a given QoS cues is indicated by a threshold 406. Therefore, the target frequency operating point 404 can be calculated as the maximum frequency operating point corresponding to an advantage value 408, indicated by the advantage model 400, that is greater than or equal to the threshold 406. It should be understood that, although... Figure 4 Only one threshold 406 is shown, but some aspects may provide multiple thresholds 406, each corresponding to a different QoS cues and indicating different levels of advantage suitable for that QoS cues. The target frequency operating point 404 is... Figure 4 The value 408 is illustrated by a vertical dashed line arranged along the horizontal frequency axis, while the dominant value 408 is illustrated as a point on the vertical dashed line along the vertical dominant axis.

[0049] In some aspects, the OS may further provide an EPP hint (e.g., a value from 0 to 15, as a non-limiting example) to indicate whether the OS or application software prefers a bias toward performance or energy efficiency. In such aspects, the cluster power management circuit 124 may use a first threshold to calculate the target frequency operating point F_EB corresponding to the energy balance point, and may also use a second threshold to calculate the target frequency operating point F_PERF corresponding to the performance point. The cluster power management circuit 124 then uses the logic shown in Table 3 below to calculate the frequency of the core cluster 102(0):

[0050] Table 3

[0051]

[0052] To illustrate, based on some aspects, Figure 1 The processor device 100 performs exemplary operations for autonomously managing the core cluster frequency using performance statistics. Figures 5A to 5C A flowchart illustrating exemplary operation 500 is provided. For clarity, in the description... Figures 5A to 5C When quoting Figures 1 to 4 The components. It should be understood that, in some respects, some exemplary operations in exemplary operation 500 may be performed in a different order than that illustrated herein, and / or may be omitted.

[0053] Example operation 500 in Figure 5A Starting in the middle, the cluster power management circuitry of the processor device (e.g., Figure 1 The cluster power management circuit 124 of the processor device 100 operates at one or more frequency points (such as...) within a time interval. Figure 2 and Figure 3Each frequency operating point in frequency operating points 206(0)-206(F) is selected from multiple processor cores (e.g., corresponding to the core cluster of processor device 100) of the core cluster. Figure 1 The core cluster 102(0) has multiple AMUs (processor cores 104(0)-104(C)) that collect multiple AMU statistics (such as data from...). Figure 1 AMU statistics 128(0)-128(C) of AMU 126(0)-126(C) (Box 502). Cluster power management circuit 124 generates a performance model that expresses processor performance as a function of frequency based on multiple AMU statistics 128(0)-128(C) (e.g., Figure 2 The performance model 200 (box 504). In some aspects, the operations of box 504 for generating the performance model 200 may include calculating performance values ​​(such as...) for each of one or more frequency operating points 206(0)-206(F). Figure 2 The performance values ​​are 204(0)-204(F), which include the frequency operating point multiplied by the average number of instructions per clock cycle at that frequency operating point (box 506).

[0054] The cluster power management circuit 124 is also based on performance model 200 and power consumption measurements (such as... Figure 1 The power consumption measurement 130) generates an EI model that expresses per-instruction power consumption as a function of frequency (e.g., Figure 3 EI model 300 (box 508). Several aspects are provided: the operations of box 508 for generating EI model 300 include calculating EI values ​​(such as...) for each of one or more frequency operating points 206(0)-206(F). Figure 3 The EI value is 304(0)-304(F), which includes the quotient of power consumption at the frequency operating point divided by the performance value at that frequency operating point (box 510). Exemplary operation 500 in Figure 5B Continue at frame 512.

[0055] Now refer to Figure 5B The cluster power management circuit 124 then generates an advantage model (e.g., based on the first rate of change of performance model 200 as a function of frequency and the second rate of change of EI model 300 as a function of frequency). Figure 4 The dominant model 400 (box 512). According to some aspects, the operation of box 512 for generating the dominant model 400 includes calculating the quotient of a first rate of change of the performance model 200 as a function of frequency divided by a second rate of change of the EI model 300 as a function of frequency (box 514).

[0056] The cluster power management circuit 124 then identifies the target frequency operating point (such as) based on the advantage model 400. Figure 4 The target frequency operating point 404 (box 516). In some aspects, the operation of box 516 for identifying the target frequency operating point 404 may include determining the value corresponding to a threshold indicated by the dominance model 400 as greater than or equal to a threshold (e.g., Figure 4 The advantage value of the threshold 406) (e.g., Figure 4 The maximum frequency operating point (box 518) is determined by the optimal value 408. Some aspects of this include: the operation of box 518 for determining the maximum frequency operating point includes determining the maximum frequency operating point based on the EPP cue as an interpolation between the energy-balanced frequency operating point (i.e., the frequency operating point selected to balance power consumption and processor performance) and the performance frequency operating point (i.e., the frequency operating point selected to prioritize processor performance) (box 520). After identifying the target frequency operating point 404, the cluster power management circuitry 124 sends the target frequency operating point 404 to the DVFS circuitry of the processor device 100 (such as...). Figure 1 The DVFS circuit 122 (box 522). In some aspects, exemplary operation 500 then in Figure 5C Continue at frame 524.

[0057] Now go to Figure 5C In some aspects, the DVFS circuit 122 receives a target frequency operating point 404 (block 524) from the cluster power management circuit 124. The DVFS circuit 122 then sets the frequency of the core cluster 102(0) based on the target frequency operating point 404 (block 526). Depending on some aspects where the target frequency operating point 404 is higher than the current frequency of the core cluster 102(0), the operation of block 526 for setting the frequency of the core cluster 102(0) may include immediately increasing the frequency of the core cluster 102(0) (block 528). Depending on some aspects where the target frequency operating point 404 is lower than the current frequency of the core cluster 102(0), the operation of block 526 for setting the frequency of the core cluster 102(0) may include incrementally decreasing the frequency of the core cluster 102(0) (block 530).

[0058] Based on the information disclosed in this article and referenced Figure 1The processor devices discussed in these aspects can be provided in or integrated into any processor-based device. Examples, without limitation, include: set-top boxes, entertainment units, navigation devices, communication devices, fixed location data units, mobile location data units, Global Positioning System (GPS) devices, mobile phones, cellular phones, smartphones, Session Initiation Protocol (SIP) phones, tablets, phablets, servers, computers, portable computers, mobile computing devices, laptops, wearable computing devices (e.g., smartwatches, health or fitness trackers, glasses, etc.), desktop computers, personal digital assistants (PDAs), monitors, computer monitors, televisions, tuners, radios, satellite radios, music players, digital music players, portable music players, digital video players, video players, digital video disc (DVD) players, portable digital video players, automobiles, vehicle components, avionics systems, drones, and multi-rotor aircraft.

[0059] In this respect, Figure 6 An example of a processor-based device 600 is illustrated, which is as follows: Figure 1 As illustrated and described. In this example, processor-based device 600 includes processor device 602, which functionally corresponds to Figure 1 The processor device 100 includes one or more processor cores 604 coupled to a cache memory 606. The processor core 604 is also coupled to a system bus 608 and can interactively couple to devices included in the processor-based device 600. As is well known, the processor core 604 communicates with these other devices by exchanging address, control, and data information on the system bus 608. For example, the processor core 604 can communicate bus transaction requests to the memory controller 610. Although in Figure 6 Not illustrated, but multiple system buses 608 may be provided, each of which constitutes a different architecture.

[0060] Other devices can be connected to system bus 608. For example... Figure 6As illustrated, these devices may include a memory system 612, one or more input devices 614, one or more output devices 616, one or more network interface devices 618, and one or more display controllers 620. Input devices 614 may include any type of input device, including but not limited to input keys, switches, voice processors, etc. Output devices 616 may include any type of output device, including but not limited to audio, video, other visual indicators, etc. Network interface devices 618 may be any device configured to allow data exchange to and from network 622. Network 622 may be any type of network, including but not limited to wired or wireless networks, private or public networks, local area networks (LANs), wireless local area networks (WLANs), wide area networks (WANs), and Bluetooth. ™ Networks and the Internet. Network interface device 618 can be configured to support any type of communication protocol desired. Memory system 612 may include a memory controller 610 coupled to one or more memory arrays 624. Display controller may include, for example... Figure 1 GPU 106.

[0061] The processor core 604 can also be configured to access the display controller 620 via the system bus 608 to control the transmission of information to one or more displays 630. The display controller 620 transmits information to be displayed to the displays 630 via one or more video processors 632, which process the information to be displayed into a format suitable for the displays 630. The displays 630 may include any type of display, including but not limited to cathode ray tube (CRT), liquid crystal display (LCD), plasma display, light-emitting diode (LED) display, etc.

[0062] Those skilled in the art will further understand that the various exemplary logic blocks, modules, circuits, and algorithms described in connection with the aspects disclosed herein can be implemented as electronic hardware, stored in memory or another computer-readable medium and executed by a processor or other processing device, or a combination of both. As an example, the master and slave devices described herein can be employed in any circuit, hardware component, integrated circuit (IC), or IC chip. The memory disclosed herein can be of any type and size and can be configured to store any type of information desired. To clearly illustrate this interchangeability, the functionality of the various exemplary components, blocks, modules, circuits, and steps has been generally described above. How such functionality is implemented depends on the specific application, design choices, and / or design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure.

[0063] The various exemplary logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein may be implemented or executed using a processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof, designed to perform the functions described herein. The processor may be a microprocessor, but in alternative embodiments, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration).

[0064] The aspects disclosed herein may be embodied in hardware and instructions stored in the hardware, and may reside in, for example, random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disks, removable disks, CD-ROMs, or any other form of computer-readable medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a remote station. Alternatively, the processor and storage medium may reside as discrete components in a remote station, base station, or server.

[0065] It should also be noted that the operational steps described in any of the exemplary aspects of this document are described for the purpose of providing examples and discussion. The described operations may be performed in many different orders other than the order illustrated. Furthermore, the operations described in a single operational step may actually be performed in multiple different steps. Additionally, one or more operational steps discussed in the exemplary aspects may be combined. It should be understood that, as will be apparent to those skilled in the art, many different modifications may be made to the operational steps illustrated in the flowcharts. Those skilled in the art will also understand that any of a variety of different techniques and arts can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0066] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein can be applied to other variations. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0067] Specific implementation examples are described in the following numbered clauses:

[0068] 1. A processor device, the processor device comprising:

[0069] The core cluster includes:

[0070] Multiple processor cores, wherein the multiple processor cores include multiple corresponding activity management units (AMUs);

[0071] Cluster power management circuitry; and

[0072] Dynamic voltage and frequency scaling (DVFS) circuitry; and

[0073] The cluster power management circuit is configured as follows:

[0074] Collect multiple AMU statistics from the plurality of AMUs for each of one or more frequency operating points within a time interval;

[0075] A performance model that expresses processor performance as a function of frequency is generated based on the aforementioned AMU statistics.

[0076] Based on the performance model and power consumption measurements, a per-instruction energy consumption (EI) model is generated that expresses per-instruction energy consumption as a function of frequency.

[0077] An advantage model is generated based on the first rate of change of the performance model as a function of frequency and the second rate of change of the EI model as a function of frequency.

[0078] The target frequency operating point is identified based on the aforementioned advantage model; and

[0079] The target frequency operating point is sent to the DVFS circuit.

[0080] 2. The processor device according to Clause 1, wherein the DVFS circuitry is configured as follows:

[0081] Receive the target frequency operating point from the cluster power management circuit; and

[0082] The frequency of the core cluster is set based on the target frequency operating point.

[0083] 3. The processor device according to Clause 2, wherein:

[0084] The target frequency operating point is higher than the current frequency of the core cluster; and

[0085] The DVFS circuit is configured to set the frequency of the core cluster based on the target frequency operating point by being configured to immediately increase the frequency of the core cluster.

[0086] 4. The processor device according to Clause 2, wherein:

[0087] The target frequency operating point is lower than the current frequency of the core cluster; and

[0088] The DVFS circuit is configured to set the frequency of the core cluster based on the target frequency operating point by being configured to incrementally decrease the frequency of the core cluster.

[0089] 5. The processor device according to any one of clauses 1 to 4, wherein the time interval includes a programmable time interval.

[0090] 6. The processor device according to any one of Clauses 1 to 5, wherein the plurality of AMU statistics include one or more of the following: a count of processor frequency cycles, a count of constant frequency cycles, a count of decommissioning instructions, a count of front-end pauses, a count of memory pause cycles, a count of total demand misses in the last level cache (LLC), a count of memory-constrained front-end pauses, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculative execution instructions.

[0091] 7. The processor device according to any one of clauses 1 to 6, wherein the cluster power management circuitry is configured to generate the performance model by being configured to calculate a performance value for each of the one or more frequency operating points, the performance value comprising the product of the frequency operating point multiplied by the average number of instructions per clock cycle at the frequency operating point.

[0092] 8. The processor device according to any one of clauses 1 to 7, wherein the cluster power management circuitry is configured to generate the EI model by being configured to calculate an EI value for each of the one or more frequency operating points, the EI value comprising the quotient of power consumption at the frequency operating point divided by the performance value of the frequency operating point.

[0093] 9. The processor device according to any one of clauses 1 to 8, wherein the cluster power management circuitry is configured to generate the advantage model by dividing the first rate of change of the performance model as a function of frequency by the second rate of change of the EI model as a function of frequency.

[0094] 10. The processor device according to any one of clauses 1 to 9, wherein the cluster power management circuitry is configured to identify the target frequency operating point based on the dominance model by being configured to determine the maximum frequency operating point corresponding to a dominance value indicated by the dominance model as greater than or equal to a threshold.

[0095] 11. The processor device according to Clause 10, wherein the threshold corresponds to a Quality of Service (QoS) prompt provided by an operating system (OS) executed by the processor device.

[0096] 12. The processor device according to any one of clauses 10 to 11, wherein the cluster power management circuitry is configured to determine the maximum frequency operating point by being configured to determine the maximum frequency operating point as an interpolation between the energy balance frequency operating point and the performance frequency operating point based on an energy performance preference (EPP) cues.

[0097] 13. The processor device according to any one of Clauses 1 to 12, wherein the processor device is integrated into a device selected from the group consisting of: set-top boxes; entertainment units; navigation devices; communication devices; fixed location data units; mobile location data units; global positioning system (GPS) devices; mobile phones; cellular phones; smartphones; session initiation protocol (SIP) phones; tablet computers; tablet phones; servers; computers; portable computers; mobile computing devices; wearable computing devices; desktop computers; personal digital assistants (PDAs); monitors; computer monitors; televisions; tuners; radios; satellite radios; music players; digital music players; portable music players; digital video players; video players; digital video disc (DVD) players; portable digital video players; automobiles; vehicle components; avionics systems; unmanned aerial vehicles; and multirotor aircraft.

[0098] 14. A processor device, the processor device comprising:

[0099] A component for collecting multiple AMU statistics from multiple Activity Management Units (AMUs) of multiple processor cores corresponding to the core cluster of the processor device for each of one or more frequency operating points within a time interval.

[0100] A component for generating a performance model that represents processor performance as a function of frequency based on the plurality of AMU statistics;

[0101] A component for generating a per-instruction energy consumption (EI) model that expresses per-instruction energy consumption as a function of frequency based on the performance model and power consumption measurements;

[0102] A component for generating an advantage model based on a first rate of change as a function of frequency of the performance model and a second rate of change as a function of frequency of the EI model;

[0103] Components for identifying the target frequency operating point based on the aforementioned advantage model; and

[0104] Components for sending the target frequency operating point to the dynamic voltage and frequency scaling (DVFS) circuitry of the processor device.

[0105] 15. A method for autonomously managing the core cluster frequency in a processor-based device, the method comprising:

[0106] The cluster power management circuitry of the processor device collects multiple AMU statistics from multiple activity management units (AMUs) of multiple processor cores corresponding to the core cluster of the processor device for each of one or more frequency operating points within a time interval.

[0107] The cluster power management circuit generates a performance model that represents processor performance as a function of frequency based on the statistical data of the multiple AMUs;

[0108] The cluster power management circuit generates a per-instruction energy consumption (EI) model that expresses per-instruction energy consumption as a function of frequency based on the performance model and power consumption measurements.

[0109] The cluster power management circuit generates an advantage model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency.

[0110] The cluster power management circuit identifies the target frequency operating point based on the advantage model; and

[0111] The cluster power management circuit sends the target frequency operating point to the dynamic voltage and frequency scaling (DVFS) circuit of the processor device.

[0112] 16. The method according to Clause 15, further comprising:

[0113] The target frequency operating point is received by the DVFS circuit from the cluster power management circuit; and

[0114] The frequency of the core cluster is set by the DVFS circuit based on the target frequency operating point.

[0115] 17. The method according to Clause 16, wherein:

[0116] The target frequency operating point is higher than the current frequency of the core cluster; and

[0117] Setting the frequency of the core cluster based on the target frequency operating point includes immediately increasing the frequency of the core cluster.

[0118] 18. The method according to Clause 16, wherein:

[0119] The target frequency operating point is lower than the current frequency of the core cluster; and

[0120] Setting the frequency of the core cluster based on the target frequency operating point includes incrementally decreasing the frequency of the core cluster.

[0121] 19. The method according to any one of Clauses 15 to 18, wherein the time interval includes a programmable time interval.

[0122] 20. The method according to any one of Clauses 15 to 19, wherein the plurality of AMU statistics include one or more of the following: a count of processor frequency cycles, a count of constant frequency cycles, a count of retirement instructions, a count of front-end pauses, a count of memory pause cycles, a count of total demand misses in the last level cache (LLC), a count of memory-constrained front-end pauses, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculative execution instructions.

[0123] 21. The method according to any one of Clauses 15 to 20, wherein generating the performance model comprises calculating a performance value for each of the one or more frequency operating points, the performance value comprising the product of the frequency operating point multiplied by the average number of instructions per clock cycle at the frequency operating point.

[0124] 22. The method according to any one of Clauses 15 to 21, wherein generating the EI model comprises calculating an EI value for each of the one or more frequency operating points, the EI value comprising the quotient of power consumption at the frequency operating point divided by the performance value of the frequency operating point.

[0125] 23. The method according to any one of clauses 15 to 22, wherein generating the advantage model comprises calculating the first rate of change of the performance model as a function of frequency by dividing the second rate of change of the EI model as a function of frequency.

[0126] 24. The method according to any one of clauses 15 to 23, wherein identifying the target frequency operating point based on the dominance model includes determining the maximum frequency operating point corresponding to a dominance value indicated by the dominance model as greater than or equal to a threshold.

[0127] 25. The method according to Clause 24, wherein the threshold corresponds to a Quality of Service (QoS) prompt provided by an operating system (OS) executed by the processor device.

[0128] 26. The method according to any one of Clauses 24 to 25, wherein determining the maximum frequency operating point includes determining the maximum frequency operating point as an interpolation between the energy balance frequency operating point and the performance frequency operating point based on an energy performance preference (EPP) cue.

[0129] 27. A non-transitory computer-readable medium having stored thereon computer-executable instructions, which, when executed, cause a processor device of a processor-based device to:

[0130] Within a time interval, multiple AMU statistics are collected from multiple Activity Management Units (AMUs) of multiple processor cores corresponding to the core cluster of the processor device for each of one or more frequency operating points;

[0131] A performance model that expresses processor performance as a function of frequency is generated based on the aforementioned AMU statistics.

[0132] Based on the performance model and power consumption measurements, a per-instruction energy consumption (EI) model is generated that expresses per-instruction energy consumption as a function of frequency.

[0133] An advantage model is generated based on the first rate of change of the performance model as a function of frequency and the second rate of change of the EI model as a function of frequency.

[0134] The target frequency operating point is identified based on the aforementioned advantage model; and

[0135] The target frequency operating point is sent to the Dynamic Voltage and Frequency Scaling (DVFS) circuitry of the processor device.

[0136] 28. The non-transitory computer-readable medium according to Clause 27, wherein the computer-executable instructions further cause the processor device to:

[0137] Receive the target frequency operating point; and

[0138] The frequency of the core cluster is set based on the target frequency operating point.

[0139] 29. The non-transitory computer-readable medium as described in Clause 28, wherein:

[0140] The target frequency operating point is higher than the current frequency of the core cluster; and

[0141] The computer-executable instructions cause the processor device to set the frequency of the core cluster based on the target frequency operating point by causing the processor device to immediately increase the frequency of the core cluster.

[0142] 30. The non-transitory computer-readable medium as described in Clause 28, wherein:

[0143] The target frequency operating point is lower than the current frequency of the core cluster; and

[0144] The computer-executable instructions cause the processor device to set the frequency of the core cluster based on the target frequency operating point by causing the processor device to incrementally decrease the frequency of the core cluster.

[0145] 31. The non-transitory computer-readable medium according to any one of Clauses 27 to 30, wherein the time interval includes a programmable time interval.

[0146] 32. A nontransitory computer-readable medium according to any one of Clauses 27 to 31, wherein the plurality of AMU statistics include one or more of the following: a count of processor frequency cycles, a count of constant frequency cycles, a count of retirement instructions, a count of front-end pauses, a count of memory pause cycles, a count of total demand misses in the last level cache (LLC), a count of memory-constrained front-end pauses, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculative execution instructions.

[0147] 33. The non-transitory computer-readable medium according to any one of clauses 27 to 32, wherein the computer-executable instructions cause the processor device to generate the performance model by causing the processor device to calculate a performance value for each of the one or more frequency operating points, the performance value comprising the product of the frequency operating point multiplied by the average number of instructions per clock cycle at the frequency operating point.

[0148] 34. The non-transitory computer-readable medium according to any one of clauses 27 to 33, wherein the computer-executable instructions cause the processor device to generate the EI model by causing the processor device to calculate an EI value for each of the one or more frequency operating points, the EI value comprising the quotient of power consumption at the frequency operating point divided by the performance value of the frequency operating point.

[0149] 35. A non-transitory computer-readable medium according to any one of clauses 27 to 34, wherein the computer-executable instructions cause the processor device to generate the advantage model by causing the processor device to calculate the first rate of change of the performance model as a function of frequency divided by the second rate of change of the EI model as a function of frequency.

[0150] 36. A non-transitory computer-readable medium according to any one of clauses 27 to 35, wherein the computer-executable instructions cause the processor device to identify the target frequency operating point based on the dominance model by causing the processor device to determine a maximum frequency operating point corresponding to a dominance value indicated by the dominance model as greater than or equal to a threshold.

[0151] 37. The non-transitory computer-readable medium as described in Clause 36, wherein the threshold corresponds to a Quality of Service (QoS) prompt provided by an operating system (OS) executed by the processor device.

[0152] 38. The nontransitory computer-readable medium according to any one of clauses 36 to 37, wherein the computer-executable instructions cause the processor device to determine the maximum frequency operating point by causing the processor device to determine the maximum frequency operating point as an interpolation between the energy balance frequency operating point and the performance frequency operating point based on the energy performance preference (EPP) cues.

Claims

1. A processor device comprising: a core cluster comprising: a plurality of processor cores comprising a corresponding plurality of activity management units (AMUs); a cluster power management circuit; and a dynamic voltage and frequency scaling (DVFS) circuit; and the cluster power management circuit is configured to: collect, over a time interval, a plurality of AMU statistics from the plurality of AMUs for each of one or more frequency operating points; generate, based on the plurality of AMU statistics, a performance model representing processor performance as a function of frequency; generate, based on the performance model and a power consumption measurement, an energy per instruction (EI) model representing energy per instruction as a function of frequency; generate, based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency, a dominance model; identify, based on the dominance model, a target frequency operating point; and send the target frequency operating point to the DVFS circuit.

2. The processor device of claim 1, wherein the DVFS circuit is configured to: receive the target frequency operating point from the cluster power management circuit; and set a frequency of the core cluster based on the target frequency operating point.

3. The processor device of claim 2, wherein: the target frequency operating point is higher than a current frequency of the core cluster; and the DVFS circuit is configured to set the frequency of the core cluster based on the target frequency operating point by being configured to immediately increase the frequency of the core cluster.

4. The processor device of claim 2, wherein: the target frequency operating point is lower than a current frequency of the core cluster; and the DVFS circuit is configured to set the frequency of the core cluster based on the target frequency operating point by being configured to incrementally decrease the frequency of the core cluster.

5. The processor device of claim 1, wherein the time interval comprises a programmable time interval.

6. The processor device of claim 1, wherein the plurality of AMU statistics comprises one or more of: a count of processor frequency cycles, a count of constant frequency cycles, a count of retired instructions, a count of front-end stalls, a count of memory stall cycles, a count of total demand misses of a last level cache (LLC), a count of front-end stalls constrained by memory, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculatively executed instructions.

7. The processor device of claim 1, wherein the cluster power management circuit is configured to generate the performance model by being configured to calculate, for each of the one or more frequency operating points, a performance value comprising a product of the frequency operating point multiplied by an average number of instructions per clock cycle at the frequency operating point.

8. The processor device of claim 1, wherein the cluster power management circuit is configured to generate the EI model by being configured to compute an EI value for each frequency operating point of the one or more frequency operating points, the EI value comprising a quotient of a power consumption at the frequency operating point divided by a performance value of the frequency operating point.

9. The processor device of claim 1, wherein the cluster power management circuit is configured to generate the advantage model by being configured to compute the first rate of change of the performance model as a function of frequency divided by the second rate of change of the EI model as a function of frequency.

10. The processor device of claim 1, wherein the cluster power management circuit is configured to identify the target frequency operating point based on the advantage model by being configured to determine a maximum frequency operating point corresponding to an advantage value indicated by the advantage model as being greater than or equal to a threshold value.

11. The processor device of claim 10, wherein the threshold value corresponds to a quality of service (QoS) hint provided by an operating system (OS) executed by the processor device.

12. The processor device of claim 10, wherein the cluster power management circuit is configured to determine the maximum frequency operating point by being configured to determine the maximum frequency operating point as an interpolated value between an energy balanced frequency operating point and a performance frequency operating point based on an energy performance preference (EPP) hint.

13. The processor device of claim 1 integrated into a device selected from a group consisting of: a set top box; an entertainment unit; a navigation device; a communications device; a fixed location data unit; a mobile location data unit; a global positioning system (GPS) device; a mobile phone; a cellular phone; a smartphone; a session initiation protocol (SIP) phone; a tablet; a phablet; a server; a computer; a portable computer; a mobile computing device; a wearable computer; a desktop computer; a personal digital assistant (PDA); a monitor; a computer monitor; a television; a tuner; a radio; a satellite radio; a music player; a digital music player; a portable music player; a digital video player; a video player; a digital video disc (DVD) player; a portable digital video player; an automobile; a transportation component; avionics; a drone; and a multicopter.

14. A processor device comprising: means for collecting, for each frequency operating point of one or more frequency operating points, a plurality of AMU statistics from a plurality of activity management units (AMUs) corresponding to a plurality of processor cores of a core cluster of the processor device for a time interval; means for generating, based on the plurality of AMU statistics, a performance model representing processor performance as a function of frequency; means for generating, based on the performance model and a power consumption measurement, an energy per instruction (EI) model representing energy per instruction energy consumption as a function of frequency; means for generating an advantage model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency; means for identifying a target frequency operating point based on the advantage model; and means for sending the target frequency operating point to a dynamic voltage and frequency scaling (DVFS) circuit of the processor device.

15. A method for autonomously managing a core cluster frequency in a processor-based device, the method comprising: collecting, by a cluster power management circuit of a processor device, a plurality of activity management unit (AMU) statistics from a plurality of AMUs corresponding to a plurality of processor cores of a core cluster of the processor device for each of one or more frequency operating points over a time interval; generating, by the cluster power management circuit, a performance model representing processor performance as a function of frequency based on the plurality of AMU statistics; generating, by the cluster power management circuit, an energy per instruction (EI) model representing energy per instruction consumption as a function of frequency based on the performance model and power consumption measurements; generating, by the cluster power management circuit, an advantage model based on a first rate of change of the performance model as a function of frequency and a second rate of change of the EI model as a function of frequency; identifying, by the cluster power management circuit, a target frequency operating point based on the advantage model; and sending, by the cluster power management circuit, the target frequency operating point to a dynamic voltage and frequency scaling (DVFS) circuit of the processor device.

16. The method of claim 15, further comprising: receiving, by the DVFS circuit, the target frequency operating point from the cluster power management circuit; and setting, by the DVFS circuit, a frequency of the core cluster based on the target frequency operating point.

17. The method of claim 16, wherein: the target frequency operating point is higher than a current frequency of the core cluster; and setting the frequency of the core cluster based on the target frequency operating point comprises immediately raising the frequency of the core cluster.

18. The method of claim 16, wherein: the target frequency operating point is lower than a current frequency of the core cluster; and setting the frequency of the core cluster based on the target frequency operating point comprises incrementally lowering the frequency of the core cluster.

19. The method of claim 15, wherein the time interval comprises a programmable time interval.

20. The method of claim 15, wherein the plurality of AMU statistics comprises one or more of a count of processor frequency cycles, a count of constant frequency cycles, a count of retired instructions, a count of front-end stalls, a count of memory stall cycles, a count of total demand misses of a last level cache (LLC), a count of front-end stalls constrained by memory, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculatively executed instructions.

21. The method of claim 15, wherein generating the performance model comprises calculating, for each frequency operating point of the one or more frequency operating points, a performance value comprising a product of the frequency operating point multiplied by an average number of instructions per clock cycle at the frequency operating point.

22. The method of claim 15, wherein generating the El model comprises calculating, for each frequency operating point of the one or more frequency operating points, an El value comprising a quotient of a power consumption at the frequency operating point divided by a performance value of the frequency operating point.

23. The method of claim 15, wherein generating the advantage model comprises calculating the first rate of change of the performance model as a function of frequency divided by the second rate of change of the El model as a function of frequency.

24. The method of claim 15, wherein identifying the target frequency operating point based on the advantage model comprises determining a maximum frequency operating point corresponding to an advantage value indicated by the advantage model as being greater than or equal to a threshold value.

25. The method of claim 24, wherein the threshold value corresponds to a quality of service (QoS) hint provided by an operating system (OS) executed by the processor device.

26. The method of claim 24, wherein determining the maximum frequency operating point comprises determining the maximum frequency operating point as an interpolated value between an energy balanced frequency operating point and a performance frequency operating point based on an energy performance preference (EPP) hint.

27. A non-transitory computer-readable medium having stored thereon computer- executable instructions, which, when executed, cause a processor device of a processor-based device to: collect, for each frequency operating point of one or more frequency operating points, a plurality of AMU statistics from a plurality of activity management units (AMUs) of a plurality of processor cores corresponding to a core cluster of the processor device over a time interval; generate, based on the plurality of AMU statistics, a performance model representing processor performance as a function of frequency; generate, based on the performance model and power consumption measurements, an energy per instruction (El) model representing energy per instruction as a function of frequency; generate, based on a first rate of change of the performance model as a function of frequency and a second rate of change of the El model as a function of frequency, an advantage model; identify, based on the advantage model, a target frequency operating point; and send the target frequency operating point to a dynamic voltage and frequency scaling (DVFS) circuit of the processor device.

28. The non-transitory computer-readable medium of claim 27, wherein the computer- executable instructions further cause the processor device to: receive the target frequency operating point; and set a frequency of the core cluster based on the target frequency operating point.

29. The non-transitory computer-readable medium of claim 28, wherein: the target frequency operating point is higher than a current frequency of the core cluster; and the target frequency operating point is lower than the current frequency of the core cluster. The computer-executable instructions cause the processor device to set the frequency of the core cluster based on the target frequency operating point by causing the processor device to incrementally increase the frequency of the core cluster.

30. The non-transitory computer-readable medium of claim 28, wherein: the target frequency operating point is lower than a current frequency of the core cluster; and the computer-executable instructions cause the processor device to set the frequency of the core cluster based on the target frequency operating point by causing the processor device to incrementally decrease the frequency of the core cluster.

31. The non-transitory computer-readable medium of claim 27, wherein the time interval comprises a programmable time interval.

32. The non-transitory computer-readable medium of claim 27, wherein the plurality of AMU statistics comprises one or more of a count of processor frequency cycles, a count of constant frequency cycles, a count of retire instructions, a count of front end stalls, a count of memory stall cycles, a count of total demand misses of a last level cache (LLC), a count of front end stalls that are memory bound, a count of LLC demand accesses, a count of bus accesses, a count of bus access cycles, and a count of speculatively executed instructions.

33. The non-transitory computer-readable medium of claim 27, wherein the computer- executable instructions cause the processor device to generate the performance model by causing the processor device to calculate, for each of the one or more frequency operating points, a performance value comprising a product of the frequency operating point multiplied by an average number of instructions per clock cycle at the frequency operating point.

34. The non-transitory computer-readable medium of claim 27, wherein the computer- executable instructions cause the processor device to generate the EI model by causing the processor device to calculate, for each of the one or more frequency operating points, an EI value comprising a quotient of a power consumption at the frequency operating point divided by a performance value of the frequency operating point.

35. The non-transitory computer-readable medium of claim 27, wherein the computer- executable instructions cause the processor device to generate the advantage model by causing the processor device to calculate, as a function of frequency, a first rate of change of the performance model divided by a second rate of change of the EI model.

36. The non-transitory computer-readable medium of claim 27, wherein the computer- executable instructions cause the processor device to identify the target frequency operating point based on the advantage model by causing the processor device to determine a maximum frequency operating point corresponding to an advantage value indicated by the advantage model as being greater than or equal to a threshold value.

37. The non-transitory computer-readable medium of claim 36, wherein the threshold value corresponds to a quality of service (QoS) hint provided by an operating system (OS) executed by the processor device.

38. The non-transitory computer readable medium of claim 36, wherein the computer executable instructions cause the processor device to determine the maximum frequency operating point by causing the processor device to determine the maximum frequency operating point as an interpolated value between an energy balanced frequency operating point and a performance frequency operating point based on an energy performance preference (EPP) hint.