Performing dynamic microarchitecture throttling of processor cores based on quality of service (QoS) levels in processor devices
By introducing throttling selection circuits and microarchitecture throttling technology into processor devices, the throttling level of the processor core is dynamically adjusted based on QoS level and performance status, solving the problem of increased power consumption caused by frequency changes in synchronous core clusters, and achieving the effect of reducing power consumption without degrading performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2026-04-10
AI Technical Summary
In conventional processor devices, frequency changes in the synchronous core cluster affect all processor cores, leading to increased power consumption. This is especially true under workload scheduling with different Quality of Service (QoS) levels, where workloads with lower QoS levels still need to be executed at the highest QoS level frequency, resulting in unnecessary power consumption.
By introducing throttling selection circuitry into the processor device, the throttling level of the processor core is determined based on QoS level and performance status. Microarchitectural throttling techniques, such as inserting no-operation instructions, are used to adjust instruction execution efficiency without changing frequency or voltage, thereby achieving dynamic power management.
While ensuring performance requirements, reduce the power consumption of the processor core, improve the overall energy efficiency of the processor device, and avoid unnecessary power consumption.
Smart Images

Figure CN121844277A_ABST
Abstract
Description
Priority Application
[0001] This application claims priority to U.S. Patent Application Serial No. 18 / 469,630, entitled “PERFORMING DYNAMIC MICROARCHITECTURAL THROTTLING OF PROCESSOR CORES BASED ON QUALITY-OF-SERVICE (QoS) LEVELS IN PROCESSOR DEVICES,” filed September 19, 2023, which is incorporated by reference herein in its entirety. TECHNICAL FIELD
[0002] The technology of the present disclosure generally relates to power and performance management in multi-core processor-based devices, and specifically relates to frequency management for clusters of processor cores of processor devices. BACKGROUND
[0003] Conventional processor devices can implement the Advanced Configuration and Power Interface (ACPI) specification, which defines an open industry standard that includes power management across hardware, operating systems (OS), and application software of a processor device. Using functionality defined by the ACPI specification, a processor device can perform frequency management to modify its performance and power consumption. For example, when a workload executed by a processor device does not require enhanced performance and / or does not involve a user experience that requires higher performance, the frequency of the processor device can be reduced. Reducing the frequency of the processor device can reduce power consumption. Conversely, if a workload executed by a processor device requires enhanced performance and / or involves a user experience that requires higher performance, the frequency of the processor device can be increased. However, increasing the frequency of the processor device also increases the power consumption of the processor device.
[0004] Some conventional processor devices are implemented as multiple processor cores organized into core clusters. Each core cluster can be "synchronous" in that all processor cores of the core cluster are clocked using a single clock source, such as a phase-locked loop (PLL). Because the processor cores all share the same clock source, a frequency change of the core cluster affects all processor cores within the core cluster. However, when an operating system (OS) scheduler executing on the core cluster schedules workloads on processor cores associated with different quality of service (QoS) levels, the power consumption of the core cluster can be negatively impacted. Because each QoS level corresponds to a different frequency and power expectation, the frequency of the core cluster can be determined by the highest QoS level of all workloads executing on the multiple processor cores. Thus, workloads requiring a lower QoS level must still execute at the frequency required by the highest QoS level of all workloads executing on the processor cores, resulting in increased power consumption of the core cluster. SUMMARY
[0005] Aspects disclosed in the detailed description include performing dynamic micro- architecture throttling of a processor core based on a quality of service (QoS) level in a processor device. Related apparatuses, methods, and computer-readable media are also disclosed. In this regard, a processor device includes a synchronous core cluster including a plurality of processor cores, a throttle selection circuit, and a throttle circuit. The throttle selection circuit of the synchronous core cluster is configured to determine a performance state of a processor core of the plurality of processor cores and receive, from the processor core, a QoS level associated with a workload scheduled for execution by the processor core. Subsequently (e.g., at periodic intervals), the throttle selection circuit determines a throttle level of the processor core based on the QoS level and the performance state and provides the throttle level to the throttle circuit. Upon receiving the throttle level, the throttle circuit performs micro-architecture throttling of the processor core based on the throttle level. As used herein, "micro-architecture throttling" refers to modifying the efficiency of instruction execution of a processing core (e.g., by inserting no-operation (NOP) instructions for execution by the processor core, as non-limiting examples) without changing the frequency or voltage of the synchronous core cluster. In this manner, lower performance threads executing on the processor core consume less power without compromising performance requirements and further make power available to the processor device as a whole.
[0006] Some aspects may specify that the throttling selection circuitry determines the Energy Performance Preference (EPP) level corresponding to the QoS level, and determines the throttling level based on the QoS level and performance level by determining the throttling level based on the EPP level and performance level. In some aspects, during each periodic interval, the throttling selection circuitry populates each of a plurality of throttling level lookup tables (LUTs) corresponding to a plurality of processor cores. For example, the throttling selection circuitry may calculate the average core frequency corresponding to each of the plurality of EPP levels. Then, for each of the plurality of throttling levels, the throttling selection circuitry calculates the corresponding performance state of the processor core that requires the throttling level to achieve at least the average core frequency. According to some aspects, the determination of the EPP level corresponding to the QoS level is based on a mapping register that maps the QoS level to the EPP level. Some aspects may specify that determining the EPP level corresponding to the QoS level includes mapping the QoS level to the EPP level based on the performance state of the processor core.
[0007] In some aspects, determining the throttling level based on the EPP level and the performance state may include selecting the row corresponding to the EPP level from among a plurality of throttling level LUTs corresponding to the processor core. In such aspects, the throttling selection circuitry then determines the throttling level based on the column of the lowest performance state in that row that is greater than or equal to the performance state of the processor core.
[0008] Some aspects may specify that the Dynamic Voltage and Frequency Scaling (DVFS) aggregator circuitry of the synchronous core cluster receives multiple corresponding EPP cues from the multiple processor cores. The DVFS aggregator circuitry selects a cluster performance state for the synchronous core cluster based on these multiple EPP cues (e.g., by selecting the highest performance state indicated by multiple mapped LUTs corresponding to the multiple processor cores). The DVFS aggregator circuitry sends the cluster performance state to the DVFS circuitry of the synchronous core cluster, which then sets the frequency and voltage of the synchronous core cluster based on the cluster performance state.
[0009] In another aspect, a processor device is provided. The processor device includes a synchronous core cluster comprising a plurality of processor cores, a throttling selection circuit, and a throttling circuit. The throttling selection circuit is configured to determine a performance state of one of the plurality of processor cores. The throttling selection circuit is further configured to receive from the processor core a QoS level associated with a workload scheduled for execution by the processor core. The throttling selection circuit is also configured to determine a throttling level for the processor core based on the QoS level and the performance state. The throttling selection circuit is additionally configured to provide the throttling level to the throttling circuit. The throttling circuit is configured to receive the throttling level and perform microarchitectural throttling of the processor core based on the throttling level.
[0010] In another aspect, a processor device is provided. The processor device includes components for determining the performance state of a processor core among a plurality of processor cores in a synchronized core cluster of the processor device. The processor device also includes components for receiving from the processor core a QoS level associated with a workload scheduled for execution by the processor core. The processor device further includes components for determining a throttling level for the processor core based on the QoS level and the performance state. The processor device additionally includes components for performing microarchitectural throttling of the processor core based on the throttling level.
[0011] On the other hand, a method is provided for performing dynamic microarchitectural throttling of a processor core based on a QoS level. The method includes determining the performance state of a processor core among a plurality of processor cores in a synchronous core cluster of a processor device by a throttling selection circuit of the synchronous core cluster. The method also includes receiving, by the throttling selection circuit, a QoS level associated with a workload scheduled for execution by the processor core. The method further includes determining a throttling level for the processor core by the throttling selection circuit based on the QoS level and the performance state. The method additionally includes providing the throttling level to a throttling circuit of the synchronous core cluster by the throttling selection circuit. The method also includes receiving the throttling level by the throttling circuit. Finally, the method includes performing microarchitectural throttling of the processor core by the throttling circuit based on the throttling level.
[0012] In another aspect, a non-transitory computer-readable medium is disclosed. This non-transitory computer-readable medium stores computer-executable instructions that, when executed, cause a processor device of a processor-based device to determine the performance state of a processor core among a plurality of processor cores in a synchronous core cluster of the processor device. The computer-executable instructions also cause the processor device to receive a QoS level associated with a workload scheduled for execution by the processor core. The computer-executable instructions further cause the processor device to determine a throttling level for the processor core based on the QoS level and the performance state. The computer-executable instructions additionally cause the processor device to perform microarchitectural throttling of the processor core based on the throttling level. Attached Figure Description
[0013] Figure 1 This is a block diagram of an exemplary processor-based device, which includes a synchronous core cluster configured to perform dynamic microarchitectural throttling of processor cores based on quality of service (QoS) levels.
[0014] Figure 2 This provides a more detailed example based on several aspects. Figure 1 A diagram of one of the synchronization core clusters, which includes throttling selection circuitry and dynamic voltage and frequency scaling (DVFS) circuitry.
[0015] Figure 3 This is an example based on some aspects. Figure 2 The throttling selection circuit uses an exemplary throttling level lookup table (LUT) diagram to map energy performance preference (EPP) levels and performance status to throttling levels.
[0016] Figure 4 This is an example based on some aspects Figure 2 throttling selection circuit and Figure 2 and Figure 3 A diagram illustrating an exemplary physical implementation of a throttling level LUT.
[0017] Figures 5A to 5D Examples are provided based on some aspects. Figure 1 A flowchart illustrating an exemplary operation performed by a processor device to perform dynamic microarchitectural throttling of the processor core based on QoS levels.
[0018] Figure 6 Yes, it can include Figure 1 A block diagram of an exemplary processor-based device. Detailed Implementation
[0019] Several exemplary aspects of this disclosure will now be described with reference to the accompanying drawings. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or superior to other aspects. The terms “first,” “second,” etc., are used herein to distinguish similarly named elements and are not to be construed as indicating an ordering relationship between such elements unless expressly stated herein.
[0020] The aspects disclosed in the detailed description include performing dynamic microarchitectural throttling of processor cores based on the Quality of Service (QoS) level in the processor device. Related apparatus, methods, and computer-readable media are also disclosed. In this regard, the processor device includes a synchronous core cluster comprising a plurality of processor cores, a throttling selection circuit, and a throttling circuit. The throttling selection circuit of the synchronous core cluster is configured to determine the performance state of one of the plurality of processor cores and receive from the processor core a QoS level associated with a workload scheduled for execution by the processor core. Subsequently (e.g., at periodic intervals), the throttling selection circuit determines a throttling level for the processor core based on the QoS level and the performance state, and provides the throttling level to the throttling circuit. Upon receiving the throttling level, the throttling circuit performs microarchitectural throttling of the processor core based on the throttling level. As used herein, "microarchitectural throttling" refers to modifying the efficiency of instruction execution of a processing core (e.g., by inserting no-operation (NOP) instructions for execution by the processor core, as a non-limiting example) without changing the frequency or voltage of the synchronous core cluster. In this way, lower-performance threads executing on the processor core consume less power without compromising performance requirements, and further make power available for the processor device as a whole.
[0021] Some aspects may specify that the throttling selection circuitry determines the Energy Performance Preference (EPP) level corresponding to the QoS level, and determines the throttling level based on the QoS level and performance level by determining the throttling level based on the EPP level and performance level. In some aspects, during each periodic interval, the throttling selection circuitry populates each of a plurality of throttling level lookup tables (LUTs) corresponding to a plurality of processor cores. For example, the throttling selection circuitry may calculate the average core frequency corresponding to each of the plurality of EPP levels. Then, for each of the plurality of throttling levels, the throttling selection circuitry calculates the corresponding performance state of the processor core that requires the throttling level to achieve at least the average core frequency. According to some aspects, the determination of the EPP level corresponding to the QoS level is based on a mapping register that maps the QoS level to the EPP level. Some aspects may specify that determining the EPP level corresponding to the QoS level includes mapping the QoS level to the EPP level based on the performance state of the processor core.
[0022] In some aspects, determining the throttling level based on the EPP level and the performance state may include selecting the row corresponding to the EPP level from among a plurality of throttling level LUTs corresponding to the processor core. In such aspects, the throttling selection circuitry then determines the throttling level based on the column of the lowest performance state in that row that is greater than or equal to the performance state of the processor core.
[0023] Some aspects may specify that the Dynamic Voltage and Frequency Scaling (DVFS) aggregator circuitry of the synchronous core cluster receives multiple corresponding EPP cues from the multiple processor cores. The DVFS aggregator circuitry selects a cluster performance state for the synchronous core cluster based on these multiple EPP cues (e.g., by selecting the highest performance state indicated by multiple mapped LUTs corresponding to the multiple processor cores). The DVFS aggregator circuitry sends the cluster performance state to the DVFS circuitry of the synchronous core cluster, which then sets the frequency and voltage of the synchronous core cluster based on the cluster performance state.
[0024] In this respect, Figure 1 This is a block diagram of an exemplary processor device 100 (also referred to as a “processor” or “CPU”). The processor device 100 may include ordered or unordered processors (OoP), and / or may be one of a plurality of processor devices 100. Examples of processor devices 100 may include, but are not limited to, digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (PGAs), or other equivalent integrated or discrete logic circuits.
[0025] like Figure 1 As seen, processor device 100 includes multiple synchronous core clusters (in Figure 1 The clusters are labeled “synchronous core clusters” (102(0)-102(X)), each of which includes multiple processor cores (not shown). Figure 1 The processor device 100 in the example also includes a graphics processing unit (GPU) 104 for performing graphics operations. As a non-limiting example, the GPU 104 may include a dedicated hardware unit with fixed functionality and programmable components for rendering graphics and executing GPU applications. The GPU 104 may also include a DSP, a general-purpose microprocessor, an ASIC, an FPGA, or other equivalent integrated or discrete logic circuitry, for clarity. Figure 1 Not shown in the image.
[0026] Figure 1The processor device 100 also includes additional exemplary components, including an artificial intelligence (AI) engine 106, mobile device management (MDM) circuitry 108, power management circuitry 110, on-chip network (NoC) 112, and memory device 114. As a non-limiting example, the AI engine 106 of the processor device 100 includes circuitry and logic for providing AI-based functionality such as search, speech recognition, text and / or image generation. The MDM circuitry 108 provides functionality for provisioning, configuring, updating, and / or protecting mobile devices to which the processor device 100 is integrated. The power management circuitry 110 provides advanced performance and power management functionality for the processor device 100 as a whole, while the NoC 112 is configured to manage communication between different devices including the processor device 100. Finally, the memory device 114 provides storage and access to data used by the processor device 100, and in some aspects, as a non-limiting example, may include a double data rate (DDR) synchronous dynamic random access memory (SDRAM) device.
[0027] Figure 1 The processor device 100 may encompass any of known digital logic elements, semiconductor circuits, processing cores, and / or memory structures, as well as other elements or combinations thereof. The aspects described herein are not limited to any particular arrangement of elements, and the disclosed techniques can be readily extended to various structures and layouts on semiconductor dies or packages. It should be understood that some aspects of the processor device 100 may include, in addition to… Figure 1 Elements other than those exemplified, and / or may include Figure 1 The illustrated components may be more or fewer. For example, processor device 100 may also include cache, controller, communication bus and / or persistent storage devices, which, for clarity, are described in... Figure 2 The middle part is omitted.
[0028] Figure 1 More detailed examples Figure 2 An exemplary component of the synchronous core cluster 102(0). Figure 2 As seen, the synchronous core cluster 102(0) includes multiple processor cores 200(0)-200(C). The processor cores 200(0)-200(C) of the synchronous core cluster 102(0) are communicatively coupled to a last-level cache (LLC) that stores frequently accessed data. Figure 2 The LLC202 is labeled as such for faster access and is communicatively coupled to a phase-locked loop (PLL) that provides clock signals to processor cores 200(0)-200(C) and LLC202. Figure 1The term “PLL” is used herein to refer to the synchronous core cluster 102(0). As used herein, the term “synchronous” as used with respect to the synchronous core cluster 102(0) refers to the fact that because processor cores 200(0)-200(C) and LLC 202 all receive the same clock signal provided by PLL 204, processor cores 200(0)-200(C) and LLC 202 all operate at the same frequency. The synchronous core cluster 102(0) can be placed in one of several performance states, each corresponding to a combination of frequencies and voltages at which all processor cores 200(0)-200(C) operate. The frequencies and voltages at which processor cores 200(0)-200(C) of the synchronous core cluster 102(0) operate are controlled by DVFS circuitry 206, and performance and power management for the synchronous core cluster 102(0) is handled by cluster power management circuitry 208. It will be understood that, although Figure 2 Only exemplary elements of the synchronization core cluster 102(0) are shown, but each of the synchronization core clusters 102(0)-102(C) includes elements corresponding to the illustrated elements of the synchronization core cluster 102(0). It will be further understood that the synchronization core cluster 102(0) may include elements not shown for clarity. Figure 1 Additional components are shown in the example.
[0029] Because processor cores 200(0)-200(C) all operate at the same frequency, a frequency change in the synchronous core cluster 102(0) (e.g., caused by a change in performance state) affects all processor cores 200(0)-200(C) within the synchronous core cluster 102(0). Figure 2 When the operating system (OS) scheduler running on the processor device 100 schedules workloads associated with different QoS levels on processor cores 200(0)-200(C), the frequency of the synchronous core cluster 102(0) can be determined by the highest QoS level of all workloads executed on processor cores 200(0)-200(C). For example, if a workload executed on processor core 200(0) is associated with a high QoS level, the DVFS circuitry 206 and / or the cluster power management circuitry 208 can place the synchronous core cluster 102(0) in a higher performance state (corresponding to a higher frequency) to meet the high QoS requirements. This causes all processor cores 200(0)-200(C) to operate at a higher processor frequency. If, for example, processor core 200(C) is concurrently executing workloads requiring a low QoS level, then processor core 200(C) must still execute at a higher processor frequency, which results in unnecessary power consumption for processor core 200(C).
[0030] In this regard, the synchronous core cluster 102(0) provides throttling selection circuitry 210 and throttling circuitry 212, which are configured to provide dynamic microarchitectural throttling of processor cores 200(0)-200(C) based on QoS levels. As used herein, “microarchitectural throttling” refers to modifying the efficiency of instruction execution on one or more of the processor cores 200(0)-200(C) without modifying the frequency or voltage at which the synchronous core cluster 102(0) is operating. It will be understood that, although Figure 2 The throttling selection circuit 210 is shown as a separate element from the cluster power management circuit 208, but some aspects may specify that the throttling selection circuit 210 and the cluster power management circuit 208 may be integrated into a single element.
[0031] Using processor core 200(0) as an example, in exemplary operation, throttling selection circuit 210 determines the performance state of processor core 200(0) (in... Figure 2 The state is labeled "Performance Status" 214. Performance Status 214 represents the current frequency and voltage combination under which the synchronization core cluster 102(0) (and, through extension, processor cores 200(0)-200(C)) is currently operating. Throttling selection circuit 210 also receives QoS level 216 associated with the workload scheduled for execution by processor core 200(0). QoS level 216 is specified by the OS executing on processor device 100 and represents the quality of service level requested by the OS for the execution workload.
[0032] Then, the throttling selection circuit 210 performs a series of operations, which in some respects may occur periodically. For example, the throttling selection circuit 210 may determine an EPP level 218 corresponding to QoS level 216, wherein EPP level 218 includes an indicator having a value defined by processor device 100 as representing a bias toward performance or energy efficiency of the system, wherein different values for EPP level 218 are associated with different frequency and voltage preferences. Because the number of EPP levels supported by synchronous core cluster 102(0) and the number of QoS levels supported by the OS can vary, the throttling selection circuit 210 may include multiple mapped registers corresponding to processor cores 200(0)-200(C) (in Figure 3The mapping registers 220(0)-220(C) are labeled as "mapping registers". Each of the mapping registers 220(0)-220(C) can be periodically updated by the throttling selection circuit 210 to map the current QoS level of the workload executed on the corresponding processor cores 200(0)-200(C) to an EPP level supported by the synchronous core cluster 102(0). Thus, for example, mapping register 220(0) can map QoS level 216 to EPP level 218 for processor core 200(0). Alternatively or additionally, in some aspects, the throttling selection circuit 210 can map QoS level 216 to EPP level 218 based on the current performance state 214 of processor core 200(0).
[0033] Then, the throttling selection circuit 210 determines the throttling level 222 of the processor core 200(0) based, for example, on EPP level 218 and performance state 214, and on QoS level 216 and performance state 214. The throttling level 222 represents the degree to which the performance of the processor core 200(0) should be reduced such that, when operating in performance state 214, the instruction execution rate of the processor core 200(0) corresponds to EPP level 218. In some aspects, the throttling level 222 may include a value between zero (0) and 15, where a value of zero (0) indicates no throttling, and a value of 15 indicates the highest throttling level (i.e., the lowest instruction execution rate).
[0034] Throttling selection circuit 210 provides throttling level 222 to throttling circuit 212 of synchronous core cluster 102(0), which then performs microarchitectural throttling of processor core 200(0) based on throttling level 222. This causes processor core 200(0) to execute instructions at a slower effective rate than would have occurred under current performance state 214, and reduces the power consumption of processor core 200(0). For example, throttling circuit 212 can perform microarchitectural throttling by inserting NOP instructions (not shown) for processor core 200(0) to execute. When executed by processor core 200(0), the NOP instruction delays the execution of other instructions by processor core 200(0) (thus producing an effective rate of instruction execution corresponding to EPP level 218), while causing processor core 200(0) to consume less power.
[0035] In some respects, the throttling selection circuit 210 may use multiple throttling level LUTs 224(0)-224(C) corresponding to processor cores 200(0)-200(C) to determine the throttling level 222. The following is relative to... Figure 2The throttling levels LUT 224(0)-224(C) and exemplary aspects of the operations for accessing and filling the throttling levels LUT 224(0)-224(C) are discussed in more detail.
[0036] In some aspects, the DVFS aggregator circuit 226 may be further specified to determine the performance state in which the synchronization core cluster 102(0) is set to operate. In such aspects, processor cores 200(0)-200(C) provide corresponding multiple EPP hints 228 to indicate the desired EPP level for each of the processor cores 200(0)-200(C). Upon receiving an EPP hint 228, the DVFS aggregator circuit 226 selects a cluster performance state for the synchronization core cluster 102(0) based on the EPP hint 228 (in... Figure 3 The cluster performance state is labeled as “cluster performance state” 230. According to some of these aspects, the DVFS aggregator circuit 226 may employ multiple mapping LUTs 232(0)-232(C) corresponding to multiple processor cores 200(0)-200(C), and map each of the EPP hints 228 to the performance state of the corresponding processor core 200(0)-200(C). Therefore, selecting the cluster performance state 230 may include selecting the highest performance state indicated by the multiple mapping LUTs 232(0)-232(C) based on the EPP hints 228. The DVFS aggregator circuit 226 sends the cluster performance state 230 to the DVFS circuit 206, which then sets the frequency and voltage of the synchronization core cluster 102(0) based on the cluster performance state 230.
[0037] Figure 2 More detailed examples are given based on several aspects. Figure 3 The throttling level LUT 224(0). For example... Figure 2 As seen, the throttling level LUT 224(0) comprises multiple rows 300(0)-300(E), each corresponding to EPP levels 302(0)-302(E). Thus, for example, EPP level 302(0) may represent an energy-saving level associated with a lower frequency and voltage combination, EPP level 302(1) may represent an energy balance level associated with an intermediate frequency and voltage combination, and EPP level 302(E) may represent a performance level associated with a higher frequency and voltage combination. The columns of the throttling level LUT 224(0) represent multiple throttling levels 304(0)-304(T), arranged in order of decreasing throttling levels. Thus, throttling level 304(0) may represent the lowest throttling level (e.g., no throttling), while throttling level 304(T) may represent the highest throttling level.
[0038] The entries in the throttling level LUT 224(0) represent the throttling level LUT 224(0). Figure 3 The performance status of processor core 200(0) (in) Figure 2 The states are labeled “Performance States” (306(0,0)-306(E,T). Each performance state in performance states 306(0,0)-306(E,T) indicates the performance state of processor core 200(0), which, when microarchitectural throttling is applied at the level indicated by the corresponding throttling levels 304(0)-304(T), will result in corresponding EPP levels 302(0)-302(E). For example, if processor core 200(0) is placed in performance state 306(0,1) and the corresponding throttling level 304(2) is applied, the resulting performance of processor core 200(0) will correspond to EPP level 302(0).
[0039] As mentioned above, in contrast to Figure 2 As noted, in some aspects, when determining the throttling level 222 of processor core 200(0) based on EPP level 302(0) and performance state 214 (e.g., a performance state representing the current combination of frequency and voltage under which processor core 200(0) is operating), throttling selection circuit 210 may use throttling level LUT 224(0). In such aspects, determining throttling level 222 may involve throttling selection circuit 210 selecting and... Figure 3 The throttling level 222 is determined based on the column of the lowest performance state in row 300(0) of performance state 214 that is greater than or equal to the performance state 214 of processor core 200(0). Thus, for example, if performance state 306(0,2) in row 300(0) is the lowest among performance states 306(0,0)-306(0,T) that are greater than or equal to performance state 214, the throttling selection circuit 210 will select... Figure 2 The throttling level 304(1) is used as Figure 4 The throttling level is 222.
[0040] In some aspects, the throttling selection circuit 210 may also be specified to periodically fill each of the plurality of throttling levels LUT 224(0)-224(C) corresponding to the plurality of processor cores 200(0)-200(C). In such aspects, the throttling selection circuit 210 may perform a series of operations for each of the plurality of EPP levels 302(0)-302(E). The throttling selection circuit 210 may first calculate the average core frequency corresponding to each EPP level 302(0)-302(E). Then, the throttling selection circuit 210 calculates the corresponding performance states 306(0,0)-306(E,T) of the processor core 200(0) that require the throttling level to achieve at least the average core frequency for each of the plurality of throttling levels 304(0)-304(T).
[0041] Figure 2 Examples are given based on some aspects Figure 2 The throttling selection circuit 210 and Figure 3 and Figure 4 An exemplary physical implementation of the throttling level LUT224(0). Figure 3 In the LUT 224(0), the throttling level comprises k register rows 400(0)-400(k-1), which includes multiple registers 402(0,0)-402(k-1,15), where registers 402(0,0)-402(0,15) constitute register row 400(0), registers 402(i,0)-402(i,15) constitute register row 400(i), and registers 402(k-1,0)-402(k-1,15) constitute register row 400(k-1). Register rows 400(0)-400(k-1) correspond to Figure 4 Multiple rows 300(0)-300(E), each associated with an EPP level. Figure 4 In the example, register line 400(0) is associated with EPP level 0, which represents the energy-optimized EPP level. Register line 400(i) is associated with EPP level i, which represents the energy-balanced EPP level, and register line 400(k-1) is associated with EPP level k-1, which represents the performance EPP level.
[0042] like Figure 2As seen, the location of each register in registers 402(0,0)-402(k-1,15) within register rows 400(0)-400(k-1) corresponds to a throttling level, arranged in descending order. Thus, for example, register 402(0,0) is associated with the lowest throttling level (e.g., no throttling), while register 402(0,15) could represent the highest throttling level (e.g., throttling level 15 / 16, indicating that the processor core performance is throttled to 1 / 16 of the unthrottled performance). Each register in registers 402(0,0)-402(k-1,15) is filled with indicators... Figure 4 The performance status of processor core 200(0) (in) Figure 1 The index marked "P" corresponds to the throttling level associated with that register. The performance state indicates the performance state of processor core 200(0), which, when microarchitectural throttling is applied at the level indicated by the corresponding throttling level, will produce the corresponding EPP level associated with register lines 400(0)-400(k-1). Therefore, for example, if processor core 200(0) is placed in performance state 306(0,1) and the corresponding throttling level 0 / 16 is applied, the resulting performance of processor core 200(0) will correspond to EPP level 0 associated with register line 400(0).
[0043] In an exemplary operation, as indicated by arrow 406, the throttling selection circuit 210 inputs the EPP level corresponding to the QoS associated with the workload scheduled for execution by processor core 200(0) into selection logic 404. In this example, selection logic 404 determines that register line 400(i) corresponds to the EPP level and therefore selects register line 400(i) for further processing. Then, as indicated by arrow 410, the throttling selection circuit 210 inputs the performance state of processor core 200(0) into comparison logic elements 408(0)-408(15) corresponding to registers 402(i,0)-402(i,15). Each of the comparison logic elements 408(0)-408(15) determines whether the performance state of processor core 200(0) is greater than or equal to the performance state stored in the corresponding registers 402(i,0)-402(i,15) (i.e., performance state P[i,0]-P[i,15]). The result is routed as a 16-bit value to throttling selection logic 412, where the bit with a value of 1 indicates that the performance state of processor core 200(0) is greater than or equal to the performance state stored in the corresponding registers 402(i,0)-402(i,15). Throttling selection logic 412 determines which of the performance states P[i,0]-P[i,15] is the lowest performance state greater than or equal to the performance state of the processor core, and outputs a 16-bit threshold level with a bit set to a value of 1 to indicate which throttling level should be applied, as indicated by arrow 414.
[0044] To illustrate, based on some aspects, Figures 5A to 5D The processor device 100 performs exemplary operations for performing dynamic microarchitectural throttling of the processor core based on QoS levels. Figures 5A to 5D A flowchart illustrating exemplary operation 500 is provided. For clarity, in the description... Figures 1 to 3 When quoting Figure 5A The components. It will be understood that, in some respects, some operations in exemplary operation 500 may be performed in a different order than that illustrated herein and / or may be omitted.
[0045] Example operation 500 in Figure 1 It begins in processor devices (such as...) Figure 1 The synchronous core cluster of the processor device 100 (e.g., Figure 2 and Figure 2 The throttling selection circuit of the synchronous core cluster 102(0) (e.g., Figure 2 The throttling selection circuit 210) determines the processor cores (e.g., among the multiple processor cores of the synchronous core cluster 102(0)) Figure 2The performance status of processor core 200(0) in multiple processor cores 200(0)-200(C) (e.g., Figure 2 Performance status 214 (box 502). Throttling selection circuit 210 receives the QoS level (e.g., performance status 214) associated with the workload scheduled for execution by processor core 200(0). Figure 3 QoS level 216 (box 504).
[0046] In some aspects, a series of operations are then performed at periodic intervals (box 506). According to some of these aspects, the throttling selection circuit 210 fills multiple throttling level LUTs (such as LUTs) corresponding to the multiple processor cores 200(0)-200(C). Figure 3 Each throttling LUT (box 508) in the throttling level LUTs 224(0)-224(C)). Some such aspects may specify the operation of box 508 for filling each throttling LUT including for multiple EPP levels (such as Figure 3 A series of operations are performed for each EPP level (block 510) in the EPP levels 302(0)-302(E). In this respect, the throttling selection circuit 210 can calculate the average core frequency corresponding to the EPP level (block 512). Then, the throttling selection circuit 210 targets multiple throttling levels (e.g., Figure 5B For each throttling level in throttling levels 304(0)-304(T)), calculate the corresponding performance state (such as) of processor core 200(0) that requires a throttling level to achieve at least the average core frequency. Figure 5B Performance states 306(0,0)-306(E,T) (Box 514). Exemplary operation 500 in Figure 2 Continue at frame 516.
[0047] Now go to Figure 2 The operation continues at periodic intervals based on some aspects (box 506). The throttling selection circuit 210 determines the throttling level of the processor core 200(0) based on the QoS level 216 and the performance state 214 (e.g., Figure 2 The throttling level 222 (box 516). Some aspects may specify that the operation of box 516 for determining the throttling level 222 may include throttling selection circuitry 210 determining the EPP level corresponding to QoS level 216 (such as...). Figure 2 EPP level 218 (box 518). According to some aspects, the operation of box 518 for determining EPP level 218 may include based on a mapping register that maps QoS level 216 to EPP level 218 (e.g., Figure 2The mapping registers 220(0)-220(C) are used to determine EPP level 218 (box 520). Some aspects may specify that the operations of box 518 for determining EPP level 218 include based on the performance state of processor core 200(0) (e.g., Figure 2 Performance status 214) maps QoS level 216 to EPP level 218 (box 522).
[0048] In some aspects, the operation of block 516 for determining throttling level 222 may include throttling selection circuitry 210 at multiple throttling level LUTs (e.g., Figure 3 The throttling level LUT corresponding to processor core 200(0) in the throttling level LUT 224(0)-224(C)) (e.g., Figure 3 Select multiple rows of throttling level LUT 224(0) in the throttling level LUT 224(0) (e.g., Figure 5C The rows corresponding to EPP level 218 in rows 300(0)-300(E)) (such as Figure 5C Row 300(0) (box 524). In this respect, the throttling selection circuit 210 then determines the throttling level 222 (box 526) based on the column of the lowest performance state in row 300(0) of the performance state 214 that is greater than or equal to the processor core 200(0). Exemplary operation 500 then Figure 2 Continue at box 528.
[0049] Now for reference Figure 2 In some respects, the operations performed at periodic intervals continue (box 506). Throttling selection circuit 210 throttles the synchronization core cluster 102(0) (e.g., Figure 2 The throttling circuit 212 provides a throttling level 222 (box 528). The throttling circuit 212 receives the throttling level 222 (box 530). The throttling circuit 212 then performs microarchitecture throttling of the processor core 200(0) based on the throttling level 222 (box 532). In some aspects, the operation of the throttling circuit 212 for performing microarchitecture throttling may include the throttling circuit 212 inserting NOP instructions for execution by the processor core 200(0) (box 534).
[0050] Some aspects may specify the DVFS aggregator circuitry of the synchronization core cluster 102(0) (e.g., Figure 5D The DVFS aggregator circuit 226 receives multiple corresponding EPP prompts (such as...) from multiple processor cores 200(0)-200(C). Figure 5D EPP prompt 228 (box 536). Exemplary operation 500 then Figure 2 Continue at frame 538.
[0051] Now go to Figure 2 The DVFS aggregator circuit 226 selects the cluster performance state of the synchronization core cluster 102(0) based on multiple EPP hints 228 (e.g., Figure 2 Cluster performance state 230 (box 538). Depending on some aspects, the operations in box 538 for selecting cluster performance state 230 may include multiple mapped LUTs (such as...) corresponding to multiple processor cores 200(0)-200(C). Figure 1 The cluster performance state 230 is selected based on the mapping LUTs 232(0)-232(C) (box 540). In some aspects, the operation of box 540 for selecting the cluster performance state 230 based on the mapping LUTs 232(0)-232(C) may include selecting the highest performance state indicated by a plurality of mapping LUTs 232(0)-232(C) (box 542).
[0052] DVFS aggregator circuit 226 directs DVFS circuits (e.g., to the synchronization core cluster 102(0)) to the core cluster 102(0). Figure 6 The DVFS circuit 206 sends cluster performance status 230 (box 544). The DVFS circuit 206 receives cluster performance status 230 from the DVFS aggregator circuit 226 (box 546). The DVFS circuit 206 then sets the frequency and voltage of the synchronization core cluster 102(0) based on the cluster performance status 230 (box 548).
[0053] Based on the information disclosed in this article and referenced Figure 1 The processor devices discussed in these aspects can be located in or integrated into any processor-based device. Examples, without limitation, include: set-top boxes, entertainment units, navigation devices, communication devices, fixed location data units, mobile location data units, Global Positioning System (GPS) devices, mobile phones, cellular phones, smartphones, Session Initiation Protocol (SIP) phones, tablets, phablets, servers, computers, portable computers, mobile computing devices, laptops, wearable computing devices (e.g., smartwatches, health or fitness trackers, glasses, etc.), desktop computers, personal digital assistants (PDAs), monitors, computer monitors, televisions, tuners, radios, satellite radios, music players, digital music players, portable music players, digital video players, video players, digital video disc (DVD) players, portable digital video players, automobiles, vehicle components, avionics systems, drones, and multirotor aircraft.
[0054] In this respect, Figure 1 An example of a processor-based device 600 is illustrated, which is, relative to... Figure 6As illustrated and described. In this example, processor-based device 600 includes processor device 602, which functionally corresponds to Figure 6 The processor device 100 includes one or more processor cores 604 coupled to a cache memory 606. The processor core 604 is also coupled to a system bus 608 and can interactively couple to devices included in the processor-based device 600. As is well known, the processor core 604 communicates with these other devices by exchanging address, control, and data information on the system bus 608. For example, the processor core 604 can communicate bus transaction requests to the memory controller 610. Although in Figure 1 Not illustrated, but multiple system buses 608 may be provided, each of which constitutes a different architecture.
[0055] Other devices can be connected to system bus 608. For example... As illustrated, these devices may include a memory system 612, one or more input devices 614, one or more output devices 616, one or more network interface devices 618, and one or more display controllers 620. Input devices 614 may include any type of input device, including but not limited to input keys, switches, voice processors, etc. Output devices 616 may include any type of output device, including but not limited to audio, video, other visual indicators, etc. Network interface devices 618 may be any device configured to allow data exchange to and from network 622. Network 622 may be any type of network, including but not limited to wired or wireless networks, private or public networks, local area networks (LANs), wireless local area networks (WLANs), wide area networks (WANs), and Bluetooth. ™ Networks and the Internet. Network interface device 618 can be configured to support any type of communication protocol desired. Memory system 612 may include a memory controller 610 coupled to one or more memory arrays 624. Display controller may include, for example... GPU 104.
[0056] The processor core 604 can also be configured to access the display controller 620 via the system bus 608 to control the transmission of information to one or more displays 626. The display controller 620 transmits information to be displayed to the displays 626 via one or more video processors 628, which process the information to be displayed into a format suitable for the displays 626. The displays 626 may include any type of display, including but not limited to cathode ray tube (CRT), liquid crystal display (LCD), plasma display, light-emitting diode (LED) display, etc.
[0057] Those skilled in the art will further understand that the various exemplary logic blocks, modules, circuits, and algorithms described in connection with the aspects disclosed herein can be implemented as electronic hardware, stored in memory or another computer-readable medium and executed by a processor or other processing device, or a combination of both. As an example, the master and slave devices described herein can be employed in any circuit, hardware component, integrated circuit (IC), or IC chip. The memory disclosed herein can be of any type and size and can be configured to store any type of information desired. To clearly illustrate this interchangeability, the functionality of the various exemplary components, blocks, modules, circuits, and steps has been generally described above. How such functionality is implemented depends on the specific application, design choices, and / or design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure.
[0058] The various exemplary logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein may be implemented or executed using a processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic unit, discrete hardware component, or any combination thereof, designed to perform the functions described herein. The processor may be a microprocessor, but in alternative embodiments, it may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration).
[0059] The aspects disclosed herein may be embodied in hardware and instructions stored in the hardware, and may reside in, for example, random access memory (RAM), flash memory, read-only memory (ROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disks, removable disks, CD-ROMs, or any other form of computer-readable medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a remote station. Alternatively, the processor and storage medium may reside as discrete components in a remote station, base station, or server.
[0060] It should also be noted that the operational steps described in any of the exemplary aspects of this document are described for the purpose of providing examples and discussion. The described operations may be performed in many different orders other than the order illustrated. Furthermore, the operations described in a single operational step may actually be performed in multiple different steps. Additionally, one or more operational steps discussed in the exemplary aspects may be combined. It should be understood that, as will be apparent to those skilled in the art, many different modifications may be made to the operational steps illustrated in the flowcharts. Those skilled in the art will also understand that any of a variety of different techniques and arts can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.
[0061] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein can be applied to other variations. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0062] Specific implementation examples are described in the following numbered clauses.
[0063] 1. A processor device, the processor device comprising: Synchronization core cluster, the synchronization core cluster includes: Multiple processor cores; Throttling selection circuit; and Throttling circuit; The throttling selection circuit is configured as follows: Determine the performance status of the processor cores among the plurality of processor cores; Receive from the processor core the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; The throttling level of the processor core is determined based on the QoS level and the performance status; and The throttling level is provided to the throttling circuit; and The throttling circuit is configured as follows: Receive the throttling level; and The processor core's microarchitectural throttling is performed based on the throttling level.
[0064] 2. The processor device according to Clause 1, wherein the throttling selection circuit is configured to determine the throttling level and provide the throttling level to the throttling circuit at periodic intervals.
[0065] 3. The processor device according to any one of clauses 1 to 2, wherein: The throttling selection circuit is further configured to determine an Energy Performance Preference (EPP) level corresponding to the QoS level; and The throttling selection circuit is configured to determine the throttling level of the processor core based on the QoS level and the performance state by being configured to determine the throttling level of the processor core based on the EPP level and the performance state.
[0066] 4. The processor device according to Clause 3, wherein: The throttling selection circuit includes a mapping register that maps the QoS level to the EPP level; and The throttling selection circuit is configured to determine the EPP level based on the mapping register.
[0067] 5. The processor device according to any one of clauses 3 to 4, wherein the throttling selection circuit is configured to determine the EPP level by being configured to map the QoS level to the EPP level based on the performance state of the processor core.
[0068] 6. The processor device according to any one of clauses 3 to 5, wherein: The synchronous core cluster also includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels. The throttling selection circuit is configured to determine the throttling level of the processor core by being configured to perform the following operations: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; and The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
[0069] 7. The processor device according to Clause 6, wherein the throttling selection circuitry is further configured to populate each of the plurality of throttling level LUTs by being configured to perform the following operations: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
[0070] 8. A processor device according to any one of clauses 1 to 7, wherein the throttling circuit is configured to perform the microarchitectural throttling of the processor core by being configured to insert a No-Operation (NOP) instruction for execution by the processor core.
[0071] 9. The processor device according to any one of clauses 1 to 8, wherein the synchronization core cluster further comprises: Dynamic Voltage and Frequency Scaling (DVFS) aggregator circuits; and DVFS circuit; The DVFS aggregator circuit is configured as follows: Receive multiple corresponding EPP prompts from the multiple processor cores; The cluster performance state of the synchronization core cluster is selected based on the multiple EPP prompts; and Send the cluster performance status to the DVFS circuit; and The DVFS circuit is configured as follows: Receive the cluster performance status from the DVFS aggregator circuit; and The frequency and voltage of the synchronization core cluster are set based on the cluster performance status.
[0072] 10. The processor device according to Clause 9, wherein: The synchronization core cluster also includes multiple mapping lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints from the plurality of EPP hints to the corresponding performance state; and The DVFS aggregator circuit selects the cluster performance state based on the multiple mapped LUTs.
[0073] 11. The processor device according to Clause 10, wherein the DVFS aggregator circuitry is configured to select the cluster performance state based on the plurality of mapped LUTs by being configured to select the highest performance state indicated by the plurality of mapped LUTs.
[0074] 12. The processor device according to any one of Clauses 1 to 11, wherein the processor device is integrated into a device selected from the group consisting of: set-top boxes; entertainment units; navigation devices; communication devices; fixed location data units; mobile location data units; global positioning system (GPS) devices; mobile phones; cellular phones; smartphones; session initiation protocol (SIP) phones; tablet computers; tablet phones; servers; computers; portable computers; mobile computing devices; wearable computing devices; desktop computers; personal digital assistants (PDAs); monitors; computer monitors; televisions; tuners; radios; satellite radios; music players; digital music players; portable music players; digital video players; video players; digital video disc (DVD) players; portable digital video players; automobiles; vehicle components; avionics systems; unmanned aerial vehicles; and multirotor aircraft.
[0075] 13. A processor device, the processor device comprising: A component used to determine the performance status of a processor core among multiple processor cores in a synchronous core cluster of the processor device. A component for receiving from the processor core the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; Components for determining the throttling level of the processor core based on the QoS level and the performance state; and A component for performing microarchitectural throttling of the processor core based on the throttling level.
[0076] 14. A method for performing dynamic microarchitectural throttling in a processor core based on Quality of Service (QoS) levels, the method comprising: The performance state of a processor core among multiple processor cores in the synchronous core cluster is determined by the throttling selection circuit of the synchronous core cluster of the processor device. The QoS level associated with the workload scheduled for execution by the processor core is received by the throttling selection circuit. The throttling selection circuit determines the throttling level of the processor core based on the QoS level and the performance status; The throttling level is provided by the throttling selection circuit to the throttling circuit of the synchronization core cluster; The throttling level is received by the throttling circuit; and The throttling circuit performs microarchitectural throttling of the processor core based on the throttling level.
[0077] 15. The method according to Clause 14, the method further comprising determining the throttling level and providing the throttling level to the throttling circuit at periodic intervals.
[0078] 16. The method according to any one of Clauses 14 to 15, the method further comprising determining an Energy Performance Preference (EPP) level corresponding to the QoS level; Determining the throttling level of the processor core based on the QoS level and the performance status includes determining the throttling level of the processor core based on the EPP level and the performance status.
[0079] 17. The method according to Clause 16, wherein: The throttling selection circuit includes a mapping register that maps the QoS level to the EPP level; and The EPP level is determined based on the mapping register.
[0080] 18. The method according to any one of Clauses 16 to 17, wherein determining the EPP level includes mapping the QoS level to the EPP level based on the performance state of the processor core.
[0081] 19. The method according to any one of Clauses 16 to 18, wherein: The synchronous core cluster includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required for a corresponding throttling level among the multiple throttling levels to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels; and Determining the throttling level of the processor core includes: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; and The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
[0082] 20. The method according to Clause 19, further comprising populating each of the plurality of throttling level LUTs by: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
[0083] 21. The method according to any one of Clauses 14 to 20, wherein performing microarchitectural throttling of the processor core includes inserting No-Operation (NOP) instructions for the processor core to execute.
[0084] 22. The method according to any one of clauses 14 to 21, wherein the method further comprises: The synchronous core cluster receives multiple corresponding EPP prompts from the multiple processor cores via the dynamic voltage and frequency scaling (DVFS) aggregator circuitry. The DVFS aggregator circuit selects the cluster performance state of the synchronization core cluster based on the multiple EPP prompts; The DVFS aggregator circuit sends the cluster performance status to the DVFS circuit of the synchronization core cluster; The cluster performance status is received by the DVFS circuit from the DVFS aggregator circuit; and The DVFS circuit sets the frequency and voltage of the synchronization core cluster based on the cluster performance status.
[0085] 23. The method described according to Clause 22, wherein: The synchronization core cluster includes multiple lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints from the plurality of EPP hints to the corresponding performance state; and The cluster performance status is selected based on the multiple mapped LUTs.
[0086] 24. The method according to Clause 23, wherein selecting the cluster performance state based on the plurality of mapped LUTs includes selecting the highest performance state indicated by the plurality of mapped LUTs.
[0087] 25. A non-transitory computer-readable medium having stored thereon computer-executable instructions, which, when executed, cause a processor device of a processor-based device to: Determine the performance status of the processor cores among the multiple processor cores in the synchronous core cluster of the processor device; Receive the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; The throttling level of the processor core is determined based on the QoS level and the performance status; and The processor core's microarchitectural throttling is performed based on the throttling level.
[0088] 26. The non-transitory computer-readable medium according to Clause 25, wherein the computer-executable instructions cause the processor device to determine the throttling level and provide the throttling level to the throttling circuit at periodic intervals.
[0089] 27. A non-transitory computer-readable medium according to any one of clauses 25 to 26, wherein: The computer-executable instructions further cause the processor device to determine an Energy Performance Preference (EPP) level corresponding to the QoS level; and The computer-executable instructions cause the processor device to determine the throttling level of the processor core based on the QoS level and the performance state by causing the processor device to determine the throttling level of the processor core based on the EPP level and the performance state.
[0090] 28. The non-transitory computer-readable medium according to Clause 27, wherein the computer-executable instructions cause the processor device to determine the EPP level based on a mapping register that maps the QoS level to the EPP level.
[0091] 29. A nontransitory computer-readable medium according to any one of Clauses 27 to 28, wherein the computer-executable instructions cause the processor device to determine the EPP level by causing the processor device to map the QoS level to the EPP level based on the performance state of the processor core.
[0092] 30. A non-transitory computer-readable medium according to any one of clauses 27 to 29, wherein: The synchronous core cluster includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required for a corresponding throttling level among the multiple throttling levels to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels; and The computer-executable instructions cause the processor device to determine the throttling level of the processor core by causing the processor device to perform the following operations: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; and The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
[0093] 31. The non-transitory computer-readable medium according to Clause 30, wherein the computer-executable instructions further cause the processor device to populate each of the plurality of throttling-level LUTs by causing the processor device to perform the following operations: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
[0094] 32. A nontransitory computer-readable medium according to any one of Clauses 25 to 31, wherein the computer-executable instructions cause the processor device to perform microarchitectural throttling of the processor core by causing the processor device to insert no-operation (NOP) instructions for execution by the processor core.
[0095] 33. A non-transitory computer-readable medium according to any one of clauses 25 to 32, wherein the computer-executable instructions further cause the processor device to: Receive multiple corresponding EPP prompts from the multiple processor cores; The cluster performance state of the synchronization core cluster is selected based on the multiple EPP prompts; and The frequency and voltage of the synchronization core cluster are set based on the cluster performance status.
[0096] 34. The non-transitory computer-readable medium as described in Clause 33, wherein: The synchronization core cluster includes multiple lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints from the plurality of EPP hints to the corresponding performance state; and The computer-executable instructions cause the processor device to select the cluster performance state based on the plurality of mapped LUTs.
[0097] 35. The non-transitory computer-readable medium according to Clause 34, wherein the computer-executable instructions cause the processor device to select the cluster performance state based on the plurality of mapped LUTs by causing the processor device to select the highest performance state indicated by the plurality of mapped LUTs.
Claims
1. A processor device, the processor device comprising: Synchronization core cluster, the synchronization core cluster includes: Multiple processor cores; Throttling selection circuit; and Throttling circuit; The throttling selection circuit is configured as follows: Determine the performance status of the processor cores among the plurality of processor cores; Receive from the processor core the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; The throttling level of the processor core is determined based on the QoS level and the performance status; and The throttling level is provided to the throttling circuit; and The throttling circuit is configured as follows: Receive the throttling level; and The processor core's microarchitectural throttling is performed based on the throttling level.
2. The processor device of claim 1, wherein the throttling selection circuit is configured to determine the throttling level and provide the throttling level to the throttling circuit at periodic intervals.
3. The processor device according to claim 1, wherein: The throttling selection circuit is further configured to determine an Energy Performance Preference (EPP) level corresponding to the QoS level; and The throttling selection circuit is configured to determine the throttling level of the processor core based on the QoS level and the performance state by being configured to determine the throttling level of the processor core based on the EPP level and the performance state.
4. The processor device according to claim 3, wherein: The throttling selection circuit includes a mapping register that maps the QoS level to the EPP level; and The throttling selection circuit is configured to determine the EPP level based on the mapping register.
5. The processor device of claim 3, wherein the throttling selection circuit is configured to determine the EPP level by being configured to map the QoS level to the EPP level based on the performance state of the processor core.
6. The processor device according to claim 3, wherein: The synchronous core cluster also includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels. The throttling selection circuit is configured to determine the throttling level of the processor core by being configured to perform the following operations: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; as well as The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
7. The processor device of claim 6, wherein the throttling selection circuitry is further configured to populate each of the plurality of throttling level LUTs by being configured to perform the following operations: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
8. The processor device of claim 1, wherein the throttling circuit is configured to perform the microarchitectural throttling of the processor core by being configured to insert a No-Operation (NOP) instruction for execution by the processor core.
9. The processor device of claim 1, wherein the synchronization core cluster further comprises: Dynamic voltage and frequency scaling (DVFS) aggregator circuit; as well as DVFS circuit; The DVFS aggregator circuit is configured as follows: Receive multiple corresponding EPP prompts from the multiple processor cores; The cluster performance status of the synchronization core cluster is selected based on the multiple EPP prompts; as well as Send the cluster performance status to the DVFS circuit; and The DVFS circuit is configured as follows: Receive the cluster performance status from the DVFS aggregator circuit; as well as The frequency and voltage of the synchronization core cluster are set based on the cluster performance status.
10. The processor device according to claim 9, wherein: The synchronization core cluster also includes multiple mapping lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints in the plurality of EPP hints to the corresponding performance state; and The DVFS aggregator circuit selects the cluster performance state based on the multiple mapped LUTs.
11. The processor device of claim 10, wherein the DVFS aggregator circuitry is configured to select the cluster performance state based on the plurality of mapped LUTs by being configured to select the highest performance state indicated by the plurality of mapped LUTs.
12. The processor device of claim 1, wherein the processor device is integrated into a device selected from the group consisting of: set-top boxes; entertainment units; navigation devices; communication devices; fixed location data units; mobile location data units; global positioning system (GPS) devices; mobile phones; cellular phones; smartphones; session initiation protocol (SIP) phones; tablet computers; tablet phones; servers; computers; portable computers; mobile computing devices; wearable computing devices; desktop computers; personal digital assistants (PDAs); monitors; computer monitors; televisions; tuners; radios; satellite radios; music players; digital music players; portable music players; digital video players; video players; digital video disc (DVD) players; portable digital video players; automobiles; vehicle components; avionics systems; unmanned aerial vehicles; and multirotor aircraft.
13. A processor device, the processor device comprising: A component used to determine the performance status of a processor core among multiple processor cores in a synchronous core cluster of the processor device. A component for receiving from the processor core the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; A component for determining the throttling level of the processor core based on the QoS level and the performance status; as well as A component for performing microarchitectural throttling of the processor core based on the throttling level.
14. A method for performing dynamic microarchitectural throttling in a processor core based on Quality of Service (QoS) levels, the method comprising: The performance state of a processor core among multiple processor cores in the synchronous core cluster is determined by the throttling selection circuit of the synchronous core cluster of the processor device. The QoS level associated with the workload scheduled for execution by the processor core is received by the throttling selection circuit. The throttling selection circuit determines the throttling level of the processor core based on the QoS level and the performance status; The throttling level is provided by the throttling selection circuit to the throttling circuit of the synchronization core cluster; The throttling level is received by the throttling circuit; as well as The throttling circuit performs microarchitectural throttling of the processor core based on the throttling level.
15. The method of claim 14, further comprising determining the throttling level and providing the throttling level to the throttling circuit at periodic intervals.
16. The method of claim 14, further comprising determining an Energy Performance Preference (EPP) level corresponding to the QoS level; Determining the throttling level of the processor core based on the QoS level and the performance status includes determining the throttling level of the processor core based on the EPP level and the performance status.
17. The method of claim 16, wherein: The throttling selection circuit includes a mapping register that maps the QoS level to the EPP level; and The EPP level is determined based on the mapping register.
18. The method of claim 16, wherein determining the EPP level includes mapping the QoS level to the EPP level based on the performance state of the processor core.
19. The method of claim 16, wherein: The synchronous core cluster includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required for a corresponding throttling level among the multiple throttling levels to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels. and Determining the throttling level of the processor core includes: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; as well as The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
20. The method of claim 19, further comprising populating each of the plurality of throttling level LUTs by: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
21. The method of claim 14, wherein performing microarchitectural throttling of the processor core includes inserting no-operation (NOP) instructions for the processor core to execute.
22. The method according to claim 14, further comprising: The synchronous core cluster receives multiple corresponding EPP prompts from the multiple processor cores via the dynamic voltage and frequency scaling (DVFS) aggregator circuitry. The DVFS aggregator circuit selects the cluster performance state of the synchronization core cluster based on the multiple EPP prompts; The DVFS aggregator circuit sends the cluster performance status to the DVFS circuit of the synchronization core cluster; The cluster performance status is received by the DVFS circuit from the DVFS aggregator circuit; as well as The DVFS circuit sets the frequency and voltage of the synchronization core cluster based on the cluster performance status.
23. The method according to claim 22, wherein: The synchronization core cluster includes multiple lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints in the plurality of EPP hints to the corresponding performance state; and The cluster performance status is selected based on the multiple mapped LUTs.
24. The method of claim 23, wherein selecting the cluster performance state based on the plurality of mapped LUTs includes selecting the highest performance state indicated by the plurality of mapped LUTs.
25. A non-transitory computer-readable medium having stored thereon computer-executable instructions, which, when executed, cause a processor device of a processor-based device to: Determine the performance status of the processor cores among the multiple processor cores in the synchronous core cluster of the processor device; Receive the Quality of Service (QoS) level associated with the workload scheduled for execution by the processor core; The throttling level of the processor core is determined based on the QoS level and the performance status; as well as The processor core's microarchitectural throttling is performed based on the throttling level.
26. The non-transitory computer-readable medium of claim 25, wherein the computer-executable instructions cause the processor device to determine the throttling level and provide the throttling level to the throttling circuit at periodic intervals.
27. The non-transitory computer-readable medium according to claim 25, wherein: The computer-executable instructions further cause the processor device to determine an Energy Performance Preference (EPP) level corresponding to the QoS level; and The computer-executable instructions cause the processor device to determine the throttling level of the processor core based on the QoS level and the performance state by causing the processor device to determine the throttling level of the processor core based on the EPP level and the performance state.
28. The non-transitory computer-readable medium of claim 27, wherein the computer-executable instructions cause the processor device to determine the EPP level based on a mapping register that maps the QoS level to the EPP level.
29. The non-transitory computer-readable medium of claim 27, wherein the computer-executable instructions cause the processor device to determine the EPP level by causing the processor device to map the QoS level to the EPP level based on the performance state of the processor core.
30. The non-transitory computer-readable medium according to claim 27, wherein: The synchronous core cluster includes multiple throttling level lookup tables (LUTs) corresponding to the multiple processor cores. Each throttling level LUT includes multiple entries, which are organized into multiple rows corresponding to multiple EPP levels and multiple columns corresponding to multiple throttling levels. Each entry indicates the minimum performance state required for a corresponding throttling level among the multiple throttling levels to achieve the average core frequency of the corresponding EPP level among the multiple EPP levels; and The computer-executable instructions cause the processor device to determine the throttling level of the processor core by causing the processor device to perform the following operations: In the plurality of throttling level LUTs corresponding to the processor core, select the row corresponding to the EPP level from the plurality of rows of the throttling level LUT; as well as The throttling level is determined based on the column of the lowest performance state in the row of the performance states that are greater than or equal to those of the processor core.
31. The non-transitory computer-readable medium of claim 30, wherein the computer-executable instructions further cause the processor device to fill each of the plurality of throttling level LUTs by causing the processor device to perform the following operations: For each of the multiple EPP levels: Calculate the average core frequency corresponding to the EPP level; and For each of the plurality of throttling levels, the corresponding performance state of the processor core required to achieve at least the average core frequency is calculated.
32. The non-transitory computer-readable medium of claim 25, wherein the computer-executable instructions cause the processor device to perform microarchitectural throttling of the processor core by causing the processor device to insert no-operation (NOP) instructions for execution by the processor core.
33. The non-transitory computer-readable medium of claim 25, wherein the computer-executable instructions further cause the processor device to: Receive multiple corresponding EPP prompts from the multiple processor cores; The cluster performance status of the synchronization core cluster is selected based on the multiple EPP prompts; as well as The frequency and voltage of the synchronization core cluster are set based on the cluster performance status.
34. The non-transitory computer-readable medium according to claim 33, wherein: The synchronization core cluster includes multiple lookup tables (LUTs) corresponding to the multiple processor cores. Each of the plurality of mapping LUTs maps the EPP hints in the plurality of EPP hints to the corresponding performance state; and The computer-executable instructions cause the processor device to select the cluster performance state based on the plurality of mapped LUTs.
35. The non-transitory computer-readable medium of claim 34, wherein the computer-executable instructions cause the processor device to select the cluster performance state based on the plurality of mapped LUTs by causing the processor device to select the highest performance state indicated by the plurality of mapped LUTs.