Systems, apparatus, and methods for optimal processor throttling
Optimal throttling of multi-core processor elements based on user activity and power capabilities addresses the inefficiencies of conventional systems by maintaining performance and optimizing power usage without immediate degradation.
Patent Information
- Application Number
- JP2024000031
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-02
- Filing Date
- 2024-01-04
- Publication Date
- 2025-10-06
- Estimated Expiration
- 2039-04-02
AI Technical Summary
Conventional power management systems in multi-core processors often force processors to operate at their lowest operating point in response to platform events, leading to undesirable performance penalties and inefficiencies.
A processor is configured to optimally throttle processing elements based on user activity and power capabilities, allowing for dynamic control of operating points to maintain performance without sacrificing protection mechanisms, and enabling throttling at levels higher than the minimum operating point when necessary.
This approach maintains appropriate performance levels while preserving protection mechanisms, avoiding unnecessary performance degradation and optimizing power usage.
Smart Images

Figure 0007749709000001 
Figure 0007749709000002 
Figure 0007749709000003
Abstract
Description
[Technical Field]
[0001] FIELD Embodiments relate to power management in systems, and more particularly to power management in multi-core processors. [Background technology]
[0002] Advances in semiconductor processes and logic design have enabled an increase in the amount of logic that can reside on an integrated circuit device. As a result, computer system configurations have evolved from single or multiple integrated circuits within a system to multiple hardware threads, multiple cores, multiple devices, and / or complete systems on individual integrated circuits. Furthermore, as integrated circuit density has increased, the power requirements of computing systems (from embedded systems to servers) have also gradually increased. Furthermore, software inefficiencies and their hardware requirements have also increased computing device energy consumption. In fact, several studies have shown that in countries such as the United States, computing devices consume a significant percentage of the total electricity supply. As a result, there is a critical need for energy efficiency and management related to integrated circuits. These needs will increase as servers, desktop computers, notebooks, Ultrabooks™, tablets, mobile phones, processors, embedded systems, and the like become more widespread (from their inclusion in standard computers, automobiles, and televisions to biotechnology). [Brief explanation of the drawings]
[0003] [Figure 1] 1 is a block diagram of a portion of a system according to one embodiment of the present invention.
[0004] [Figure 2] FIG. 2 is a block diagram of a processor according to one embodiment of the present invention.
[0005] [Figure 3] FIG. 2 is a block diagram of a multi-domain processor according to another embodiment of the present invention.
[0006] [Figure 4] 1 is an embodiment of a processor including multiple cores.
[0007] [Figure 5] FIG. 2 is a block diagram of the microarchitecture of a processor core according to one embodiment of the present invention.
[0008] [Figure 6] FIG. 2 is a block diagram of the microarchitecture of a processor core according to another embodiment.
[0009] [Figure 7] FIG. 10 is a block diagram of the microarchitecture of a processor core according to yet another embodiment.
[0010] [Figure 8] FIG. 10 is a block diagram of the microarchitecture of a processor core according to yet another embodiment.
[0011] [Figure 9] FIG. 4 is a block diagram of a processor according to another embodiment of the present invention.
[0012] [Figure 10] FIG. 1 is a block diagram of an exemplary SoC according to one embodiment of the present invention.
[0013] [Figure 11] FIG. 2 is a block diagram of another exemplary SoC according to one embodiment of the present invention.
[0014] [Figure 12] 1 is a block diagram of an exemplary system according to one embodiment of the present invention.
[0015] [Figure 13] FIG. 2 is a block diagram of another exemplary system in which embodiments may be implemented.
[0016] [Figure 14] FIG. 1 is a block diagram of a representative computer system.
[0017] [Figure 15] 1 is a block diagram of a system according to one embodiment of the present invention.
[0018] [Figure 16] FIG. 1 is a block diagram of a system according to one embodiment.
[0019] [Figure 17] FIG. 2 is a flow diagram of a method according to one embodiment of the present invention.
[0020] [Figure 18] FIG. 4 is a flow diagram of a method according to another embodiment of the present invention.
[0021] [Figure 19] FIG. 4 is a flow diagram of a method according to another embodiment of the present invention.
[0022] [Figure 20] FIG. 4 is a flow diagram of a method according to another embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] In various embodiments, a processor, such as a multi-core processor or other system-on-chip (SoC), may be configured to optimally throttle one or more processing elements of the processor in response to a platform event. That is, rather than immediately throttling all operations within the processor to a minimum operating point, processor performance may be optimized in a throttled condition without sacrificing protection mechanisms that protect against electrical faults in light of a platform event. More specifically, as described herein, a first agent, i.e., a platform target agent, may proactively determine an optimal throttle power target for the entire processor. Meanwhile, another agent, referred to herein as a power management agent, may proactively determine optimal throttle targets for each processing element based at least in part on user activity, priority information, etc. An additional agent, also referred to herein as a throttle agent, may then dynamically control the operating points of the processing elements based at least in part on these optimized processing element throttle targets and the overall optimal throttle power target. Such dynamically controlled operating points may be optimized. As a result, the processor does not have to immediately degrade to its minimum operating point.
[0024] According to the embodiments described herein, appropriate throttling may be maintained for the processor while preserving performance capabilities. That is, the throttle level optimized as described herein may be set at a level where a typical load can still operate at an appropriate performance level. However, it is understood that in worst-case conditions, such as virus load levels, a full throttling mechanism may be implemented to force all processing elements to operate at their minimum operating points.
[0025] In an embodiment, such optimized throttle operation may occur based on input power capability, i.e., when a given power source, such as a battery power source, is fully charged, throttle operation need not be limited to the minimum operating point because the battery power source can provide enough power to operate at throttle levels greater than the minimum operating point.
[0026] Conversely, conventional platforms typically force the processor to operate at its lowest operating point (e.g., lowest frequency and / or voltage mode) in response to a platform event, which can have an undesirable impact on performance. Instead, in embodiments, based at least in part on a priori knowledge of input power capabilities (e.g., charging capacity), processor throttling in response to a platform event may be at a level higher than the lowest operating point, thereby avoiding unnecessary performance penalties.
[0027] When the platform detects that there is a potential malfunction, the platform generates a fast signal to quickly throttle processor power to avoid the malfunction. In this situation, a processor according to an embodiment may be configured to optimize performance when throttled without sacrificing protection mechanisms (to avoid the malfunction). To this end, a processor mechanism is provided to proactively set an optimal throttle power target for the processor and to proactively set an optimal operating point for each processing element based at least in part on user activity. After these proactive settings are generated, when a throttle signal is received (i.e., in response to detecting a potential malfunction), the processing elements may be controlled with low latency to operate at a point no higher than the throttle power target of the respective processing element.
[0028] Although the following embodiments are described with reference to particular integrated circuits, such as computing platforms or processors, other embodiments are applicable to other types of integrated circuits and logic devices. Techniques and teachings similar to those of the embodiments described herein may be applied to other types of circuits or semiconductor devices that benefit from good energy efficiency and energy conservation. For example, the disclosed embodiments are not limited to any particular type of computer system. That is, the disclosed embodiments may be used in many different system types, ranging from server computers (e.g., tower, rack, blade, microserver, etc.), communications systems, storage systems, desktop computers of any configuration, laptops, notebooks, and tablet computers (including 2:1 tablets, phablets, etc.), and may be used within other devices, such as handheld devices, systems-on-chips (SoCs), and embedded applications. Some examples of handheld devices include cellular phones, such as smartphones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and handheld PCs. Embedded applications may typically include microcontrollers, digital signal processors (DSPs), network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, wearable devices, or any other system capable of performing the functions and operations taught below. Furthermore, embodiments may be implemented in mobile terminals with standard voice capabilities, such as mobile phones, smartphones, and phablets, and / or non-mobile terminals without standard wireless voice communication capabilities, such as many wearables, tablets, notebooks, desktops, microservers, servers, etc. Furthermore, the apparatus, methods, and systems described herein are not limited to physical computing devices, but may also relate to software optimization.
[0029] Referring to Figure 1, a block diagram of a portion of a system according to one embodiment of the present invention is shown. As shown in Figure 1, system 100 may include various components, including processor 110, which as shown is a multi-core processor. Processor 110 may be coupled to power supply 150 via external voltage regulator 160. External voltage regulator 160 may perform a first voltage conversion to provide a primary regulated voltage to processor 110.
[0030] As shown, processor 110 may be a single-die processor including multiple cores 120a-120n. Additionally, each core may be associated with an integrated voltage regulator (IVR) 125a-125n. The IVR receives a primary regulated voltage and generates an operating voltage to be provided to one or more processor agents associated with the IVR. Accordingly, an IVR implementation may be provided to enable fine-grained control of the voltage, and therefore the power and performance, of individual cores. Thus, each core can operate at an independent voltage and frequency, allowing for great flexibility and providing extensive opportunities for balancing power consumption and performance. In some embodiments, the use of multiple IVRs allows components to be grouped into separate power planes. As a result, power is regulated by the IVR and supplied only to components within the group. During power management, a given power plane of one IVR may be powered down or turned off when the processor is placed into a particular low-power state, while another power plane of another IVR remains active or fully powered.
[0031] 1 , additional components may be present within the processor, including an input / output interface 132, another interface 134, and an integrated memory controller 136. As shown, each of these components may be powered by another integrated voltage regulator 125x. In one embodiment, interface 132 may enable operation of an Intel® Quick Path Interconnect (QPI) interconnect. The QPI interconnect provides a point-to-point (PtP) link in a cache coherent protocol that includes multiple layers, including a physical layer, a link layer, and a protocol layer. Additionally, interface 134 may communicate via the PCIe™ (Peripheral Component Interconnect Express) protocol.
[0032] Also shown is a power control unit (PCU) 138, which may include hardware, software, and / or firmware that performs power management operations with respect to processor 110. As shown, PCU 138 provides control information to external voltage regulator 160 via a digital interface, causing the voltage regulator to generate an appropriate regulated voltage. PCU 138 also provides control information to IVR 125 via another digital interface to control the generated operating voltage (or cause the corresponding IVR to be disabled in a low-power mode). In various embodiments, PCU 138 may include various power management logic units to perform hardware-based power management. Such power management may be fully processor controlled (e.g., by various processor hardware) (and may be triggered by load and / or power, thermal, or other processor constraints). And / or power management may be performed in response to an external source (e.g., platform or supervisory power management source or system software).
[0033] In embodiments herein, PCU 138 may be configured to control domains within processor 110 to operate at one of a plurality of different operating points. More specifically, PCU 138 may determine the operating points of the different domains. Additionally, a high-speed mechanism, e.g., external to PCU 138, may be involved in performing throttling actions in response to platform events. Accordingly, as further shown in FIG. 1 , throttle circuit 139 (which may in some cases be a distributed hardware circuit) may further communicate with IVR 125 (and / or cores 120 themselves). In this manner, throttle circuit 139 translates processor throttle power levels into corresponding throttle levels for individual domains (e.g., cores 120, etc.) based at least in part on translation information received from PCU 138, as further described below.
[0034] Although not shown for ease of explanation, it is understood that additional components, such as uncore logic, and other components, such as internal memory, e.g., one or more levels of a cache memory hierarchy, etc., may be present within processor 110. Furthermore, although the implementation of FIG. 1 is shown with an integrated voltage regulator, embodiments are not limited thereto.
[0035] The processors described herein may utilize power management techniques that may be independent of and complementary to operating system (OS)-based power management (OSPM) mechanisms. According to one exemplary OSPM technique, a processor can operate in various performance states or levels, so-called P-states, i.e., P0 through P-N. Typically, the P1 performance state may correspond to the highest guaranteed performance state that may be requested by the OS. In addition to this P1 state, the OS may also request a higher performance state, i.e., P0 state. This P0 state therefore allows the processor hardware to configure the processor, or at least portions thereof, to operate at frequencies higher than the guaranteed frequency when power and / or thermal budget is available. In many implementations, a processor may include multiple so-called bin frequencies, either as fuses or programmed into the processor during manufacturing, that exceed the maximum peak frequency of a particular processor above the P1 guaranteed maximum frequency. Furthermore, according to one OSPM mechanism, a processor can operate in various power states or levels. With respect to power states, the OSPM mechanism may specify different power consumption states, generally referred to as C-states, C0, C1, through Cn states. A core executes in the C0 state when active, and when idle, the core may be placed in a core low-power state, also called a non-zero C-state (e.g., C1-C6 states), with each C-state being a lower power consumption level (so C6 is a deeper low-power state than C1, etc.).
[0036] It will be appreciated that many different types of power management techniques may be used individually or in combination in different embodiments. As a representative example, a power controller may control a processor to be power managed through a particular form of dynamic voltage frequency scaling (DVFS). In DVFS, the operating voltage and / or operating frequency of one or more cores or other processor logic may be dynamically controlled to reduce power consumption in certain circumstances. In one example, DVFS may be implemented using Enhanced Intel SpeedStep™ technology available from Intel Corporation of Santa Clara, CA to provide optimal performance at the lowest power consumption level. In another example, DVFS may be implemented using Intel TurboBoost™ technology to cause one or more cores or other computing engines to operate at higher than their guaranteed operating frequency based on conditions (e.g., load and availability).
[0037] Another power management technique that may be used in certain examples is dynamic swapping of loads between different compute engines. For example, a processor may include asymmetric cores or other processing engines that operate at different power consumption levels. Thus, in power-constrained situations, one or more loads can be dynamically switched to run on a lower-power core or other compute engine. Another exemplary power management technique is hardware duty cycling (HDC). HDC periodically enables and disables cores and / or other compute engines according to a duty cycle, such that one or more cores may be inactive during inactive periods of the duty cycle and active during active periods of the duty cycle.
[0038] Embodiments may be implemented in processors for various markets, including server processors, desktop processors, mobile processors, etc. Referring to Figure 2, a block diagram of a processor according to one embodiment of the present invention is shown. As shown in Figure 2, processor 200 may be a multi-core processor including multiple cores 210a-210n. In one embodiment, each such core may be of an independent power domain and may be configured to enter and exit an active state and / or a maximum performance state based on load.
[0039] The various cores may be coupled via interconnect 215 to an uncore 220, which may include a system agent, or various components. As shown, uncore 220 may include a shared cache 230, which may be a last-level cache. Additionally, uncore 220 may include an integrated memory controller 240 for communicating with system memory (not shown in FIG. 2 ), e.g., via a memory bus. Uncore 220 also includes various interfaces 250, a performance monitoring unit (PMU) 260, and a power control unit 255, which may include logic to perform the power management techniques described herein. Additionally, power control unit 255 may include throttle control circuitry 256. Throttle control circuitry 256 may dynamically determine translation information used by throttle circuitry 258 based at least in part on hint information to enable the circuitry to dynamically determine throttle power levels for individual cores 210 (and potentially other components of processor 200) and enable low-latency operating point changes in response to platform events.
[0040] Additionally, interfaces 250a-250n may provide connections to various off-chip components such as peripherals, mass storage, etc. While the embodiment of Figure 2 is illustrated with this particular implementation, the scope of the present invention is not limited in this respect.
[0041] Referring to Figure 3, a block diagram of a multi-domain processor according to another embodiment of the present invention is shown. As shown in the embodiment of Figure 3, a processor 300 includes multiple domains. Specifically, a core domain 310 includes multiple cores 3100-310. n , graphics domain 320 may include one or more graphics engines, and system agent domain 350 may also be present. In some embodiments, system agent domain 350 may run at a frequency independent of the core domains and may remain powered on at all times to handle power control events. As a result, domains 310 and 320 may be controlled to dynamically enter and exit higher and lower power states. Each of domains 310 and 320 may operate at different voltages and / or powers. It should be noted that while only three domains are shown, the scope of the invention is not limited in this respect, and it is understood that additional domains may be present in other embodiments. For example, there may be multiple core domains, each including at least one core.
[0042] Typically, each core 310 may further include a lower level cache and additional processing elements in addition to various execution units. The various cores may also communicate with each other and with multiple units 3400-3400 of last level cache (LLC). n In various embodiments, LLC 340 may be shared among the cores and the graphics engine, as well as various media processing circuits. As shown, ring interconnect 330 thus couples the cores together and provides interconnection between the cores, graphics domain 320, and system agent circuitry 350. In one embodiment, interconnect 330 may be part of the core domain; however, in other embodiments, the ring interconnect may be its own domain.
[0043] As further shown, system agent domain 350 may include a display control 352 that may provide control and interface for an associated display. As further shown, system agent domain 350 may include a power control unit 355 that may include a throttle control circuit 356. Throttle control circuit 356 determines conversion information based at least in part on hint information regarding the workloads running on cores 310, graphics engine 320, etc. As described herein, PCU 355 also provides such information to throttle circuit 358 to enable throttle circuit 358 to determine throttle operating points for the corresponding cores / graphics engines and enable low-latency transitions to such updated operating points.
[0044] As further illustrated in FIG. 3, processor 300 may further include an integrated memory controller (IMC) 370 that may provide an interface to system memory, such as dynamic random access memory (DRAM). n may be present to enable interconnection between the processor and other circuitry. For example, in one embodiment, at least one direct media interface (DMI) interface may be provided along with one or more PCIe™ interfaces. Furthermore, one or more QPI interfaces may also be provided to provide communication between other agents, such as additional processors or other circuitry. While shown at a high level in the embodiment of FIG. 3, it will be understood that the scope of the present invention is not limited in this respect.
[0045] 4 illustrates one embodiment of a processor including multiple cores. Processor 400 may include any processor or processing device, such as a microprocessor, embedded processor, digital signal processor (DSP), network processor, handheld processor, application processor, co-processor, system-on-chip (SoC), or other device that executes code. Processor 400, in one embodiment, includes at least two cores, cores 401 and 402, which may include asymmetric or symmetric cores (in the illustrated embodiment). However, processor 400 may include any number of processing elements, which may be symmetric or asymmetric.
[0046] In one embodiment, a processing element represents hardware or logic that supports a software thread. Examples of hardware processing elements include a thread unit, a thread slot, a thread, a processing unit, a context, a context unit, a logical processor, a hardware thread, a core, and / or any other element capable of maintaining processor state, such as execution state or architectural state. In other words, a processing element, in one embodiment, represents any hardware capable of being independently associated with code, such as a software thread, an operating system, an application, or other code. A physical processor typically represents an integrated circuit that may contain any number of other processing elements, such as cores or hardware threads.
[0047] A core sometimes represents logic located on an integrated circuit capable of maintaining independent architectural states, where each independently maintained architectural state is associated with at least some dedicated execution resources. In contrast to a core, a hardware thread typically represents any logic located on an integrated circuit capable of maintaining independent architectural states, where the independently maintained architectural states share access to executing resources. As shown, when certain resources are shared and others are dedicated to architectural states, the lines between the representations of hardware threads and cores overlap. Furthermore, cores and hardware threads are viewed by the operating system as separate logical processors, where the operating system can schedule operations on each logical processor independently.
[0048] Physical processor 400 as shown in FIG. 4 includes two cores, cores 401 and 402. Here, cores 401 and 402 are considered to be symmetric cores, i.e., cores with the same configuration, functional units, and / or logic. In another embodiment, core 401 includes an out-of-order processor core, and core 402 includes an in-order processor core. However, cores 401 and 402 may be independently selected from any type of core, such as a native core, a software-managed core, a core adapted to execute a native instruction set architecture (ISA), a core adapted to execute a translated ISA, a co-designed core, or other known core. For further discussion, the functional units shown in core 401 are described in further detail below, and the units in core 402 operate similarly.
[0049] As shown, core 401 includes two hardware threads 401a and 401b, which may also be referred to as hardware thread slots 401a and 401b. Thus, in one embodiment, a software entity, such as an operating system, views processor 400 as four individual processors, i.e., four logical processors or four software threads, as processing elements capable of simultaneously executing. As described above, a first thread is associated with architecture state register 401a, a second thread is associated with architecture state register 401b, a third thread is associated with architecture state register 402a, and a fourth thread is associated with architecture state register 402b. Here, each of the architecture state registers (401a, 401b, 402a, 402b) may be referred to as a processing element, thread slot, or thread unit, as described above. As shown, architecture state register 401a is replicated in architecture state register 401b. Thus, individual architecture state / context can be stored for logical processor 401a and logical processor 401b. Within core 401, other smaller resources, such as the instruction pointer and remaining logic in allocator and renamer block 430, may also be duplicated for threads 401a and 401b. Some resources, such as the reloader buffer, ILTB 420, load / store buffers, and queues in reloader / retirement unit 435, may be shared through partitioning. Other resources, such as general-purpose internal registers, page-table base registers, low-level data caches, and portions of data TLB 415, execution unit 440, and out-of-order unit 435, may be fully shared.
[0050] Processor 400 sometimes includes other resources that may be fully shared, shared through partitioning, or dedicated by / to a processing element. One embodiment of a purely exemplary processor is shown in FIG. 4 , with the processor's illustrative logical units / resources. The processor may include or omit any of these functional units and may include any other known functional units, logic, or firmware not shown. As shown, core 401 includes a simplified representative out-of-order (OOO) processor core. However, in different embodiments, an in-order processor may be utilized. The OOO core includes a branch target buffer 420 that predicts branches to be executed / taken, and an instruction-translation buffer (I-TLB) 420 that stores address translation entries for instructions.
[0051] Core 401 further includes a decode module 425 coupled to fetch unit 420 to decode fetched elements. The fetch logic, in one embodiment, includes individual sequences associated with thread slots 401a, 401b, respectively. Typically, core 401 is associated with a first ISA, which defines / specifies instructions executable on processor 400. Sometimes, machine code instructions that are part of the first ISA include a portion of the instruction (called an opcode) that references / specifies the instruction or operation to be performed. Decode logic 425 includes circuitry that recognizes these instructions from their opcodes and passes the decoded instructions down the pipeline for processing as defined by the first ISA. For example, decoder 425, in one embodiment, includes logic designed or adapted to recognize specific instructions, such as transactional instructions. As a result of recognition by decoder 425, the architecture or core 401 performs a specific, predetermined action to execute the task associated with the appropriate instruction. It is important to note that any of the tasks, blocks, operations, and methods described herein may be performed in response to single or multiple instructions, some of which may be new or old.
[0052] In one example, allocator and renamer block 430 includes an allocator that reserves resources such as a register file to store instruction processing results. However, threads 401a and 401b may be capable of out-of-order execution. Here, allocator and renamer block 430 also reserves other resources such as a reorder buffer to track instruction results. Unit 430 may also include a register renamer to rename program / instruction reference registers to other registers internal to processor 400. Reloader / retirement unit 435 includes components such as the reorder buffer, load buffer, and store buffer described above to support out-of-order execution and in-order retirement of instructions that are subsequently executed out-of-order.
[0053] Scheduler and execution units block 440, in one embodiment, includes a scheduler unit that schedules instructions / operations on the execution units. For example, floating point instructions are scheduled to a subset of the execution units that have an available floating point execution unit. Register files associated with the execution units are also included for storing information instruction processing results. Exemplary execution units include a floating point execution unit, an integer execution unit, a jump execution unit, a load execution unit, a store execution unit, and other known execution units.
[0054] A lower level data cache and data translation buffer (interpretation) 450 is coupled to the execution unit 440. The data cache stores recently used / operated on elements such as data operands, whose state may be kept coherently in memory. The D-TLB stores recent virtual / linear to physical address translations. As a particular example, a processor may include a page table structure to decompose physical memory into multiple virtual pages.
[0055] Here, cores 401 and 402 share access to a higher-level, or further-out, cache 410 for caching recently fetched elements. Note that higher or further out refers to cache levels that are further incremented or further away from the execution units. In one embodiment, higher-level cache 410 is a last-level data cache, such as a second or third level data cache, i.e., the last cache in the memory hierarchy in processor 400. However, higher-level cache 410 is not limited as it may be associated with or include an instruction cache. A trace cache, i.e., an instruction-based cache, may alternatively be coupled after decoder 425 for storing recently decoded traces.
[0056] In the illustrated configuration, processor 400 also includes a bus interface module 405 and a power controller 460, which may perform power management in accordance with embodiments of the present invention. In this scenario, bus interface 405 communicates with devices external to processor 400, such as system memory and other components.
[0057] The memory controller 470 may interface with other devices, such as one or more memories. In one example, the bus interface 405 includes a ring interconnect with a memory controller for interfacing with memory and a graphics controller for interfacing with a graphics processor. In an SoC environment, even more devices, such as a network interface, coprocessors, memory, a graphics processor, and any other known computing device / interface, may be integrated onto a single die or integrated circuit to provide a small form factor with high functionality and low power consumption.
[0058] Referring to Figure 5, a block diagram of the microarchitecture of a processor core according to one embodiment of the present invention is shown. As shown in Figure 5, processor core 500 may be a multi-stage pipelined out-of-order processor. Core 500 may operate at various voltages based on a received operating voltage, which may be received from an integrated voltage regulator or an external voltage regulator.
[0059] 5, the core 500 includes a front-end unit 510 that can be used to fetch and prepare instructions to be executed for later use in the processor pipeline. For example, the front-end unit 510 may include a fetch unit 501, an instruction cache 503, and an instruction decoder 505. In some implementations, the front-end unit 510 may further include a trace cache, along with microcode storage and micro-operation storage. The fetch unit 501 may fetch macro-instructions, for example from memory or the instruction cache 503, and provide them to the instruction decoder 505, which decodes them into primitives, or micro-operations, for execution by the processor.
[0060] Coupled between the front-end unit 510 and the execution unit 520 is an out-of-order (OOO) engine 515, which may be used to receive microinstructions and prepare them for execution. More specifically, the OOO engine 515 may include various buffers for reordering the microinstruction flow and allocating various resources needed for execution, as well as for providing remaining logical registers in storage locations within various register files, such as register file 530 and extended register file 535. Register file 530 may include separate register files for integer and floating-point operations. A set of machine-specific registers (MSRs) 538 may also be present and accessible to various logic within (and outside) the core 500 for configuration, control, and additional operations.
[0061] For example, various resources, including various integer, floating-point, and single instruction multiple data (SIMD) logic units, among other dedicated hardware, may reside within execution unit 520. For example, such an execution unit may include one or more arithmetic logic units (ALUs) 522 and one or more vector execution units 524, among other such execution units.
[0062] Results from the execution units may be provided to retirement logic, or reorder buffer (ROB) 540. More specifically, ROB 540 may include various arrays and logic that receive information associated with instructions being executed. This information is then examined by ROB 540 to determine whether the instructions can be validly retired and whether the resulting data can be committed to the processor's architectural state, or whether one or more exceptions have occurred that prevent the instructions from being properly retired. Of course, ROB 540 may handle other operations related to retirement.
[0063] As shown in FIG. 5 , ROB 540 is coupled to cache 550. Cache 550, in one embodiment, may be a lower-level cache (e.g., an L1 cache). However, the scope of the present invention is not limited in this respect. Execution unit 520 may also be directly coupled to cache 550. From cache 550, data communication may occur with higher-level caches, system memory, etc. The performance and energy efficiency capabilities of core 500 may vary based on workload and / or processor constraints. As such, a power controller (not shown in FIG. 5 ) may dynamically determine an appropriate configuration for all or a portion of processor 500 based at least in part on thermal setpoints determined as described herein. While the embodiment of FIG. 5 is shown at a high level, it will be understood that the scope of the present invention is not limited in this respect. For example, the implementation of FIG. 5 relates to an out-of-order machine such as the Intel® x86 Instruction Set Architecture (ISA), although the scope of the present invention is not limited in this respect. That is, other embodiments may be implemented in an in-order processor, a RISC (reduced instruction set computing) processor such as an ARM-based processor, or a processor of another type of ISA that can emulate the instructions and operations of a different ISA via an emulation engine and associated logic.
[0064] Referring to FIG. 6, a block diagram of the microarchitecture of a processor core according to another embodiment is shown. In the embodiment of FIG. 6, core 600 may be a low-power core of a different microarchitecture, such as an Intel® Atom™-based processor having a relatively limited pipeline depth designed to reduce power consumption. As shown, core 600 includes an instruction cache 610 coupled to provide instructions to an instruction decoder 615. A branch predictor 605 may be coupled to instruction cache 610. Notably, instruction cache 610 may be further coupled to another level of cache memory, such as an L2 cache (not shown in FIG. 6 for ease of illustration). Additionally, instruction decoder 615 provides decoded instructions to an issue queue 620 for storage and delivery to a given execution pipeline. A microcode ROM 618 is coupled to instruction decoder 615.
[0065] Floating-point pipeline 630 includes a floating-point register file 632, which may include multiple architectural registers of a given bit length, such as 128, 256, or 512 bits. Pipeline 630 includes a floating-point scheduler 634 that schedules instructions for execution on one or more execution units of the pipeline. In the illustrated embodiment, such execution units include an ALU 635, a shuffle unit 636, and a floating-point adder 638, and results produced by these execution units may be returned to buffers and / or registers in register file 632. Of course, while shown with these few exemplary execution units, it will be understood that additional or different floating-point execution units may be present in alternative embodiments.
[0066] An integer pipeline 640 may also be provided. In the illustrated embodiment, pipeline 640 includes an integer register file 642, which may include multiple architectural registers of a given bit length, such as 128 or 256 bits. Pipeline 640 includes an integer scheduler 644 that schedules instructions for execution on one or more execution units of the pipeline. In the illustrated embodiment, such execution units include an ALU 645, a shifter unit 646, and a jump execution unit 648. Results produced by these execution units may also be returned to buffers and / or registers in register file 642. Of course, while shown with these few exemplary execution units, it will be understood that additional or different integer execution units may be present in alternative embodiments.
[0067] A memory execution scheduler 650 may schedule memory operations for execution in an address generation unit 652, which is coupled to the TLB 654. As will be appreciated, these structures may be coupled to a data cache 660, which may be an L0 and / or L1 data cache. The data cache 660 also couples to additional levels of the cache memory hierarchy, including an L2 cache memory.
[0068] To provide support for out-of-order execution, an allocator / renamer 670 may be provided, as well as a reorder buffer 680 configured to reorder instructions executed out-of-order into in-order for retirement. Note that the performance and energy efficiency capabilities of core 600 may vary based on workload and / or processor constraints. As such, a power controller (not shown in FIG. 6 ) may dynamically determine an appropriate configuration for all or a portion of processor 500 based at least in part on thermal setpoints determined as described herein. While the description of FIG. 6 is shown with this particular pipeline architecture, it will be understood that many variations and alternatives are possible.
[0069] Note that in processors with asymmetric cores, such as those according to the microarchitectures of Figures 5 and 6, workloads may be dynamically swapped between cores for power management reasons, and these cores may have different pipeline designs and depths, but the same or related ISA. Such dynamic core swapping may be performed in a manner that is transparent to user applications (and possibly also to the kernel).
[0070] Referring to FIG. 7, a block diagram of the microarchitecture of a processor core according to yet another embodiment is shown. As shown in FIG. 7, core 700 may include a multi-stage in-order pipeline that executes at very low power consumption levels. As one such example, processor 700 may have a microarchitecture according to the ARM Cortex A53 design available from ARM Holdings LTD., Sunnyvale, CA. In one implementation, an eight-stage pipeline configured to execute both 32-bit and 64-bit code may be provided. Core 700 includes a fetch unit 710 configured to fetch instructions and provide them to a decode unit 715. Decode unit 715 may decode instructions, e.g., macro-instructions of a given ISA, such as the ARMv8 ISA. Note further that a queue 730 may be coupled to decode unit 715 for storing the decoded instructions. The decoded instructions are provided to issue logic 725, where they may be issued to a given one of a plurality of execution units.
[0071] With further reference to FIG. 7 , issue logic 725 may issue a name to one of multiple execution units. In the illustrated embodiment, these execution units include an integer unit 735, a multiply unit 740, a floating point / vector unit 750, a dual issue unit 760, and a load / store unit 770. The results of these different execution units may be provided to a writeback unit 780. While a single writeback unit is shown for simplicity of explanation, it will be understood that in some implementations, a separate writeback unit may be associated with each of the execution units. Furthermore, while each of the units and logic shown in FIG. 7 is shown at a high level, a particular implementation may include more or different structures. A processor designed with one or more cores having a pipeline like that of FIG. 7 may be implemented in many different end products, from mobile devices to server systems.
[0072] Referring to FIG. 8, a block diagram of the microarchitecture of a processor core according to yet another embodiment is shown. As shown in FIG. 8, core 800 may include a multi-stage, multi-issue, out-of-order pipeline that executes at very high performance levels (potentially at higher power consumption levels than core 700 of FIG. 7). As one such example, processor 800 may have a microarchitecture that conforms to the AMR Cortex A57 design. In one implementation, a 15 (or more) stage pipeline configured to execute both 32-bit and 64-bit code may be provided. Furthermore, the pipeline may provide three (or more) wide and three (or more) instruction operations. Core 800 includes a fetch unit 810 configured to fetch instructions and provide them to a decoder / renamer / dispatcher 815. Decoder / renamer / dispatcher 815 may decode instructions, for example macro-instructions of the ARMv8 instruction set architecture, rename register references within the instructions, and (eventually) dispatch the instructions to a selected execution unit. The decoded instructions may be stored in queue 825. Note that for ease of explanation, a single queue structure is shown in Figure 8, but it will be understood that a separate queue may be provided for each of multiple different types of execution units.
[0073] 8 also shows issue logic 830, which may issue the decoded instructions stored in queue 825 to selected execution units. In particular embodiments, issue logic 830 may be implemented by separate issue logic for each of multiple different types of execution units to which issue logic 830 is coupled.
[0074] The decoded instruction may be issued to a given one of multiple execution units. In the illustrated embodiment, these execution units include one or more integer unit 835, multiply unit 840, floating point / vector unit 850, branch unit 860, and load / store unit 870. In one embodiment, floating point / vector unit 850 may be configured to process 128- or 256-bit SIMD or vector data. Furthermore, floating point / vector execution unit 850 may perform IEEE-754 double-precision floating-point operations. The results of these different execution units may be provided to writeback unit 880. It should be noted that in some implementations, a separate writeback unit may be associated with each of the execution units. Furthermore, although each of the units and logic shown in FIG. 8 is shown at a high level, a particular implementation may include more or different structures.
[0075] Note that in processors with asymmetric cores, such as those according to the microarchitectures of Figures 7 and 8, loads may be dynamically swapped for power management reasons, and these cores may have different pipeline designs and depths, but the same or related ISA. Such dynamic core swapping may be performed in a manner that is transparent to user applications (and possibly also to the kernel).
[0076] Processors designed with one or more cores having a pipeline such as in any one or more of FIGS. 5-8 may be implemented in many different end products, from mobile devices to server systems. Referring to FIG. 9, a block diagram of a processor according to another embodiment of the present invention is shown. In the embodiment of FIG. 9, processor 900 may be an SoC including multiple domains. Each of the domains may be controlled to operate at an independent voltage and operating frequency. As a specific illustrative example, processor 900 may be an Intel® Architecture Core™-based processor, such as the i3, i5, i7, or another such processor available from Intel Corporation. However, other low-power processors available from Advanced Micro Devices, Inc. (AMD) of Sunnyvale, CA, ARM-based designs by ARM Holdings, Ltd. or their licensees, or MIPS-based designs from MIPS Technologies, Inc. of Sunnyvale, CA, or their licensees or adopters, may alternatively be present, such as the Apple A7 processor, a Qualcomm Snapdragon processor, or a Texas Instruments OMAP processor, in other embodiments. Such SoCs may be used in low-power systems such as smartphones, tablet computers, phablet computers, Ultrabook™ computers, or other portable computing devices, or in-vehicle computing systems.
[0077] In the high-level diagram of FIG. 9, a processor 900 includes multiple core units 9100-910 nEach core unit may include one or more processor cores, one or more cache memories, and other circuitry. Each core unit 910 may support one or more instruction sets (e.g., the x86 instruction set (with any extensions added by newer versions)), the MIPS instruction set, the ARM instruction set (with any additional extensions such as NEON), or other instruction sets, or combinations thereof. Some of the core units may be heterogeneous (e.g., different designs). Additionally, each such core may be coupled to a cache memory (not shown), which in one embodiment may be a shared level-2 (L2) cache memory. Non-volatile storage 930 may be used to store various programs and other data. For example, this storage may be used to store at least a portion of microcode, boot information such as a BIOS, other system software, and other information. In an embodiment, non-volatile storage 930 may store multiple configurations, such as those described herein, which may be prioritized for use by firmware.
[0078] Each core unit 910 may also include an interface, such as a bus interface unit, that allows interconnection to additional circuitry of the processor. In one embodiment, each core unit 910 couples a coherent fabric, which may act as a first-level cache, to a coherent on-die interconnect. The on-die interconnect also couples to a memory controller 935, which in turn controls communication with memory, such as DRAM (not shown in FIG. 9 for ease of illustration).
[0079] In addition to the core units, additional processing engines are present in the processor, including at least one graphics processing unit (GPU) 920 that performs graphics processing and possibly general-purpose operations on a graphics processor (so-called GPGPU operations). Additionally, at least one image signal processor 925 may be present. The signal processor 925 may be configured to process input image data received from one or more capture devices within or off-chip in the SoC.
[0080] Other accelerators may also be present. In the illustration of FIG. 9, a video coder 950 may perform coding operations, including encoding and decoding video information, to provide hardware acceleration support, for example, for high-definition video content. A display controller 955 may also be provided to accelerate display operations, including providing support for internal and external displays in the system. Additionally, a security processor 945 may be present to perform security operations, such as secure boot operations, various encryption operations, etc.
[0081] Each unit may have its own power consumption, which is controlled via power manager 940. Power manager 940 may include control logic to implement the various power management techniques described herein, including dynamic determination of the appropriate configuration based on thermal point selection.
[0082] In some embodiments, SoC 900 may further include a non-coherent fabric coupled to a coherent fabric to which various peripheral devices may be coupled. One or more interfaces 960a-960d may communicate with one or more off-chip devices. Such communication may be via various communication protocols such as PCIe, GPIO, USB, I2C, UART, MIPI, SDIO, DDR, SPI, HDMI, among other types of communication protocols. While illustrated at a high level in the embodiment of FIG. 9, it will be understood that the scope of the present invention is not limited in this respect.
[0083] Referring to FIG. 10 , a block diagram of an exemplary SoC is shown. In the illustrated embodiment, SoC 1000 may be a multi-core SoC configured for optimized low-power operation for incorporation into other low-power devices, such as smartphones or tablet computers, or other portable or commercial computing devices. As an example, SoC 1000 may be implemented using asymmetric or different types of cores, such as a combination of high-power and / or low-power cores, e.g., out-of-order cores and in-order cores. In different embodiments, these cores may be based on an Intel® Architecture™ core design or an ARM architecture design. In yet other embodiments, a mix of Intel and ARM cores may be implemented in a given SoC.
[0084] As shown in FIG. 10 , SoC 1000 includes a first core domain 1010 having a plurality of first cores 10120-10123. In one example, these cores may be low-power cores, such as in-order cores, as described herein. In one embodiment, these first cores may be implemented as ARM Cortex A53 cores. These cores are also coupled to cache memory 1015n of core domain 1010. SoC 1000 also includes a second core domain 1020. In the illustration of FIG. 10 , second core domain 1020 has a plurality of second cores 10220-10223. In one example, these cores may be higher power consumption cores than first core 1012. In one embodiment, the second cores may be out-of-order cores, such as ARM Cortex A57 cores. These cores are also coupled to cache memory 1015n of core domain 1020. It should be noted that while the example shown in FIG. 10 includes four cores in each domain, it is understood that in other examples, more or fewer cores may be present in a given domain.
[0085] 10, a graphics domain 1030 is also provided that may include, for example, one or more graphics processing units (GPUs) configured to independently execute the graphics workload provided by one or more cores of core domains 1010 and 1020. As an example, GPU domain 1030 may be used to provide display support for various screen sizes in addition to providing graphics and display rendering operations.
[0086] As shown, the various domains couple to a coherent interconnect 1040, which in one embodiment may be a cache coherent interconnect fabric that couples to a unified memory controller 1050. The coherent interconnect 1040 may, in some examples, include a shared cache memory, such as an L3 cache. In one embodiment, the memory controller 1050 may be a direct memory controller that provides multiple channels of communication with off-chip memory, such as multiple channels of DRAM (not shown in FIG. 10 for ease of illustration).
[0087] In different examples, the number of core domains may vary. For example, in a low-power SoC suitable for incorporation into a mobile computing device, there may be a limited number of core domains, as shown in FIG. 10 . Furthermore, in such a low-power SoC, a core domain 1020 including higher power cores may have a fewer number of such cores. For example, in one implementation, two cores 1022 may be provided to enable operation at reduced power consumption levels. Furthermore, different core domains may also be coupled to interrupt controllers to enable dynamic swapping of loads between different domains.
[0088] In still other embodiments, a larger number of core domains and additional optional IP logic may be present, where the SoC can be scaled to higher performance (and power) levels for incorporation into other computing devices such as desktops, servers, high-performance computing systems, base stations, etc. As one such example, four core domains may be provided, each having a given number of out-of-order cores. Still further, in addition to optional GPU support (in the form of a GPGPU, as an example), one or more accelerators may also be provided to provide optimal hardware support for specific functions (e.g., web serving, network processing, switching, etc.). Furthermore, input / output interfaces may be present for coupling such accelerators to off-chip components.
[0089] Referring to FIG. 11 , a block diagram of another exemplary SoC is shown. In the embodiment of FIG. 11 , SoC 1100 may include various circuits to enable high performance for multimedia applications, communications, and other functions. Thus, SoC 1100 is suitable for incorporation into a wide variety of portable and other devices, such as smartphones, tablet computers, smart TVs, in-vehicle computing systems, and the like. In the illustrated example, SoC 1100 includes a central processing unit (CPU) domain 1110. In one embodiment, multiple individual processor cores may reside within CPU domain 1110. As an example, CPU domain 1110 may be a quad-core processor with four multithreaded cores. Such a processor may be a homogeneous or heterogeneous processor, e.g., a mix of low-power and high-power processor cores.
[0090] Additionally, GPU domain 1120 is provided for performing advanced graphics processing within one or more GPUs for processing graphics and computing APIs. DSP 1130 may provide one or more low-power DSPs for processing low-power multimedia applications such as music playback, audio / video, etc., in addition to advanced computations that may occur during the execution of multimedia instructions. Additionally, communication unit 1140 may include various components for providing connectivity via various wireless protocols such as cellular communications (including 3G / 4G LTE), Bluetooth, wireless local area protocols, IEEE 802.11, etc.
[0091] Furthermore, multimedia processor 1150 may be used to perform capture and playback of high-resolution video and audio content, including processing of user gestures. Sensor unit 1160 may include multiple sensors and / or a sensor controller that interfaces with various off-chip sensors present within a given platform. Image signal processor 1170 may be provided with one or more separate ISPs to perform image processing on captured content from one or more cameras of the platform, including still and video cameras.
[0092] Display processor 1180 may provide support for connection to high-resolution displays of a given pixel density, including the ability to wirelessly communicate content for playback on such displays. Furthermore, position unit 1190 may include a GPS receiver supporting multiple GPS satellites to provide applications with highly accurate positioning information obtained using the GPS receiver. While the example of Figure 11 is shown with this particular set of components, it will be understood that many variations and alternatives are possible.
[0093] FIG. 12 shows a block diagram of an exemplary system in which embodiments can be used. As shown, system 1200 may be a smartphone or other wireless communication device. Baseband processor 1205 is configured to perform various signal processing on communication signals to be transmitted from or received by the system. Baseband processor 1205 is also coupled to application processor 1210, which may be the main CPU of the system, for running the OS and other system software, as well as user applications such as many well-known social media and multimedia apps. Application processor 1210 may include power control and throttling circuitry as described herein and may be further configured to perform various other computing operations for the device.
[0094] The application processor 1210 may also be coupled to a user interface / display 1220, such as a touchscreen display. The application processor 1210 may further be coupled to a memory system including non-volatile memory, i.e., flash memory 1230, and system memory, i.e., dynamic random access memory (DRAM) 1235. As further shown, the application processor 1210 is coupled to a capture device 1240, such as one or more image capture devices capable of recording video and / or still images.
[0095] 12, a universal integrated circuit card (UICC) 1240 including a subscriber identity module and possibly secure storage and a crypto processor is also coupled to the application processor 1210. The system 1200 may further include a security processor 1250 that may be coupled to the application processor 1210. A number of sensors 1225 may be coupled to the application processor 1210 to allow for input of various sensed information, such as an accelerometer and other environmental information. An audio output device 1295 may provide an interface for outputting audio, for example, in the form of voice communication, playback or streaming of audio data, etc.
[0096] Further shown is a near field communication (NFC) contactless interface 1260 that communicates over NFC short range via an NFC antenna 1265. While separate antennas are shown in Figure 12, it will be appreciated that in some implementations a single antenna or different sets of antennas may be provided to enable various wireless functions.
[0097] A power management integrated circuit (PMIC) 1215 couples to the application processor 1210 to perform platform-level power management. To this end, the PMIC 1215 may issue power management requests to the application processor 1210 to enter specific low-power states as needed. Furthermore, based on platform constraints, the PMIC 1215 may also control the power levels of other components of the system 1200. Furthermore, as described herein, the PMIC 1215 may send a throttle signal to the application processor 1210 in response to a given platform event.
[0098] Various circuits may be coupled between the baseband processor 1205 and the antenna 1290 to transmit and receive communications. Specifically, a radio frequency (RF) transceiver 1270 and a wireless local area network (WLAN) transceiver 1275 may be present. Typically, the RF transceiver 1270 may be used to receive and transmit wireless data and calls according to a given wireless communication protocol, such as a 3G or 4G wireless communication protocol, such as according to code division multiple access (CDMA), global system for mobile communication (GSM), long term evolution (LTE), or other protocols. Additionally, a GPS sensor 1280 may be present. Other wireless communications, such as reception or transmission of radio signals, e.g., AM / FM and other signals, may also be provided. Additionally, local wireless communications may be realized via the WLAN transceiver 1275.
[0099] Figure 13 shows a block diagram of another exemplary system in which embodiments can be used. In the diagram of Figure 13, system 1300 may be a mobile, low-power system such as a tablet computer, a 2:1 tablet, a phablet, or other convertible or standalone tablet system. As shown, SoC 1310 is present and may be configured to operate as the device's application processor and may include power controls and throttle circuitry as described herein.
[0100] Various devices may be coupled to SoC 1310. In the illustrated diagram, a memory subsystem includes flash memory 1340 and DRAM 1345 coupled to SoC 1310. Additionally, a touch panel 1320 is coupled to SoC 1310 to provide display capabilities and tactile user input, including providing a virtual keyboard on the display of touch panel 1320. To provide wired network connectivity, SoC 1310 couples to Ethernet interface 1330. A peripheral hub 1325 is coupled to SoC 1310 to enable interfacing with various peripheral devices that may be coupled to system 1300 by any of a variety of ports or other connectors.
[0101] In addition to the internal power control circuits and functions within the SoC 1310, the PMIC 1380 is coupled to the SoC 1310 to provide platform-based power management, for example, based on whether the system is powered by a battery 1390 or an AC power source via an AC adapter 1395. In addition to this power source-based power management, the PMIC 1380 may also perform platform power management activities based on environmental and usage conditions. Furthermore, the PMIC 1380 may communicate control and status information to the SoC 1310 to effect various power management operations within the SoC 1310. Furthermore, as described herein, the PMIC 1380 may send a throttle signal to the SoC 1310 in response to a given platform event.
[0102] 13, a WLAN unit 1350 is coupled to the SoC 1310 and also coupled to an antenna 1355 to provide wireless capabilities. In various implementations, the WLAN unit 1350 may provide communications according to one or more wireless protocols.
[0103] As further shown, multiple sensors 1360 may be coupled to SoC 1310. These sensors may include various accelerometers, environmental and other sensors, including user gesture sensors. Finally, audio codec 1365 is coupled to SoC 1310 to provide an interface to audio output device 1370. Of course, while the example of FIG. 13 is shown with this particular implementation, it will be understood that many variations and alternatives are possible.
[0104] Referring to FIG. 14 , a block diagram of a representative computer system, such as a notebook, Ultrabook, or other small form factor system, is shown. Processor 1410, in one embodiment, includes a microprocessor, a multi-core processor, a multi-threaded processor, an ultra-low voltage processor, an embedded processor, or other known processing element. In the illustrated implementation, processor 1410 acts as a main processing unit and a central hub for communicating with many of the various components of system 1400. As an example, processor 1400 may be implemented as an SoC and may include power control and throttling circuitry as described herein.
[0105] The processor 1410, in one embodiment, is in communication with a system memory 1415. As an illustrative example, the system memory 1415 is implemented with multiple memory devices or modules to provide a given amount of system memory.
[0106] Mass storage 1420 may also be coupled to processor 1410 for permanent storage of information such as data, applications, one or more operating systems, etc. In various embodiments, this mass storage may be implemented with SSDs to improve system responsiveness and enable thin-and-light system designs. Alternatively, mass storage may be implemented primarily using hard disk drives (HDDs), with a smaller amount of SSD storage acting as an SSD cache, allowing nonvolatile storage of context state and other such information during power-down events. This may result in faster startup upon resumption of system activity. Also shown in FIG. 14, a flash device 1422 may be coupled to processor 1410, for example, via a serial peripheral interface (SPI). This flash device may provide nonvolatile storage of system software, including basic input / output software (BIOS) and other system firmware.
[0107] Various input / output (I / O) devices may be present in system 1400. Specifically, the embodiment of FIG. 14 shows a display 1424, which may be a high-resolution LCD or LED panel that also provides a touchscreen 1425. In one embodiment, the display 1424 may be coupled to the processor 1410 via a display interconnect, which may be implemented as a high-performance graphics interconnect. The touchscreen 1425 may be coupled to the processor 1410 via another interconnect, which in one embodiment may be an I2C interconnect. As further shown in FIG. 14 , in addition to the touchscreen 1425, tactile user input may also occur via a touchpad 1430. The touchpad 1430 may be configured within the housing and may be coupled to the same I2C interconnect as the touchscreen 1425.
[0108] For perceptual computing and other purposes, various sensors may be present in the system and may be coupled to the processor 1410 in different ways. Certain internal and environmental sensors may be coupled to the processor 1410 through a sensor hub 1440, for example, via an I2C interconnect. In the embodiment shown in FIG. 14 , these sensors may include an accelerometer 1441, an ambient light sensor (ALS) 1442, a compass 1443, and a gyroscope 1444. Other environmental sensors may include one or more thermal sensors 1446, which in some embodiments are coupled to the processor 1410 via a system management bus (SMBus).
[0109] Also in FIG. 14, various peripherals may be coupled to processor 1410 via a low pin count (LPC) interconnect. In the illustrated embodiment, various components may be coupled through embedded controller 1435. Such components may include a keyboard 1436 (e.g., coupled via a PS2 interface), a fan 1437, and a thermal sensor 1439. In some embodiments, touchpad 1430 may also be coupled to EC 1435 via a PS2 interface. Additionally, a security processor, such as a trusted platform module (TPM) 1438, may also be coupled to processor 1410 via this LPC interconnect.
[0110] System 1400 can communicate with external devices in a variety of ways, including wirelessly. In the embodiment shown in FIG. 14 , there are various wireless modules, each of which may correspond to a radio configured for a particular wireless communication protocol. One method for short-range, such as near-field, wireless communication may be via NFC unit 1445. NFC unit 1445, in one embodiment, may communicate with processor 1410 via SMBus. Note that via this NFC unit 1445, devices in close proximity to each other can communicate.
[0111] 14, the additional radio units may include other short-range radio engines including a WLAN unit 1450 and a Bluetooth unit 1452. Wi-Fi® communications may be achieved using the WLAN unit 1450, while short-range Bluetooth® communications may occur via the Bluetooth unit 1452. These units may communicate with the processor 1410 via a given link.
[0112] Additionally, wireless wide area communications, for example according to cellular or other wireless wide area protocols, may occur via WWAN unit 1456. WWAN unit 1456 may also be coupled to a subscriber identity module (SIM) 1457. Additionally, a GPS module 1455 may also be present to enable receipt and use of location information. It should be noted that in the embodiment shown in FIG. 14 , WWAN unit 1456 and an integrated capture device, such as camera module 1454, may communicate over a given link.
[0113] An integrated camera module 1454 can be incorporated within the lid. To provide audio input and output, an audio processor can be implemented by a digital signal processor (DSP) 1460. The DSP 1460 can be coupled to the processor 1410 via a high definition audio (HDA) link. Similarly, the DSP 1460 can communicate with an integrated coder / decoder (CODEC) and amplifier 1462, which can be coupled to an output speaker 1463, which can be implemented within the housing. Similarly, the amplifier and CODEC 1462 can be coupled to receive audio input from a microphone 1465. The microphone 1465, in one embodiment, can be implemented by a dual array microphone (such as a digital microphone array) to provide high-quality audio input and enable voice-activated control of various operations within the system. Note that audio output can also be provided from the amplifier / CODEC 1462 to a headphone jack 1464. While the embodiment of FIG. 14 is shown with these specific components, it will be understood that the scope of the present invention is not limited in this respect.
[0114] Embodiments may be implemented in many different system types. Referring to FIG. 15, a block diagram of a system according to one embodiment of the present invention is shown. As shown in FIG. 15, a multiprocessor system 1500 is a point-to-point interconnect system and includes a first processor 1570 and a second processor 1580 coupled via a point-to-point interconnect 1550. As shown in FIG. 15, each of the processors 1570 and 1580 may be a multimedia processor including a first processor core and a second processor core (i.e., processor cores 1574a and 1574b, and processor cores 1584a and 1584b), although in some cases more cores may be present in a processor. Each of the processors may include a PCU 1575, 1585, or other power management logic, and throttle circuits 1577, 1587.
[0115] Continuing with reference to FIG. 15 , first processor 1570 further includes a memory controller hub (MCH) 1572 and point-to-point (PP) interfaces 1576 and 1578. Similarly, second processor 1580 includes an MCH 1582 and PP interfaces 1586 and 1588. As shown in FIG. 15 , MCHs 1572 and 1582 couple the processors to their respective memories, i.e., memory 1532 and memory 1534. These memories may be portions of system memory (e.g., DRAM) locally attached to the respective processors. First processor 1570 and second processor 1580 may be coupled to chipset 1590 via PP interfaces 1562 and 1564, respectively. As shown in FIG. 15 , chipset 1590 includes PP interfaces 1594 and 1598.
[0116] Additionally, chipset 1590 includes an interface 1592 that couples chipset 1590 to a high-performance graphics engine 1538 via a PP interconnect 1539. Chipset 1590 may also be coupled to a first bus 1516 via an interface 1596. As shown in FIG. 15 , various input / output (I / O) devices 1514 may be coupled to first bus 1516, along with a bus bridge 1518 that couples first bus 1516 to a second bus 1520. Various devices may be coupled to second bus 1520, which in one embodiment includes, for example, a keyboard / mouse 1522, a communication device 1526, and a data storage unit 1528 such as a disk drive or other mass storage device that may include cord 1530. Additionally, audio I / O 1524 may be coupled to second bus 1520. Embodiments may be incorporated into other types of systems, including mobile devices such as smart cellular phones, tablet computers, netbooks, Ultrabooks, etc.
[0117] Referring to FIG. 16 , a block diagram of a system according to one embodiment is shown. More specifically, FIG. 16 illustrates a portion of a system 1600, which may be any type of computing system ranging from small portable devices such as mobile phones, tablet computers, laptop computers, etc., to larger systems such as desktop systems, server computers, etc. In a typical environment embodiment herein, the computing system may be powered, at least sometimes, by a battery-based power source. As a result, a platform electrical fault may occur. Although the scope of the invention is not limited in this respect, such an electrical fault may include a battery voltage drop below a voltage regulator's undervoltage lockout level. Another example may be a battery short-circuit protection fault. A platform event may occur in an AC power source. For example, an AC adapter also has overcurrent and short-circuit protection mechanisms. For one thing, the AC adapter limits output current and reduces voltage when a certain threshold is reached. When the voltage drops below a certain threshold, the AC adapter turns off. To restore power, the AC adapter is typically unplugged from the wall. Additionally, programmable power supplies typically require the power supply to limit current when the load draws more than a threshold current and turn off when the voltage drops below a threshold voltage. To prevent such electrical faults, a platform agent, such as a power management integrated circuit (PMIC), may issue a notification. This notification of a platform event may be effective as a trigger for a platform event signal in response to detecting a possible adverse platform event, such as an electrical fault. As described herein, this platform event signal may be used to cause throttling within the processor.
[0118] 16, system 1600 includes processor 1610, which may be a multi-core processor or other SoC. In embodiments herein, processor 1610 may be implemented as a processor package and may include one or more semiconductor dies. As shown, processor 1610 includes multiple processing elements 16201-16202. n In some examples, the processing elements 1620 may be homogenous processing elements, such as homogenous processing cores. However, in many embodiments, there may be heterogeneous processing elements 1620, including cores, graphics processors, controllers, special purpose functional units, etc. In some examples, a collection of processing elements 1620 may be grouped into so-called domains, which may operate in common or independent performance states.
[0119] 16, the processor 1610 also includes a throttle circuit 1630. In the illustrated embodiment, the throttle circuit 1630 includes a plurality of individual throttle agents 16301-1630. n As shown, throttle circuit 1630 may be a distributed hardware circuit with each throttle agent 1630 associated with a corresponding processing element 1620. However, it will be understood that the scope of the invention is not limited in this respect and that in other embodiments, a correspondence other than 1:1 is possible. In an embodiment, throttle circuit 1630 and its component throttle agents 16301-1630 n may be implemented as a hardware circuit. Further, in the implementation shown in Figure 16, each throttle circuit 1630 is associated with, but separate from, a corresponding processing element 1620, and in some examples, a given throttle agent may be included within the corresponding processing element 1620.
[0120] Continuing with reference to FIG. 16, a power target agent (PTA) 1640, a power management agent (PMA) 1650, and a usage monitor 1660 are further illustrated. While shown as external components with respect to the processor 1610, in some embodiments, the PTA 1640, the PMA 1650, and the usage monitor 1660 may be implemented at least in part in firmware and / or software executing on one or more processing engines 1620 of the processor 1610. For example, the PTA 1640 may be implemented within an operating system or other system scheduler. Thus, the PTA 1640 receives information regarding the status of power capabilities (e.g., charging capabilities). Additionally, the PMA 1650, in some embodiments, may be firmware or power management microcode and may execute on one or more of the processing elements 1620. As an example, a dedicated core or microcontroller may be used to execute the PMA 1650 to manage the power consumption of the processor 1610. And, a usage monitor 1660 may also be implemented within the operating system or other system scheduler to monitor user activity and provide hints to PMA 1650. Usage monitor 1660 may optimize processor utilization for power (when running on battery) and performance. Note that in one embodiment, usage monitor 1660 may be bundled with power target agent 1640, although these two agents perform different functions.
[0121] 16 , the system 1600 further includes a platform monitor 1670. The platform monitor 1670 may be configured to monitor platform events and assert a throttle signal in response to detecting a platform event that should trigger throttling. As an example, the platform monitor 1670 may include hardware circuitry that compares critical voltage rails to thresholds. When such rail voltages fall below the thresholds, the platform monitor 1670 may be configured to output a throttle signal. Although, of course, different implementations are possible, in one embodiment, the platform monitor 1670 may be implemented as a power management integrated circuit (PMIC), a charge controller, or other platform hardware component.
[0122] During normal operation, the PTA 1640 may receive input power capability information and, based at least in part thereon, determine a package throttle power threshold or target, which is the maximum power consumption level allowed for the processor 1610 during a platform event. It should be noted that this package throttle power target may be less than the processor's given power target or power budget. In some examples, the processor may be configured with multiple power levels. Such power levels may include one or more lower, long-term power budgets or limits (which, on average, the processor will not exceed). The processor may further be configured with an instantaneous power budget that may exceed the higher, long-term power budget. It is understood that in some embodiments, each configured power budget may also have an associated throttle power target.
[0123] From input power capability information or other platform information, such as another indication of high load on the platform device, the PTA 1640 may determine a package throttle power threshold based at least in part on the available charge of the system 1600's battery power source. As an example of another indication, the modem may momentarily throttle its processor and assert a signal when it must operate at a high power spike to transmit to a distant base station. It should be noted that, in an embodiment, the power target agent 1640 monitors the input power status for update events while the platform is powered by an AC power source. Furthermore, based at least in part on the monitored power status, the PTA 1640 may proactively set a package throttle power target. The PTA 1640 may transmit the package throttle power target to a throttle agent in the throttle circuit 1630.
[0124] Additionally, power management agent 1650 may generate conversion information, based at least in part on hint information provided by usage monitor 1660, to enable the throttle agent to convert package throttle power thresholds to processing element throttle power thresholds or targets. As an example, the hint information may correspond to priority hints regarding loads running on different types of processing elements 1620. In one embodiment, the hint information is a numerical value that enables power management agent 1650 to select appropriate parameters for processor 1610. Such hint information may be an abstract concept. As a result, the OS or usage monitor 1660 is not tied to a specific processor (a role performed by power management agent 1650). In one embodiment, usage monitor 1660 may be configured to monitor user actions and provide hint information as a numerator. For example, consider two different types of user activities: file copy and gaming. In this configuration, user monitor 1660 may output hint information having a first hint value (Hint-1) for file copy activity and a second hint value (Hint-2) for gaming activity. Based at least in part on this hint, the PMA 1650 may provide conversion information, for example in the form of a control signal, to the corresponding throttle circuit 1630. In response to this particular hint, the power management agent 1650 may set the conversion ratio of throttle agent 16301 to 10% and the conversion ratio of throttle agent 16302 to 90% if "Hint-1" is received, or set the conversion ratio of throttle agent 16301 to 50% and the conversion ratio of throttle agent 16302 to 50% if "Hint-2" is received. Each throttle agent 1630 also uses this conversion information and package throttle power thresholds to determine appropriate throttle operating points for one or more associated processing elements 1620.
[0125] It should be noted that the above-described operations of setting package throttle power targets and individual processing element throttle power targets may occur proactively during normal operation, for example, according to a given time period or execution loop. Then, during operation, when a given platform event that should trigger processor throttling is detected, a fast, low-latency, hardware-based operation occurs to throttle individual processing elements 1620. That is, when the platform monitor 1670 identifies a change in the platform, such as a switch in operating power from AC power to battery power, it sends a throttle signal to the throttle circuitry 1630 to indicate that a platform event has occurred. Then, in response to this throttle signal, the individual throttle agent 1630 immediately causes the corresponding processing element 16620 to not operate above its corresponding throttle operating point, enabling low-latency throttling. While the embodiment of FIG. 16 is shown at a high level, it is understood that many variations and alternatives are possible.
[0126] Referring to Figure 17, a flow diagram of a method according to one embodiment of the present invention is shown. More specifically, method 1700 is a method of monitoring user activity with a user monitor and controlling platform power consumption with a power management agent according to one embodiment. Thus, method 1700 may be performed by hardware circuitry, software, firmware, and / or combinations thereof. In particular embodiments, the power target agent and usage monitor may be implemented at least in part as software, and in certain embodiments, the agent may be included within an operating system or other management software.
[0127] As shown, method 1700 begins by receiving a notification of an input power capability change (block 1710). The input power capability change can be battery capability or AC power status. The battery capability may change due to the effects of battery charge state, temperature, or aging. The AC power source may be a USB power delivery hub that varies its advertised power depending on the number of ports plugged in. Such an indication of an input power capability change may be a routine notification provided to the platform targeting agent from one or more platform entities, such as a PMIC, a battery fuel gauge, a USB power delivery controller, a battery charging circuit, a voltage regulator, etc. For example, periodic notifications of battery capability (e.g., charge percentage) may be provided. Alternatively, in other examples, the power targeting agent may routinely poll for this information.
[0128] In any event, control passes from block 1710 to block 1720, where a package throttle power target may be set based at least in part on this capacity change. More specifically, a peak power threshold may be set based at least in part on the battery's charge capacity, as described herein. As one representative example, when the platform is battery-powered, the power target agent may set the package throttle power target to a level of 50 watts (W) when the battery has 80% charge and to 30 W when the battery has 20% charge. Notably, while the platform is powered by an AC power source, the power target agent may be configured to set or program the package throttle power target to 50 W when the battery has 80% charge and to program it to 30 W when the battery has 20% charge (per the representative example described above). At block 1730, this package throttle power target may be stored in a configuration memory. For example, this information may be stored in a throttle power threshold field of a configuration register of the processor. Additionally, the power management agent may send this package throttle power target to the throttle agent.
[0129] The power target agent (or another portion of the operating system or other system software) may also perform scheduling activities. Thus, as further shown in FIG. 17 , at block 1740, the power target agent may schedule loads to processing elements. For example, computationally intensive loads may be scheduled to core processing elements, graphics-intensive loads may be scheduled to graphics processing elements, network loads may be scheduled to interface controllers, etc. At block 1750, the usage monitor may provide hint information to the power management agent. More specifically, this hint information may relate to priority information regarding loads being scheduled to different types of processing elements. For example, for computationally intensive loads, core processing elements may have a higher priority, and thus this hint information may indicate the same. Also, for graphics-intensive loads, graphics processing elements may have a higher priority, and thus, in this example, corresponding higher priority hint information for these processing elements may be provided. The scope of the invention is not limited in this respect, and in one embodiment, this hint information may be in the form of a conversion ratio, etc. Although the embodiment of FIG. 17 is shown at a high level, it will be appreciated that many variations and alternatives are possible.
[0130] Referring to Figure 18, a flow diagram of a method according to another embodiment of the present invention is shown. More specifically, as shown in Figure 18, method 1800 is a method of controlling power consumption in a processor by a power management agent as described herein. Accordingly, method 1800 may be performed by hardware circuitry, software, firmware, and / or combinations thereof. In particular embodiments, the power management agent may be implemented at least in part by hardware circuitry and firmware, and in one particular embodiment, the power management agent may be a dedicated processing element, such as a dedicated core or microcontroller, that executes power management code.
[0131] As shown, method 1800 begins by determining an operating point for each processing element based at least in part on a power budget and a load (block 1810). For example, assume that a processor is configured with a given power budget, such as a TDP level or other long-term power level at which the processor may operate (and which may be stored in the processor's configuration memory). The power management agent may also have information regarding the loads to be executed on different processing elements. Thus, in one embodiment, the power management agent may be configured to determine an operating point for each processing element based at least in part on a request from the host software to provide a desired combination of power and performance. The power management agent may then perform an allocation or budgeting of this total power budget to each of the processing elements. For example, when all processing elements execute an equal load, a common operating point may be determined for each of the processing elements. Alternatively, in a more typical scenario, some processing elements may have a higher load and / or consume more power than other processing elements. Thus, asymmetric operating points (e.g., each consisting of an operating voltage and an operating frequency, among other possible parameters) may be determined for each processing element. These operating points may then be transmitted to the individual processing elements in block 1820. In response to this operating point information, the processing elements are operated at such operating points. It will be understood that these operations of blocks 1810 and 1820 may be performed repeatedly during normal operation of the platform.
[0132] 18, at diamond 1830, it is determined whether hint information has been received (e.g., from a usage monitor). If it is determined that such hint information has not been received, control then passes to diamond 1840 to determine whether an update period has expired. This update period is a period during which the power management agent may analyze the power budget and current load to determine whether the operating point should be updated. Once this update period has expired, control returns to block 1810, discussed above.
[0133] 18 , if instead it is determined that updated hint information has been received, control passes to block 1850. At block 1850, a translation may be determined. More specifically, this translation is some form of operation that takes the package throttle power target and allocates it appropriately among different processing elements. In an embodiment, the power management agent may perform this translation based at least in part on the hint information received from the usage monitor. In this manner, a more fair allocation of available power amounts to different processing elements based on their priority, criticality, user utility, etc. may occur.
[0134] In one embodiment, the conversion of throttle power from package to element may use a scaling factor. In one such example, assume that a processor includes two processing elements (a core and a graphics processor) and that the package throttle power threshold is 50 W. Further, for illustrative purposes, when a load is core-intensive (e.g., host software copying a file), the power management agent generates conversion information to set a ratio of 90% for the core and 10% for the graphics processor. With this ratio or scaling factor configuration, the resulting amount of power drawn by the processing element corresponds to a level of 45 W for the core and 5 W for the graphics processor. Alternatively, assume that a load (e.g., host software rendering an image) causes the power management agent to set a ratio of 20% for the core and 80% for the graphics processor. With this ratio or scaling factor configuration, the resulting amount of power drawn by the processing element corresponds to a level of 10 W for the core and 40 W for the graphics processor.
[0135] It should be noted that embodiments may use different techniques to enable the power management agent to provide conversion information to define the budgets dynamically calculated by each throttle agent. Accordingly, different methods of performing this conversion are possible, but in some examples, allocation may occur according to ratio information associated with each of the processing elements. Alternatively, pointer information may be provided based on the conversion operation performed by the power management agent. The pointer information may also be used to enable access to a table containing operating point information. In one example, the conversion information includes ratio information, and setting the ratio value (e.g., a coefficient) for each processing element may be based at least in part on hint information received from a usage monitor. Finally, with further reference to FIG. 18 , control may pass to block 1860, where the conversion information may be sent to the throttle agent. While the embodiment of FIG. 18 is shown at a high level, many variations and alternatives are possible.
[0136] Referring to FIG. 19 , a flow diagram of a method according to yet another embodiment of the present invention is shown. More specifically, method 1900 is a method for setting processing element power throttle limits using a throttle agent, according to one embodiment. Accordingly, method 1900 may be performed by hardware circuits, software, firmware, and / or combinations thereof. In particular embodiments, one or more hardware circuits may be provided to perform method 1900 (e.g., a distributed throttle circuit having multiple individual hardware circuits, each associated with at least one processing element). It should be noted that method 1900 may be used to set individual platform element throttle power targets, and may be performed according to a so-called slow loop, as this target setting may be performed in a non-deterministic manner, without affecting the low-latency activation of the actual throttle operation.
[0137] As shown in Figure 19, method 1900 begins by receiving a package throttle power target from a power target agent and conversion information from the power target agent (block 1910). Although a single operation is shown in Figure 19, it is understood that these two different values may be received asynchronously (from two different entities) and according to a given time period. In any event, control passes to block 1920 where a processing element throttle power target may be determined based at least in part on these two values. In one example, the throttle agent receives a processing element throttle power target (P element_throttle ) to the package throttle power target (P package_throttle ) and the count (c power_management ) is calculated, for example, according to the following formula: element_throttle =f(P package_throttle ,c power_management ). In some embodiments, there may be no specific timing requirements for the throttle agent to translate throttle power thresholds from package to element. This is because the power target agent may set the throttle power only in response to specific platform events (e.g., a change in battery status), which typically has a latency on the order of a few seconds. Furthermore, note that the power management agent may be optimized for the type of host software activity, which typically does not have hard timing requirements.
[0138] In another embodiment, the throttle agent may be configured to maintain a lookup table (which may be provided by platform firmware, e.g., the basic input / output system (BIOS)). In this embodiment, the power management agent provides a pointer to the lookup table for a given set of package throttle power thresholds.
[0139] 19, control then passes to block 1930, where a throttle operating point for a processing element may be determined based on the throttle power target for the processing element. As an example, the throttle agent may access a local lookup table based on the throttle power target for the processing element to identify a corresponding operating point (e.g., a voltage / frequency pair). Then, at block 1940, the throttle operating point may be stored in a given memory location, for example, internal to the throttle agent. As an example, the throttle agent may include configuration memory for storing the determined throttle operating point. The determined throttle operating point may be retrieved and used to control throttle activity when a throttle instruction is received, as described below.
[0140] It should be noted that, according to embodiments herein, the resulting operating points of different processing elements may be asymmetric, based at least in part on the priority of the processing elements for a given load. It is further understood that at least one processing element (and possibly multiple processing elements) may be operated at a throttle operating point higher than a minimum operating point. That is, based on battery capacity (e.g., charge level), the throttle operating point may be higher than a standard throttle operating point (i.e., a minimum operating point that may be stored in the processor's configuration memory). As an example, this minimum operating point may correspond to a low-frequency mode, which is the lowest frequency and voltage at which the processor may operate. While the embodiment of FIG. 19 is shown at a high level, it is understood that many variations and alternatives are possible.
[0141] In an embodiment, the throttle agent may use the calculated processing element throttle power target to operate one or more associated processing elements at a given operating point. For a core or graphics processor, the operating point may be converted to an operating voltage / frequency pair. For other processing elements, such as input / output (IO) ports, the operating point may be converted to other operating parameters, such as dynamic link width, bandwidth m, etc. Notably, the throttle agent may be configured to perform throttling to immediately avoid platform electrical failures. In some examples, if a processing element does not immediately convert to the indicated throttle operating point (e.g., due to hardware limitations), it may first transition to a safe state (e.g., an intermediate operating point lower than the current operating point but higher than the throttle operating point) before locking into the optimal throttle state.
[0142] Referring to FIG. 20 , a flow diagram of a method according to yet another embodiment of the present invention is shown. More specifically, method 2000 of FIG. 20 is a method of throttling processing element power using a throttling agent according to one embodiment. Accordingly, method 2000 may be performed by hardware circuitry, software, firmware, and / or a combination thereof. Here, according to an embodiment, method 2000 may be performed by a given hardware circuit, i.e., a given throttling agent. By using hardware circuitry (and possibly dedicated hardware circuitry associated with one or more processing elements), throttling operations can occur with very low latency. Thus, method 2000 may be performed according to a fast loop to cause low-latency throttling. For example, in one embodiment, the throttling agent may be configured to cause throttling within a sub-microsecond latency window from the indication of a platform event. It should be noted that the use of distributed hardware circuitry in embodiments avoids the overhead of the OS or other system software (and / or power control circuitry / firmware, such as a power management agent) determining appropriate throttle points and controlling the processing elements.
[0143] As shown in FIG. 20 , method 2000 begins by reading a throttle operating point (block 2010). As described above, this throttle operating point may be stored in the configuration memory of the throttle agent. Next, at diamond 2020, it is determined whether a throttle signal has been received, for example, from a platform monitor. As described herein, this throttle signal is an indication to the throttle agent that the processor is now to be controlled to operate at a power level not higher than the package throttle power threshold. Thus, (for example) removal of AC power results in an indication of a platform event (e.g., a throttle signal). In response to this indication, the processor may be configured to throttle overall power consumption to a preprogrammed power target (e.g., 50 W at 80% charge, 30 W at 20%, as set by the power target agent).
[0144] If it is determined that a throttle signal has been received, control passes to block 2030, where the processing element is operated at a throttle operating point. For example, in an example where the processing element includes an internal clock generator and associated voltage regulator, the throttle agent sends an operating point command to update one or more frequencies and operating voltages internally to the processing element. In other examples, the throttle agent may communicate with other control components, such as associated clock generators, voltage regulators, link state machines, etc., to control appropriate operating parameters. As a result, the processing element operates at a power level no higher than the processing element throttle power threshold. While the embodiment of FIG. 20 is shown at a high level, it will be understood that many variations and alternatives are possible.
[0145] The following examples relate to further embodiments.
[0146] In one example, a processor includes a plurality of processing elements that perform operations, a PMA coupled to the plurality of processing elements and that controls power consumption of the plurality of processing elements, and a throttle circuit coupled to the PMA, the throttle circuit may include a plurality of throttle agents, each associated with one of the plurality of processing elements, the PMA communicating conversion information to the throttle circuit, and each of the plurality of throttle agents determining a throttle power level for an associated one of the plurality of processing elements based at least in part on the conversion information.
[0147] In one example, the conversion information includes ratio information, and the PMA determines the conversion information based at least in part on hint information received from a usage monitor, the hint information indicating relative priorities of the multiple processing elements.
[0148] In one example, the conversion information includes pointer information, and a first throttle agent of the plurality of throttle agents accesses a lookup table using a first pointer in the pointer information to determine a first throttle power level for a first processing element.
[0149] In one example, the first throttle agent determines a first operating point for the first processing element based on the first throttle power level and sends an operating point update to the first processing element to cause the first processing element to operate at the first operating point.
[0150] In one example, the first operating point is greater than the minimum operating point.
[0151] In one example, the power target agent sets a package throttle power level for the processor based at least in part on changes in a battery's charge capacity.
[0152] In one example, a platform monitor communicates a throttle signal to the throttle circuit in response to a platform event, the platform event including a switch of the platform to battery operation, and the power target agent further communicates the package throttle power level to the throttle circuit.
[0153] In one example, the package throttle power level is less than a thermal design power of the processor, and the throttle circuit causes at least some of the processing elements to operate at an operating point greater than a minimum operating point.
[0154] In another example, the method comprises: determining, in a power controller of a system-on-chip (SoC) including a plurality of processing circuits, a conversion information for a throttle power threshold of the SoC based at least in part on the hint information in response to the platform event; and transmitting the conversion information to a plurality of throttle agents of the SoC to cause the plurality of throttle agents to update an operating point of at least one of the plurality of processing circuits to a throttle level greater than a minimum operating point based at least in part on the conversion information, each of the plurality of throttle agents being associated with at least one of the plurality of processing circuits.
[0155] In one example, the method further includes receiving the hint information from a usage monitor, the hint information including priority information for the plurality of processing circuits.
[0156] In one example, the method comprises: receiving conversion information at the plurality of throttle circuits; causing a first processing circuit of the plurality of processing circuits to operate at a first operating point based on a first conversion factor of the conversion information by a first throttle agent of the plurality of throttle agents; The method further includes a step of causing a second processing circuit of the plurality of processing circuits to operate at a second operating point based on a second conversion element of the conversion information by a second throttle agent of the plurality of throttle agents, the second operating point being greater than the first operating point, and the second processing circuit having a higher priority than the first processing circuit according to the hint information.
[0157] In one example, determining the transformation information includes generating a plurality of coefficients based at least in part on the hint information, each coefficient being associated with one of the plurality of processing circuits.
[0158] In one example, the method comprises: calculating, at a first throttle agent of the plurality of throttle agents, a first throttle power level for a first processing circuit of the plurality of processing circuits based on a first coefficient of the plurality of coefficients and a throttle power limit of the SoC; The method further includes causing the first processing circuit to operate at a first operating point based on the first throttle power level, the first operating point being greater than the minimum operating point.
[0159] In one example, the method comprises: causing the first processing circuit to operate at an intermediate operating point, the intermediate operating point being lower than a previous operating point at which the first processing circuit was operating and higher than the first operating point; causing the first processing circuit to operate at the first operating point after a hardware limit of the first processing circuit is reached.
[0160] In one example, determining the transformation information includes generating a plurality of pointers based at least in part on the hint information, each pointer associated with one of the plurality of processing circuits.
[0161] In one example, the method comprises: accessing, by a first throttle agent of the plurality of throttle agents, a table using a first pointer of the plurality of pointers to obtain a first throttle power level for a first processing circuit of the plurality of processing circuits; The method further includes causing the first processing circuit to operate at a first operating point based on the first throttle power level, the first operating point being greater than the minimum operating point.
[0162] In another example, a computer-readable medium containing instructions for performing the method of any of the above examples.
[0163] In another example, a computer-readable medium containing data is for use by at least one machine to manufacture at least one integrated circuit to perform the method of any of the above examples.
[0164] In another example, an apparatus includes means for performing the method of any of the above examples.
[0165] In yet another example, the system a first power source for providing power to the system; a second power source for powering the system, the second power source including a battery; a charging circuit for charging the second power source; a power management integrated circuit coupled to the processor, the power management integrated circuit sending a throttle signal to the processor in response to a switch of power from the first power source to the second power source, and at least one of the second power source and the charging circuit communicating a charging capability of the battery to the processor; the processor.
[0166] In one example, the processor: at least one core that executes a first instruction; at least one graphics processor that executes second instructions; a power control unit coupled to the at least one core and the at least one graphics processor, the power control unit controlling power consumption of the at least one core and the at least one graphics processor according to a power budget of the processor; and a power target agent preemptively determining a throttle power budget for the processor based at least in part on the charging capability. The power controller preemptively determines conversion information based on priorities of the at least one core and the at least one graphics processor. The processor may further include a throttle circuit coupled to the power controller preemptively determining a first throttle power budget for the at least one core and a second throttle power budget for the at least one graphics processor based at least in part on the conversion information and the throttle power budget, and operating the at least one core at a first operating point based on the first throttle power budget and the at least one graphics processor at a second operating point based on the second throttle power budget in response to the throttle signal.
[0167] In one example, the conversion information includes ratio information, and the power control unit determines the conversion information based at least in part on hint information received from a usage monitor, the hint information indicating relative priorities of the at least one core and the at least one graphics processor.
[0168] In one example, the conversion information includes pointer information, and the throttle circuit accesses a lookup table using a first pointer in the pointer information to determine the first throttle power amount for the at least one core, and accesses the lookup table using a second pointer in the pointer information to determine the second throttle power amount for the at least one graphics processor.
[0169] In one example, the power control unit causes the at least one core to operate at an intermediate operating point, the intermediate operating point being lower than a previous operating point at which the at least one core was operating and higher than the first operating point, the first operating point being higher than a minimum operating point; The at least one core is caused to operate at the first operating point after a hardware limit of the at least one core is reached.
[0170] It will be appreciated that various combinations of the above examples are possible.
[0171] It should be noted that the terms “circuit” and “circuitry” are used interchangeably herein. As used herein, these terms, and the term “logic,” alone or in any combination, are used to refer to analog circuits, digital circuits, hard-wired circuits, programmable circuits, processor circuits, microcontroller circuits, hardware logic circuits, state machine circuits, and / or any other type of physical hardware component. Embodiments may be used in many different types of systems. For example, in one embodiment, a communications device may be configured to perform the various methods and techniques described herein. Of course, the scope of the invention is not limited to communications devices; rather, other embodiments may be directed to other types of apparatus for processing instructions or one or more machine-readable media containing instructions that, when executed by a computing device, cause the device to perform one or more of the methods and techniques described herein.
[0172] Embodiments may be implemented in code and stored on a non-transitory storage medium that stores instructions that can be used to program a system to execute the instructions. Embodiments may be implemented in data and stored on a non-transitory storage medium that, when used by at least one machine, causes the at least one machine to produce at least one circuit for performing one or more operations. Yet further embodiments may be implemented in a computer-readable storage medium that includes information that, when manufactured into an SoC or other processor, configures the SoC or other processor to perform one or more operations. The storage medium may include any type of disk, including, but not limited to, floppy disks, optical disks, solid-state drives (SSDs), compact disk-read-only memories (CD-ROMs), compact disk-rewriteables (CD-RWs), and magneto-optical disks; read-only memory (ROM); random access memory (RAM); static RAM (SRAM); semiconductor devices such as erasable programmable ROMs (EPROMs), flash memory, electrically erasable programmable ROMs (EEPROMs); magnetic or optical cards; or any other type of medium suitable for storing electronic instructions.
[0173] While the present invention has been described with respect to a limited number of embodiments, those skilled in the art will appreciate numerous modifications and variations therefrom, and it is intended that the appended claims cover all such modifications and variations that fall within the spirit and scope of the invention. [Explanation of symbols]
[0174] 1630 Throttle Agent 1640 Power Target Agent 1650 Power Management Agent 1660 Monitor in use 1670 Platform Monitor
Claims
1. A device, a plurality of processing elements including a first processing core, a second processing core, and a graphics processor, the first processing core including a lower power processing core than the second processing core; a power management agent (PMA) coupled to the plurality of processing elements, the PMA controlling power consumption of the plurality of processing elements; a throttle circuit coupled to the PMA, the throttle circuit operating the processing elements at corresponding voltage / frequency pairs based on conversion information provided by the PMA, the conversion information defining a power of each individual processing element of the plurality of processing elements; wherein the PMA provides different translation information for loads having different attributes, the different attributes including different priorities associated with different loads.
2. The apparatus of claim 1 , wherein the PMA controls power consumption of the plurality of processing elements based on a power target.
3. 3. The apparatus of claim 1, wherein the different conversion information further causes the throttle circuit to configure different power levels for the graphics processor.
4. An apparatus described in any one of claims 1 to 3, wherein the different conversion information includes first conversion information based on a first priority and second conversion information based on a second priority.
5. 5. The device of claim 4, wherein the first conversion information includes information indicating first operating points of the first processing core, the second processing core, and the graphics processor, and the second conversion information includes information indicating second operating points of the first processing core, the second processing core, and the graphics processor.
6. The device of claim 5 , wherein the first operating point and the second operating point comprise a first set of frequencies and a second set of frequencies for the first processing core, the second processing core, and the graphics processor.
7. The apparatus of any one of claims 1 to 6, wherein the different conversion information comprises different corresponding ratio information.
8. The apparatus of any one of claims 1 to 7, wherein the PMA determines the different transformation information based at least in part on hint information.
9. The device of claim 8 , wherein the hint information is based on a first characteristic of a first load attribute and a second characteristic of a second load attribute.
10. 1. A method comprising: executing a load on a plurality of processing elements including a first processing core, a second processing core, and a graphics processor, wherein the first processing core comprises a lower power processing core than the second processing core; controlling power consumption of the plurality of processing elements by a power management agent (PMA); operating, by a throttle circuit, the processing elements at corresponding voltage / frequency pairs based on conversion information provided by the PMA, the conversion information defining the power of individual processing elements of the plurality of processing elements; providing, by the PMA, different translation information for loads having different attributes, the different attributes including different priorities associated with different loads; A method comprising:
11. The method of claim 10 , wherein the PMA controls power consumption of the plurality of processing elements based on a power target.
12. 12. The method of claim 10 or 11, wherein the different conversion information further causes the throttle circuit to configure different power levels for the graphics processor.
13. A method according to any one of claims 10 to 12, wherein the different conversion information includes first conversion information based on a first priority and second conversion information based on a second priority.
14. 14. The method of claim 13, wherein the first conversion information includes information indicating first operating points of the first processing core, the second processing core, and the graphics processor, and the second conversion information includes information indicating second operating points of the first processing core, the second processing core, and the graphics processor.
15. 15. The method of claim 14, wherein the first operating point and the second operating point comprise a first set of frequencies and a second set of frequencies for the first processing core, the second processing core, and the graphics processor.
16. The method according to any one of claims 10 to 15, wherein the different transformation information comprises different corresponding ratio information.
17. The method according to any of claims 10 to 16, wherein the PMA determines the different transformation information based at least in part on hint information.
18. The method of claim 17 , wherein the hint information is based on a first characteristic of a first load attribute and a second characteristic of a second load attribute.
Citation Information
Patent Citations
Multi-core processor system, control method of multi-core processor system and control program of multi-core processor system
JP2014209394A
Systems and methods for thermal management in a portable computing device that uses thermal resistance values to predict optimal power levels
JP2016510917A
Multiprocessor system, multiprocessor control method, and multiprocessor integrated circuit
WO2010137262A1
Scheduling method, design support method, and system
WO2012108058A1