Energy-Efficient Core Voltage Selection Device and Method
A core ranking scheme in SoC environments addresses Vmin variations across processor cores, reducing power consumption by up to 8% through selective core usage based on energy efficiency, enhancing battery life and responsiveness.
Patent Information
- Application Number
- JP2021120956
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-22
- Filing Date
- 2021-07-21
- Publication Date
- 2025-08-04
- Estimated Expiration
- 2041-07-21
AI Technical Summary
In system-on-chip (SoC) environments, the minimum applicable voltage (Vmin) varies across processor cores due to silicon attributes like inverse thermal dependence, leading to inefficient power consumption and potential transistor errors, with existing technologies lacking per-core voltage regulation, resulting in higher power consumption than necessary.
Implementing a core ranking scheme based on processor core energy efficiency, storing Vmin values in fuses or non-volatile memory during high volume manufacturing, and utilizing an operating system scheduler to select the core with the lowest Vmin for optimal battery life and responsiveness.
Achieves up to an 8% reduction in SoC power consumption for low-power, low-frequency workloads by selectively using cores with the lowest Vmin, maintaining performance and compatibility without software changes.
Smart Images

Figure 0007717514000003 
Figure 0007717514000004 
Figure 0007717514000005
Abstract
Description
Technical Field
[0001] This application claims the priority of U.S. Provisional Application No. 63 / 069,622, filed on August 24, 2020, entitled "ENERGY-EFFICIENT CORE VOLTAGE SELECTION APPARATUS AND METHOD", and incorporates it by reference in its entirety.
Background Art
[0002] In a power-on voltage-frequency controlled domain within a system-on-chip (SoC), such as a central processing unit (CPU), a graphics processing unit (GPU), and other IP (intellectual property) blocks, the minimum applicable voltage of the power supply varies at each operating frequency. This voltage is herein referred to as V min . If it is lower than such a V min value, transistor digital logic is prone to errors, and thus the circuit is considered not to function. Such a V min is determined in the manufacturing test process by high volume manufacturing (HVM). The V min voltage may be affected by other silicon attributes such as, for example, inverse thermal dependence (ITD). Within the same wafer, the V min value may vary from die to die, or may vary from processor core to core within a die.
Brief Description of the Drawings
[0003] The disclosed embodiments will be more fully understood from the following detailed description given herein and from the accompanying drawings of various embodiments of the disclosure, which should not be construed as limiting the disclosure to the specific embodiments, but are for explanation and understanding only.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Best Mode for Carrying Out the Invention
[0004] In a multiprocessor (MP) system, currently, multiple active processor cores cumulatively, for each of the processor cores, their respective per-core Vmin It consumes more power than is necessary to run at the voltage. This is due to the following reasons. First, most platforms do not have a voltage regulator (VR) control domain per processor core. For example, VR output and thus control can be aggregated commonly across two or more processor cores. Second, in a PCPS (Per-Core Performance State) platform, different operating frequencies are allowed among cores. However, the higher voltage (the maximum among cores) is realized as the V of the SoC at a given frequency. Third, in a platform that supports a per-core voltage domain, per-core V as an HVM feature may not be available. min in the maximum among) is realized as the V of the SoC at a given frequency. min as an HVM feature may not be available. min is not available.
[0005] Various embodiments utilize the V difference across processor cores to ultimately achieve optimal battery life and responsiveness. Based on statistical data, it has been observed that varying V for low-power, low-frequency workloads can be done, for example, by approximately 30 mV. Regarding the impact of the actual battery life on the workload, this can result in a savings of up to approximately 8% of the SoC power. Some embodiments use a core ranking scheme based on processor core energy efficiency. This is similar to the favored core approach in a multi-core processor system. However, here the favored core is a core with good energy efficiency that enables the SoC to use the core with the lowest V for energy efficiency. What is obtained is specific to the low-power consumption scenario. Such V values can be stored in appropriate registers during the HVM process. For example, the V value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. min difference across processor cores to ultimately achieve optimal battery life and responsiveness. Based on statistical data, it has been observed that varying V for low-power, low-frequency workloads can be done, for example, by approximately 30 mV. Regarding the impact of the actual battery life on the workload, this can result in a savings of up to approximately 8% of the SoC power. Some embodiments use a core ranking scheme based on processor core energy efficiency. This is similar to the favored core approach in a multi-core processor system. However, here the favored core is a core with good energy efficiency that enables the SoC to use the core with the lowest V for energy efficiency. What is obtained is specific to the low-power consumption scenario. Such V values can be stored in appropriate registers during the HVM process. For example, the V value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. min difference across processor cores to ultimately achieve optimal battery life and responsiveness. Based on statistical data, it has been observed that varying V for low-power, low-frequency workloads can be done, for example, by approximately 30 mV. Regarding the impact of the actual battery life on the workload, this can result in a savings of up to approximately 8% of the SoC power. Some embodiments use a core ranking scheme based on processor core energy efficiency. This is similar to the favored core approach in a multi-core processor system. However, here the favored core is a core with good energy efficiency that enables the SoC to use the core with the lowest V for energy efficiency. What is obtained is specific to the low-power consumption scenario. Such V values can be stored in appropriate registers during the HVM process. For example, the V value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. min for energy efficiency. What is obtained is specific to the low-power consumption scenario. Such V values can be stored in appropriate registers during the HVM process. For example, the V value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. min values can be stored in appropriate registers during the HVM process. For example, the V value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. min value can be stored in fuses, non-volatile memory (NVM), programmable registers, etc. In some embodiments, the operating system (OS) scheduler uses the core ranking information to select the core with the lowest V. minScheduling a specific application to a core having can achieve optimal energy performance.
[0006] Some embodiments capture the difference in V min values of fuses (or any suitable non-volatile memory) that store per-core V min variations between cores resulting from HVM process-logic changes to provide a mechanism for reducing the average power of the SoC. Based on this new aspect of silicon, lower power can be achieved in the SoC when operating at or near the low-frequency mode (LFM) and the efficient frequency (P e ). The efficient frequencies in the LFM range are, for example, from 400 MHz to 1 GHz. Telemetry data indicates that it is often in a lower frequency range, which is relevant to real-world usage scenarios. In some embodiments, the SoC displays energy efficiency information in a lower or other specific frequency range to the OS as applicable per core. In some embodiments, the OS enables an energy-aware scheduling scheme that incorporates considerations (e.g., efficient cores for handling low-utilization interrupts and deferred procedure call (DPC) services). min
[0007] Processor DPC time is the time the processor spends receiving and servicing DPCs. DPCs are interrupts that are executed with a lower priority than standard interrupts. Percent DPC time is a component of percent privileged time because DPCs are executed in privileged mode. If a high percent DPC time persists, there may be a processor bottleneck or application- or hardware-related problems that can significantly degrade the performance of the entire system.
[0008] There are numerous technical effects in various embodiments. For example, the energy-efficient core voltage selection scheme does not involve a compromise or trade-off in performance in the process of reducing the average power. In contrast, selecting the optimal V min helps to improve performance from a higher turbo residency in a specific scenario. The schemes of various embodiments utilize the goodness of silicon characteristics in the HVM process, ultimately bringing a competitive advantage that can even be uniquely developed without mandating software changes. This scheme utilizes the HVM process with well-established means to identify such V min differences per core. This scheme maintains compatibility with the stress and reliability requirements of the SoC. Other technical effects will become apparent from the various figures and embodiments.
[0009] In the following description, numerous details are set forth in order to provide a more thorough explanation of embodiments of the present disclosure. However, it will be apparent to those skilled in the art that embodiments of the present disclosure can be practiced without these specific details. Also, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the present disclosure.
[0010] Note that in the corresponding drawings of the embodiments, signals are represented by lines. Some lines may be made thicker to indicate a signal path that is more of a component, and / or may have arrows at one or both ends to indicate the main flow direction of information. Such indications are not intended to be limiting. Rather, they are used to assist in the easy understanding of circuits or logic units in relation to one or more exemplary embodiments. The signals represented can be determined by the design needs or preferences and can actually have one or more signals that can travel in any direction and can be implemented in any suitable type of signal scheme.
[0011] FIG. 1 shows V for various frequency bins and for the number of processor cores in a system-on-chip (SoC).min A set of plots 100 indicating ranges is illustrated. An example of the SoC is described with reference to FIG. 9. Referring back to FIG. 1, various embodiments combine hardware features such as V per core with software scheduling optimization to utilize newer power saving sources. min This describes the feature of utilizing newer power saving sources by combining hardware features such as V. A set of plots 100 illustrates plots 101 showing the range of deviation of V in an SoC determined using mass production for various frequency bins. min The plot of plot 102 illustrates statistics regarding the achievable V margin in significant units. The margin bar of plot 102 shows the possibility of achieving a reduction of up to 30 mV in the floor V at 800 MHz. This margin gradually decreases as it moves away from the low frequency mode (LFM) range (1.8 GHz). A similar trend is observed in the turbo region. min min
[0012] FIG. 2 illustrates a histogram 200 showing the average frequency distribution. Chip vendors and operating system developers (e.g., Intel and Microsoft respectively) collect telemetry data from actual deployments and absorb that knowledge. In one example, the data source is the Microsoft Inbox Data Collector on Windows® and Microsoft® Surface devices. In one case, data collected based on a sample space of over 700 real - world devices indicates that, for cumulative background activity monitor (BAM) - based workloads exceeding 10%, the median usage rate is less than 70%. Insights regarding low - power background activity are significantly higher than assumptions based on laboratory data. The Windows operating system (OS) indicates a platform that tracks and utilizes background activity at runtime, reducing the quality of service (QoS) level. Currently, this also serves as the runtime upper limit for p - states (p - state) of the processor, for example, within 1.5 GHz.
[0013] Given these telemetry indicators, some embodiments use per-core V for p-states below 1.5 GHz to increase energy efficiency. min While various embodiments are extensible up to the turbo region, the embodiments described herein and their results focus on battery life near the LFM boundary that affects the workload.
[0014] The instantaneous dynamic power of a processor is proportional to the square (approximate) of its instantaneous frequency and the corresponding voltage (V cc ). In a multi-core environment, currently all cores use a common V at a given p-state. min However, by using per-core V min , it is possible to bring about, for example, a lowest V min that is reduced by up to 30 mV in the LFM region. Using per-core V min can achieve a lower V min compared to using a single V min for all cores, resulting in a reduction of up to, for example, 8% for background activity.
[0015] With the support of HVM statistics, the following experiment converts this into an actual benefit in SoC power. Here, the voltage-frequency (VF) curve is updated to produce a delta indicating the benefit of per-core V min . In one example, the benefit of VF is limited to a conservative 1.3 GHz. Ideally, the updated VF should be lowered by only the observed margin. However, to avoid system instability and crashes, the delta is reversed without much concern for relative percent benefit analysis. Also, the simulated setup applies the new voltage to the cores. This is the ideal benefit that can be achieved when software-based scheduling is tied to select the best (more energy-efficient) cores presented to the OS by the hardware capability ranking of the cores.
[0016] In some embodiments, one or more fuses or non-volatile memories are updated or programmed as part of a high volume manufacturing (HVM) process and associated tool usage changes, per-core V min Thereby, new per-core V min becomes eligible. In this HVM setup, some low power and / or low QoS workloads are executed on the cores, and per-core V min is recorded. In some embodiments, instead of or in addition to reading from the fuses or NVM, per-core V min is transmitted to the firmware over some suitable communication medium (e.g., the Internet). Examples of NVM include ferroelectric memory, phase change memory, magnetic memory, resistive change memory, flash memory, and the like.
[0017] The results in Table 1 are based on the geometric mean scores of three runs on the computer platform.
Table 1
[0018] As expected, the benefits (percent savings) improve against the background activity. The above results are without software scheduling changes and assume a common VF curve across the cores. Software changes are considered in the alternative solution options described herein.
[0019] FIG. 3 illustrates an array 300 of per-core V min for each frequency for various p-states. FIG. 3 shows the pattern-encoded values interpreted for each row of p-states. The P-state is a state regarding the optimization of the operating voltage and CPU frequency. During code execution, the operating system and CPU can optimize power consumption through multiple different p-states (performance states). Depending on the requirements, the CPU operates at different frequencies. P0 is the highest frequency (often associated with the highest voltage).
[0020] HVM programs the VF fuse or non-volatile memory with a V min value. For example, the pattern-encoded value can be a five-point tuple of {P state, V min}. Typically, the manufacturing test flow determines the best V min for each frequency for a given component (e.g., an SoC with multiple cores). When V min per core is enabled in HVM, the relative V min delta fuse (or NVM) is also programmed for each core's V min .
[0021] Figure 4 illustrates a plot 400 showing that the V min difference is clear after applying the inverse thermal dependence (ITD). For example, additional buffer values (guard bands) can be applied to V min for various reasons such as low temperature effect compensation (LTEC) and reliability stress requirement (RSR) during runtime. ITD explains the relationship between the die temperature and the expected V cc value. However, since the V min delta per core (per p state) can still hold after ITD consideration, the V min after ITD application is not restrictive to the various embodiments herein. Also, unlike the V min variation for a given p state, the ITD variation between cores is small (e.g., a coefficient of less than 2% of the variance in the region of interest).
[0022] Figure 5 illustrates a ranking scheme 500 that converts the per-core V min array of FIG. 3 into a ranking array per p state according to some embodiments. In some embodiments, different V minAdditional routines are added during the HVM process to compare and update appropriate registers with this relative rank information. For example, a sequential rank set is assigned to a homogeneous core set under an independent voltage domain. Depending on the processor topology, further multiple or hierarchy-independent rankings can be created (e.g., different rank sets of processor cores for multi-processors and Atom processors, etc.).
[0023] For software scheduling to utilize energy efficient (EE) cores, the software scheduler needs to know which core to select from the relative EE rankings among cores in a given p-state. In one example, Intel architecture ISA extensions programming can be used for software scheduling. One such abstraction is described in the hardware (HW) feedback structure of Table 2.
Table 2
[0024] The energy efficiency capability field can be used to check the V for each core. Specifically, the structure for each core is initialized with performance capability and EE capability. The performance capability is an indicator of the performance core ranking (near turbo), and the EE capability is listed for the fixed frequency closer to LFM. The EE capability frequency is a direct function of the efficient frequency for the low power scenario also known as P min For maintaining compatibility with existing definitions in the public programming reference document, according to some embodiments, the ranking scheme (fused or stored in NVM) of the HVM definition can be read and transferred to the existing HW feedback register. e Figure 5 summarizes this ranking scheme, showing the V for each p-state for each core
[0025] Figure 5 summarizes this ranking scheme, showing the V for each p-state for each core minThe pattern code of Figure 3 in the table is mirrored in the ranking table for each core and each p-state. This ranking function is intended to be platform-specific. This is due to platform-specific attributes such as, for example, the instructions per cycle (IPC) gain per cycle of different processor cores. The ranking can be determined using a function. The arguments of one such generated ranking function are shown below: F(Vmin_Largest, Vmin_smallest, Vmin_thisCore, p-state, relative-IPC).
[0026] Here, p-state and relative-IPC are Hardware Guided Scheduling (HGS) arguments, and Vmin_Largest, Vmin_smallest, and Vmin_thisCore are additional arguments added to determine the ranking for each core. The larger the number, the higher the ranking, meaning a better EE core. The arguments here refer to the dependence on relative IPC and p-state in the existing / initial version of HGS.
[0027] Figure 6 illustrates a processor system 600 having an apparatus or mechanism for energy-efficient core voltage selection, according to some embodiments. The processor system 600 includes a processor 601 coupled to an operating system (OS) 602. The processor 601 has one or more processors 603 (individually labeled as processors 603_10 through 603_1N and 603_20 through 603_2N, where "N" is a number), a fabric 604 connecting to the processors 603, a memory 605, a memory controller 610, and a memory physical layer (MEM PHY) 611. In some embodiments, each processor 603 is a die, dielet, or chiplet. Here, the term "die" generally refers to a continuous single piece of semiconductor material (e.g., silicon) in which transistors or other components that make up a processor core may be present. A multi-core processor may have two or more processors on a single die, but alternatively, two or more processors may be provided on two or more respective dies. Each die may have a dedicated power controller or power control unit (p-unit), or a power controller or power control unit (p-unit) that can be configured dynamically or statically as a supervisor or supervisee. In some examples, the dies are of the same size and function, i.e., symmetric cores. However, the dies can also be asymmetric. For example, some dies have different sizes and / or functions from other dies. Each processor 603 may also be a dielet or chiplet. Here, the term "dielet" or "chiplet" generally refers to a physically separate semiconductor die, and typically, a fabric crossing the die boundary functions like a single fabric rather than two separate fabrics to enable connection to adjacent dies. Thus, at least some of the dies may be dielets. Each dielet can include one or more p-units that can be configured dynamically or statically as a supervisor, a supervisee, or both.
[0028] In some embodiments, fabric 604 is a set of interconnects or a single interconnect that enables various dies to communicate with each other. Here, the term "fabric" generally refers to a communication mechanism having a known set of source, destination, routing rules, topology, and other properties. The source and destination can be any type of data handling functional unit, such as a power management unit for example. The fabric can be two-dimensional (2D) spreading along the x-y plane of the die and / or three-dimensional (3D) spreading along the x-y-z plane of a stack of dies arranged vertically and horizontally. A single fabric may span multiple dies. The fabric can take on any topology, such as a mesh topology, a star topology, a daisy chain topology, etc. The fabric may be part of a network-on-chip (NoC) having multiple agents. Those agents can be any functional unit.
[0029] In some embodiments, each processor 603 may include a plurality of processor cores. One such example is illustrated with reference to processor 603_10. In this example, processor 603_10 includes a plurality of processor cores (Core) 606-1 through 606-M, where M is a number. For simplicity, the processor cores are generally referred to by the label 606. Here, the term "processor core" generally refers to an independent execution unit that can execute one program thread at a time in parallel with other cores. A processor core may include a dedicated power controller or power control unit (p-unit) that can be configured dynamically or statically as a supervisor or a supervisee. This dedicated p-unit is also referred to as an autonomous p-unit in some examples. In some examples, all processor cores are of the same size and function, i.e., symmetric cores. However, the processor cores may be asymmetric.
[0030] For example, some processor cores may have different sizes and / or functions from other processor cores. The processor core may be a virtual processor core or a physical processor core. Processor 603_10 may include an integrated voltage regulator (IVR) 607, a power control unit (p unit) 608, a phase-locked loop (PLL) and / or a frequency-locked loop (FLL) 609. Various blocks of processor 603_10 may be coupled via an interface or a fabric. Here, the term "interconnect" refers to a communication link or channel between two or more points or nodes. The interconnect may have one or more separate conductive paths, such as, for example, wires, vias, waveguides, passive components, and / or active components. The interconnect may also have a fabric. In some embodiments, the p unit 608 is coupled to the OS 602 via an interface. Here, the term "interface" generally refers to software and / or hardware used to communicate with the interconnect. The interface may include logic and I / O drivers / receivers for transmitting and receiving data on the interconnect or on one or more wires.
[0031] In some embodiments, each processor 603 is coupled to a power supply via a voltage regulator. The voltage regulator may be internal to the processor system 601 (e.g., on the package of the processor system 601) or external to the processor system 601. In some embodiments, each processor 603 includes an IVR 607 that receives a primary regulated voltage from a voltage regulator of the processor system 601 and generates an operating voltage for the agents of the processor 603. The agents of the processor 603 are various components of the processor 603, including the core 606, the IVR 607, the p unit 608, the PLL / FLL 609, and the like.
[0032] Therefore, the implementation of the IVR 607 can enable fine-grained control of the voltage, and thus the power and performance, of each of the individual cores 606. Accordingly, each core 606 can operate at an independent voltage and frequency, allowing for great flexibility and providing a wide range of opportunities for balancing power consumption with performance. In some embodiments, the use of multiple IVRs enables components to be grouped into separate power planes such that the power regulated by the IVR is supplied only to the components within that group. For example, each core 606 can include an IVR for managing the power supply to that core, and that IVR can receive an input power supply from the regulated output of the voltage regulator of the IVR 607 or the processor system 601. In power management, when a processor core 606 is placed in a particular low-power state, one IVR in a given power domain can be powered down or switched off while another IVR in a different power domain can remain active or be fully powered. Accordingly, the IVR can control the logic of a particular domain or the processor core 606. Here, the term "domain" generally refers to a logical or physical boundary that has similar characteristics (e.g., supply voltage, operating frequency, type of circuit or logic, and / or workload type) and / or is controlled by a particular agent. For example, a domain can be a group of logical or functional units that are controlled by a particular supervisor. A domain can also be referred to as an Autonomous Perimeter (AP). A domain can be the entire system-on-chip (SoC) or a part of the SoC and is governed by a p-unit.
[0033] In some embodiments, each processor 603 includes its own p-unit 608. The p-unit 608 controls the power and / or performance of the processor 603. The p-unit 608 may control the power and / or performance (e.g., IPC, frequency) of each individual core 606. In various embodiments, the p-units 608 of each processor 603 are coupled via fabric 604. Accordingly, the p-units 608 of each processor 603 communicate with others and the OS 602 and determine the optimal power state of the processor system 601 by controlling the power states of the individual cores 606 under their domains.
[0034] The p-unit 608 may include circuitry that includes hardware, software, and / or firmware for performing power management operations for the processor 603. In some embodiments, the p-unit 608 provides control information to a voltage regulator of the processor system 601 via an interface, causing the voltage regulator to generate an appropriate regulated voltage. In some embodiments, the p-unit 608 provides control information to the IVR of the core 606 via another interface to control the operating voltage generated (or to disable the corresponding IVR in a low power mode). In some embodiments, the p-unit 608 may include various power management logic units that perform hardware-based power management. Such power management may be entirely processor-controlled (e.g., by various processor hardware and may be triggered by workload and / or power, thermal or other processor constraints), and / or the power management may be performed in response to an external source (e.g., a platform or power management source or system software). In some embodiments, the p-unit 608 is implemented as a microcontroller.
[0035] The microcontroller may be a built-in microcontroller that is a dedicated controller, or may be a general-purpose controller. In some embodiments, the p unit 608 is implemented as control logic configured to execute its own dedicated power management code, herein referred to as pCode. In some embodiments, the power management operations executed by the p unit 608 may be implemented external to the processor 603, such as by a separate power management integrated circuit (PMIC) 613 or other components external to the processor system 601, for example. In still other embodiments, the power management operations executed by the p unit 608 may be implemented within the BIOS or other system software. In some embodiments, the p unit 608 of the processor 603 may assume the role of supervisor or supervisee.
[0036] Here, the term "supervisor" generally refers to a power controller or power management unit ("p-unit") that monitors and manages power and performance-related parameters for one or more associated power domains, either alone or in collaboration with one or more other p-units. Power / performance-related parameters can include, but are not limited to, domain power, platform power, voltage, voltage domain current, die current, load line, temperature, device latency, utilization, clock frequency, processing efficiency, current / future workload information, and other parameters. It can determine new power or performance parameters (limits, average operation, etc.) for one or more of its domains. And these parameters can be communicated directly to the supervised p-units or to entities to be controlled or monitored, such as VR or clock throttle control registers, via one or more fabrics and / or interconnects. The supervisor knows the workload (current and future) of one or more dies, the power measurements of one or more dies, and other parameters (e.g., platform-level power boundaries), and determines new power limits for one or more dies. And these power limits are communicated by the supervisor's p-unit to the supervised p-units via one or more fabrics and / or interconnects. In an example where a die has one p-unit, the supervisor's (Svor) p-unit is also referred to as the supervisor die.
[0037] Here, the term "supervisee" generally refers to a power controller or power management unit ("p-unit") that monitors and manages power and performance-related parameters for one or more associated power domains, either alone or in cooperation with one or more other p-units, and receives instructions from a supervisor to set power and / or performance parameters (e.g., supply voltage, operating frequency, maximum current, throttling threshold, etc.) for the associated power domains. In an example where the die has one p-unit, the p-unit of the supervisee (Svee) is also referred to as the supervisee die. Note that the p-unit can function as either Svor, Svee, or both p-units of Svor / Svee.
[0038] In various embodiments, the p-unit 608 executes firmware (referred to as pCode) that communicates with the OS 602. In various embodiments, each processor 603 includes a PLL or FLL 609 that generates a clock from the p-unit 608 and an input clock (or reference clock) for each core 606. The core 606 can include or be coupled to an independent clock generation circuit, such as one or more PLLs, to independently control the operating frequency of each core 606. In some embodiments, a memory controller (MC) 610 manages read and / or write operations in one or more memory modules 612. The memory module 612 is coupled to the processor 601 via a memory physical layer (MEM PHY) 611.
[0039] In some embodiments, the processor system 601 includes fuses or NVM that store the Vmin for each of the processor cores 606 among a plurality of processor cores. In various embodiments, the p unit 608 executes firmware (pCode) to rank each processor core 606 according to the Vmin for each processor core and assign the bootstrap processor to the processor core with the highest rank. In some embodiments, the OS 602 schedules interrupts or services for low-utilization tasks on the processor core with the highest rank. In some embodiments, the firmware assigns ranked processor core Advanced Configuration and Power Interface (APIC) identifiers (IDs). In some embodiments, the firmware shares the ranked APIC ID of the processor core with the OS 602. In some embodiments, the firmware shares the ranked APIC ID via the Advanced Configuration and Power Interface (APIC) table. ACPI is an industry specification regarding the efficient handling of power consumption in desktop and mobile computers. ACPI defines how the computer's basic input / output system, operating system, and peripheral devices communicate with each other regarding power usage. In some embodiments, the ACPI table has a Multiple APIC Description Table (MADT). The MADT describes all interrupt controllers in the system. Using it, the currently available processors or cores can be enumerated. In some embodiments, the firmware ranks each processor core 606 based on the efficiency of the processor core near the low-frequency mode frequency.
[0040] Figure 7 shows a V according to some embodiments minFIG. 700 is an example of a flowchart of energy efficient interrupt routing and / or services and low utilization thread scheduling based thereon. Ranking may be initiated by μCode (microcode) firmware. An example of μCode firmware is pCode executed by a power management unit (PMU or p unit). Further, runtime updates are made possible by supporting an interrupt mechanism as defined in the hardware (HW) feedback infrastructure. Such a mechanism is useful when some cores are offline or when other long-term events, such as a VF core based on a reliability stress restrictor (RSR), switch to a new but predefined VF curve. Some computer platforms are arranged to migrate to a different pre-programmed VF curve when the old VF curve is no longer suitable after years of use and wear.
[0041] Here, various blocks are shown in a specific order, but the order can be changed. For example, some blocks or operations can be executed before others, and some blocks or operations can be executed in parallel. These blocks can be executed by hardware, software, or a combination of hardware and software. In various embodiments, this bootstrap flow is changed to identify a bootstrap processor core (BSP). The BSP handles the initialization procedures for the overall system. Those procedures include identifying the characteristics of the system logic, checking the integrity of the memory, starting the remaining processors, and loading the operating system into memory. The BSP is marked as the most energy-efficient core of the SoC and is assigned the lowest APIC (advanced processor interrupt controller) ID (identification) value. The APIC ID value defines the target processor that receives interrupts delivered in logical destination mode in the local x2APIC. This is a 32-bit value initialized by hardware. Although APIC ID is a term in the Intel architecture, similar functions in other processor architectures can also be used to identify the bootstrap processor.
[0042] Initially, the BSP is core 0 or any core of a multi-core system. In a set of heterogeneous cores with large cores (or complex and / or high-power applications) and small cores (less complex and / or low-power applications), the BSP can be one of the small cores. This initial BSP is used to identify an energy-efficient (EE) BSP. However, in some embodiments, a large core can also be used as a DSP.
[0043] In this context, at block 701, microcode (e.g., pCode) or BIOS (basic input / output system) reads the fuses or NVM that stores the per-core V min value. These fuses are programmed in the HVM. When reading the fuses or NVM, the microcode or BIOS calculates and ranks the cores as described with reference to FIG. 5. Thus, the microcode or BIOS calculates and ranks the core APIC IDs based on the efficiency near the LFM frequency. Based on the calculated and ranked cores, at block 702, the microcode or BIOS transfers the BSP ownership to the most efficient core (e.g., having a higher ranking number) by setting a register (e.g., IA32_APIC_BASE.BSP = 1).
[0044] In some embodiments, at block 703, the microcode or BIOS then shares the APIC IDs of the cores ranked based on efficiency with the operating system (or kernel). For example, the microcode or BIOS shares the APIC IDs of the cores ranked based on efficiency with the OS via an ACPI (Advanced Configuration and Power Interface) table such as, for example, MADT (multiple interrupt controller table).
[0045] Using the reordered APIC IDs, at block 704, the OS efficiently services low-utilization tasks, interrupts, and DPCs on the core with the lowest Vmin. Note that the reordering of the APIC IDs also enables more efficient HW interrupt routing. In some embodiments, at block 705, the OS scheduler uses the efficient cores as preferred or favored cores for thread scheduling and / or for other background applications or interrupt services.
[0046] At block 706, a decision regarding the support of dynamic hardware (HW) is made. If the SoC supports dynamic HW feedback, this process proceeds to block 707, where pCode (or suitable microcode or firmware) shares the updated efficiency core ranking via a shared memory or a model-specific register (MSR) interface. If the SoC does not support dynamic HW feedback, this process proceeds to block 705. As described herein, at block 705, the OS scheduler uses an efficient core as a suitable or favored core for thread scheduling and / or for other background applications or interrupt services.
[0047] Elements of embodiments (e.g., flowcharts with reference to various embodiments) are also provided as a machine-readable medium (e.g., memory) storing computer-executable instructions (e.g., instructions for performing any other process described herein). In some embodiments, the computing platform has a memory, a processor, a machine-readable storage medium (also referred to as a tangible machine-readable medium), a communication interface (e.g., a wireless or wired interface), and a network bus, all coupled together.
[0048] In some embodiments, the processor is a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a general-purpose central processing unit (CPU), or low-power logic implementing a simple finite state machine to execute the methods with reference to the various flowcharts and / or various embodiments.
[0049] In some embodiments, the various logical blocks of the system are coupled together via a network bus. Any suitable protocol may be used to implement the network bus. In some embodiments, a machine-readable storage medium includes instructions (also referred to as program software code / instructions) for calculating or measuring the distance and relative orientation of a device with reference to other devices as described with reference to the various embodiments and flowcharts.
[0050] Program software code / instructions that are associated with the various flowcharts and / or flowcharts referenced to the various embodiments and that are executed to implement embodiments of the disclosed subject matter may be implemented as part of an operating system or a particular application, component, program, object, module, routine, or other sequence or compilation of instructions, or as firmware embedded in a processor, and are referred to as "program software code / instructions", "operating system program software code / instructions", "application program software code / instructions", or simply "software". In some embodiments, the program software code / instructions associated with the various flowcharts and / or flowcharts referenced to the various embodiments are executed by the system.
[0051] In some embodiments, the program software code / instructions associated with the various flowcharts and / or references to the various embodiments are stored on a computer-executable storage medium and executed by a processor. Here, the computer-executable storage medium is a tangible machine-readable medium that can be used to store program software code / instructions and data that, when executed by a computing device, cause one or more processors to execute one or more methods as may be described in one or more of the appended claims directed to the disclosed subject matter.
[0052] Tangible machine-readable media can include executable software program code / instructions and data storage at various tangible locations, including, for example, ROM, volatile RAM, non-volatile memory and / or cache and / or other tangible memories as referenced in this application. Portions of this program software code / instructions and / or data can be stored in any of these storage and memory devices. Also, the program software code / instructions can be obtained from other storage via, for example, a centralized server or a peer-to-peer network, including the Internet. Different portions of the software program code / instructions and data can be obtained in different communication sessions at different times or in the same communication session.
[0053] (Associated with references to various flowcharts and / or various embodiments) Software program code / instructions and data can be obtained in their entirety prior to the execution of each software program or application by a computing device. Alternatively, various portions of the software program code / instructions and data may be obtained dynamically when needed for execution, for example, just-in-time. Alternatively, some combination of these techniques for obtaining the software program code / instructions and data may be performed, for example, by way of example, for different applications, components, programs, objects, modules, routines, or other instruction sequences, or compilations of instruction sequences. Thus, it is not required that the data and instructions be entirely on a tangible machine-readable media at a particular instant.
[0054] Examples of tangible computer-readable media include, but are not limited to, volatile and non-volatile memory devices, specifically, read-only memory (ROM), random access memory (RAM), flash memory devices, floppy disks (registered trademark) and other removable disks, magnetic storage media, optical storage media (e.g., compact disc read-only memory (CD ROM), digital versatile disc (DVD), etc.), including recordable and non-recordable types of media. Software program code / instructions may be temporarily stored in a digital physical communication link and implemented in such a physical communication link with an electrical, optical, acoustic, or other form of propagated signal, such as a carrier wave, infrared signal, digital signal, etc.
[0055] Generally, a tangible machine-readable medium includes any tangible mechanism that provides information in a form accessible by a machine (i.e., a computing device), that is, stores and / or transmits in a digital form, such as, for example, data packets, and may be included in a communication device, a computing device, a network device, a portable information terminal, a manufacturing tool, a mobile communication device, such as, for example, an iPhone (registered trademark), a Galaxy (registered trademark), or the like, or any other device including a computing device, whether or not it can download and execute applications and assisted applications from a communication network such as the Internet. In one embodiment, a processor-based system is in the form of or includes a personal digital assistant (PDA), a mobile phone, a notebook computer, a tablet, a game console, a set-top box, an embedded system, a TV (television), a personal desktop computer, etc. Alternatively, in some embodiments of the disclosed subject matter, traditional communication applications and one or more assisted applications may be used.
[0056] In some embodiments, the machine-readable storage medium comprises machine-readable instructions that, when executed, cause one or more processors to execute a method having reading a fuse or NVM storing Vmin for each processor core in a multi-core system. The method further comprises changing an existing bootstrap processor by ranking each processor core of the multi-core system from a highest rank to a lowest rank according to the Vmin for each processor core and assigning a new bootstrap processor to the processor core having the highest rank. In some embodiments, the operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank. In some embodiments, the ranked processor cores have ranked APIC IDs.
[0057] In some embodiments, the method comprises sharing the ranked APIC IDs of the processor cores with the operating system. The sharing of the ranked APIC IDs is via an ACPI table. In some embodiments, the ACPI table has a MADT. In some embodiments, ranking each processor core of the multi-core system is further based on the efficiency of the processor core near a low-frequency mode frequency. In some embodiments, the fuse or NVM is programmed with the Vmin for each processor core during mass production.
[0058] Figure 8 illustrates a boot flow 800 prior to an OS handoff for efficient interrupt routing / service and benefits of OS-specific core parking, according to some embodiments. Here, boot flow 800 is shown operating from top to bottom in time. At block 801, microcode or firmware (e.g., pCode, etc.) logically assigns the first core 802-0 (i.e., core 0) as the BSP. All other cores 802-1, 802-2, and 802-3 expect the BSP core 0 802-0 to perform a comparison before they execute. Thus, cores 802-1, 802-2, and 802-3 are in MWAIT, as indicated by 803. The MWAIT instruction provides a hint for the processor (or core) to enter an implementation-dependent optimized state. In this example, a 4-core system is shown. However, embodiments are applicable to any number and type of cores. Here, the core type can include cores of different functions, sizes, etc.
[0059] In some embodiments, the bootstrap flow identifies the BSP as the most energy-efficient core of the SoC and assigns the lowest APIC ID value to that core. The APIC ID value specifies the target processor that receives interrupts delivered in logical destination mode in the local x2APIC.
[0060] For an OS that supports core parking, there is one possible implementation that uses the underlying novelty. For example, on the Windows® platform, the default mechanism for performing core parking follows the natural enumeration order. Core parking is a feature that dynamically selects a set of processors that should remain idle and not execute any threads, based on their recent utilization and current power policy. This reduces energy consumption and thus, heat and power usage.
[0061] By way of illustration, in a 4-core system, the parking order is 4, 3, 2. However, V minBy associating the rank order with the core ID enumeration order and moving BSP ownership to the most efficient core before the OS handoff, the parking logic can naturally use the ordered sequence (i.e., in this example, park the least efficient core first, which is core 802-0). Also, reordering the cores enables more efficient HW-based interrupt routing and OS-based interrupt / DPC services. In this case, the enumeration of core IDs is done based on the relative efficiency of the individual cores. Note that the degree of benefit of different implementations is specific to that implementation. Independent of the solution implementation, the core principle is relative V min It depends on the HVM manufacturing process for enumerating the core order. The implementation-specific sequences shown through 804, 805, 806, and 807 can achieve the intended ranking that can actually be enforced during enumeration and the boot flow. Among them, at block 808, the OS handoff is completed.
[0062] Figure 9 illustrates a smart device or computer system or SoC (System-on-Chip) having firmware for energy-efficient interrupt routing / service and low-utilization thread scheduling based on V min Here, any of the blocks may have logic to gate or ungate the clock according to a security key. In some embodiments, the SoC includes a cryptographic engine that generates keys for other IP blocks within the platform. It should be noted that elements of Figure 9 having the same reference numerals (or names) as elements in other figures can operate or function in a similar manner as described, but are not so limited.
[0063]
[0064] In some embodiments, device 5500 represents a suitable computing device, such as, for example, a computing tablet, a cellular phone or smartphone, a laptop, a desktop, an Internet of Things (IoT) device, a server, a wearable device, a set-top box, a wireless-enabled e-book reader, or the like. It is understood that certain components are shown schematically and not all components of such a device are shown in device 5500.
[0065] In one example, device 5500 includes a system-on-chip (SoC) 5501. An example of the boundary of SoC 5501 is shown using a dashed line in FIG. 5, and examples of some components are shown as being included within SoC 5501, but SoC 5501 can include any suitable components of device 5500.
[0066] In some embodiments, device 5500 includes a processor 5504. Processor 5504 can include one or more physical devices, such as, for example, a microprocessor, an application processor, a microcontroller, a programmable logic device, a processing core, or other processing implementations such as, for example, a non-integrated combination of multiple computing, graphics, accelerators, I / O, and / or other processing chips. The processing operations performed by processor 5504 include the execution of an operating platform or operating system on which applications and / or device functions are executed. Those processing operations include operations related to I / O (input / output) with a human user or other devices, operations related to power management, operations related to connecting computing device 5500 to other devices, and / or the like. Those processing operations can also include operations related to audio I / O and / or display I / O.
[0067] In some embodiments, processor 5504 includes a plurality of processing cores (also referred to as cores) 5508a, 5508b, 5508c. Although only three cores 5508a, 5508b, 5508c are shown in FIG. 9, processor 5504 may include any other suitable number of processing cores, such as dozens or hundreds of processing cores. Processor cores 5508a, 5508b, 5508c may be implemented on a single integrated circuit (IC) chip. Further, the chip may include one or more shared caches and / or private caches, buses or interconnects, graphics controllers and / or memory controllers, or other components.
[0068] In some embodiments, processor 5504 includes cache 5506. In one example, a section of cache 5506 may be dedicated to an individual core 5508 (e.g., the first section of cache 5506 is dedicated to core 5508a, the second section of cache 5506 is dedicated to core 5508b, and so on). In one example, one or more sections of cache 5506 may be shared among two or more of cores 5508. Cache 5506 may be divided into different hierarchies, such as a level 1 (L1) cache, a level 2 (L2) cache, a level 3 (L3) cache, and the like.
[0069] In some embodiments, the processor core 5504 may include a fetch unit that fetches instructions (including instructions with conditional branches) for execution by the core 5504. The instructions may be fetched from any storage device such as, for example, the memory 5530. The processor core 5504 may also include a decode unit that decodes the fetched instructions. For example, the decode unit may decode the fetched instructions into a plurality of micro-operations. The processor core 5504 may include a schedule unit that performs various operations associated with storing the decoded instructions. For example, the schedule unit may hold data from the decode unit until the instructions are ready for dispatch (e.g., until all source values of the decoded instructions become available). In one embodiment, the schedule unit may schedule and / or issue (or dispatch) the decoded instructions to the execution unit for execution.
[0070] The execution unit may execute the dispatched instructions after they have been decoded (e.g., by the decode unit) and dispatched (e.g., by the schedule unit). In one embodiment, the execution unit may include two or more execution units (such as, for example, an imaging calculation unit, a graphics calculation unit, a general-purpose calculation unit, etc.). The execution unit may also perform various arithmetic operations such as, for example, addition, subtraction, multiplication, and / or division, and may include one or more arithmetic logic units (ALUs). In one embodiment, a coprocessor (not shown) may perform various arithmetic operations together with the execution unit.
[0071] Also, the execution unit may issue instructions out of order. Thus, in one embodiment, the processor core 5504 may be an out-of-order processor core. The processor core 5504 may also include a retirement unit. The retirement unit may retire the executed instructions after they are committed. In one embodiment, the retirement of the executed instructions may result in, for example, the processor state being committed from the execution of the instructions and the physical registers used by the instructions being deallocated. The processor core 5504 may also include a bus unit that enables communication via one or more buses between components of the processor core 5504 and other components. The processor core 5504 may also include one or more registers that store data accessed by various components of the core 5504 (e.g., values related to assigned application priorities and / or subsystem state (mode) associations).
[0072] In some embodiments, the apparatus 5500 includes a connection circuit 5531. For example, the connection circuit 5531 includes a hardware device (e.g., a wireless and / or wired connector and communication hardware) and / or a software component (e.g., a driver, a protocol stack), and enables, for example, the apparatus 5500 to communicate with an external device. The apparatus 5500 may be separated from an external device such as, for example, another computing device, a wireless access point, or a base station.
[0073] In one example, connection circuit 5531 can include a plurality of different types of connections. Generally speaking, connection circuit 5531 may include a cellular connection circuit, a wireless connection circuit, and the like. The cellular connection circuit of connection circuit 5531 generally refers to a cellular network connection provided by a wireless carrier, such as, for example, GSM (Global System for Mobile Communications) or variations or derivatives thereof, CDMA (Code Division Multiple Access) or variations or derivatives thereof, TDM (Time Division Multiplexing) or variations or derivatives thereof, 3rd Generation Partnership Project (3GPP) UMTS (Universal Mobile Telecommunications Systems) system or variations or derivatives thereof, 3GPP Long Term Evolution (LTE) system or variations or derivatives thereof, 3GPP LTE Advanced (LTE-A) system or variations or derivatives thereof, 5th Generation (5G) wireless system or variations or derivatives thereof, 5G mobile network system or variations or derivatives thereof, 5G New Radio (NR) system or variations or derivatives thereof, or other cellular service standards. The wireless connection circuit (or wireless interface) of connection circuit 5531 refers to a non-cellular wireless connection and can include personal area networks (such as, for example, Bluetooth (registered trademark), near field, etc.), local area networks (such as, for example, Wi-Fi, etc.), and / or wide area networks (such as, for example, WiMax, etc.), and / or other wireless communications. In one example, connection circuit 5531 can include a network interface, such as, for example, a wired or wireless interface, so that the system embodiment can be incorporated into a wireless device such as a mobile phone or a personal digital assistant.
[0074] In some embodiments, the apparatus 5500 includes a control hub 5532 that represents hardware devices and / or software components related to interaction with one or more I / O devices. For example, the processor 5504 can communicate with one or more of the display 5522, one or more peripheral devices 5524, the storage device 5528, one or more other external devices 5529, etc. via the control hub 5532. The control hub 5532 can be a chipset, a platform control hub (PCH), and / or the like.
[0075] For example, the control hub 5532 exemplifies one or more connection points for additional devices connected to the apparatus 5500, through which, for example, a user can interact with the system. For example, devices that can be attached to the apparatus 5500 (e.g., device 5529) include a microphone device, a speaker or stereo system, an audio device, a video system or other display device, a keyboard or keypad device, or other I / O devices used in specific applications such as, for example, a card reader or other device.
[0076] As described above, the control hub 5532 can interact with audio devices, the display 5522, and the like. For example, an input via a microphone or other audio device can provide an input or command for one or more applications or functions of the device 5500. Further, audio output can be provided instead of, or in addition to, the display output. In another example, when the display 5522 includes a touch screen, the display 5522 also functions as an input device that can be at least partially managed by the control hub 5532. There can also be additional buttons or switches on the computing device 5500 to provide I / O functions managed by the control hub 5532. In one embodiment, the control hub 5532 manages devices such as, for example, an accelerometer, a camera, a light sensor, or other environmental sensors, or other hardware that can be included in the device 5500. The input can be part of a direct user interaction and can also provide environmental input to the system to affect its operation (e.g., noise filtering, display adjustment for luminance detection, application of a flash for a camera, or other mechanisms).
[0077] In some embodiments, the control hub 5532 can be coupled to various devices using any suitable communication protocol, such as, for example, PCIe (Peripheral Component Interconnect Express), USB (Universal Serial Bus), Thunderbolt, High-Definition Multimedia Interface (HDMI), FireWire, and the like.
[0078] In some embodiments, display 5522 represents hardware (e.g., a display device) and software (e.g., a driver) components that provide a visual display and / or a tactile display for a user to interact with device 5500. Display 5522 may include a display interface, a display screen, and / or a hardware device used to provide a display to the user. In some embodiments, display 5522 includes a touch screen (or touch pad) device that provides both output and input to the user. In one example, display 5522 may communicate directly with processor 5504. Display 5522 may be an internal display, such as in a mobile electronics device or a laptop device, or one or more external display devices attached via a display interface (e.g., DisplayPort, etc.). In one embodiment, display 5522 may be a head-mounted display (HMD), such as a stereoscopic display device used in, for example, a virtual reality (VR) application or an augmented reality (AR) application.
[0079] In some embodiments, although not shown in the figures, in addition to (or instead of) processor 5504, device 5500 may include a graphics processing unit (GPU) that includes one or more graphics processing cores capable of controlling one or more aspects of displaying content on display 5522.
[0080] Control hub 5532 (or platform controller hub) may include a hardware interface and connectors, and software components (e.g., drivers, protocol stacks) for making peripheral connections to, for example, peripheral device 5524.
[0081] It should be understood that the device 5500 may be a peripheral device for other computing devices and may also have peripheral devices connected thereto. The device 5500 may have a "docking" connector for connecting to other computing devices for purposes such as managing (e.g., downloading and / or uploading, modifying, synchronizing) the content on the device 5500. Additionally, the docking connector may enable the device 5500 to connect to certain peripheral devices that allow, for example, the computing device 5500 to control the output of content to an audio-visual system or other system.
[0082] In addition to a dedicated docking connector or other dedicated connection hardware, the device 5500 can make peripheral connections via common or standard-based connectors. Common types can include Universal Serial Bus (USB) connectors (which can include any of a number of different hardware interfaces), Mini DisplayPort (MDP), High-Definition Multimedia Interface (HDMI), FireWire, or other types.
[0083] In some embodiments, the connection circuit 5531 may be coupled to the control hub 5532 in addition to or instead of being directly coupled to the processor 5504. In some embodiments, the display 5522 may be coupled to the control hub 5532 in addition to or instead of being directly coupled to the processor 5504.
[0084] In some embodiments, the device 5500 includes a memory 5530 coupled to the processor 5504 via a memory interface 5534. The memory 5530 includes memory devices for storing information within the device 5500.
[0085] In some embodiments, memory 5530 includes an apparatus for maintaining and managing stable clock generation, as described with reference to various embodiments. The memory can include non-volatile (state does not change when power to the memory device is interrupted) and / or volatile (state becomes uncertain when power to the memory device is interrupted) memory devices. Memory device 5530 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or some other memory device having performance suitable for functioning as a process memory. In one embodiment, memory 5530 can operate as the system memory of apparatus 5500 for storing data and instructions used when one or more processors 5504 execute an application or process. Memory 5530 can store application data, user data, music, photos, documents, or other data, as well as system data (whether long-term or temporary) related to the execution of the applications and functions of apparatus 5500.
[0086] Elements of various embodiments and examples are also provided as a machine-readable medium (e.g., memory 5530) storing computer-executable instructions (e.g., instructions for executing any other process described herein). Machine-readable media (e.g., memory 5530) can include, but are not limited to, flash memory, optical disks, CD-ROMs, DVD ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, phase change memory (PCM), or other types of machine-readable media suitable for storing electronic instructions or computer-executable instructions. For example, the disclosed embodiments can be downloaded as a computer program (e.g., BIOS) that can be transferred from a remote computer (e.g., a server) to a requesting computer (e.g., a client) by a data signal via a communication link (e.g., a modem or network connection).
[0087] In some embodiments, device 5500 includes a temperature measurement circuit 5540, for example, to measure the temperature of various components of device 5500. In one example, the temperature measurement circuit 5540 can be built into, coupled to, or attached to various components whose temperature is to be measured and monitored. For example, the temperature measurement circuit 5540 can measure the temperature (or the temperature therein) of one or more of core 5508a, 5508b, 5508c, voltage regulator 5514, memory 5530, the motherboard of SoC 5501, and / or any suitable component of device 5500. In some embodiments, the temperature measurement circuit 5540 includes a low power hybrid reverse (LPHR) bandgap reference (BGR) and a digital temperature sensor (DTS), which utilize a subthreshold metal oxide semiconductor (MOS) transistor and a PNP parasitic bipolar junction transistor (BJT) device to form a reverse BGR that serves as the basis for a configurable BGR or DTS operating mode. The LPHR architecture uses low-cost MOS transistors and standard parasitic PNP devices. Based on the reverse bandgap voltage, the LPHR can operate as a configurable BGR. By comparing the configurable BGR with a scaled base-emitter voltage, this circuit can also act as a DTS with a linear transfer function by high-precision single temperature trimming.
[0088] In one embodiment, the apparatus 5500 includes a power measurement circuit 5542 for measuring, for example, the power consumed by one or more components of the apparatus 5500. In one example, in addition to or instead of measuring power, the power measurement circuit 5542 may measure voltage and / or current. In one example, the power measurement circuit 5542 may be incorporated, coupled, or attached to various components whose power, voltage, and / or current should be measured and monitored. For example, the power measurement circuit 5542 may measure the power, current, and / or voltage supplied by one or more voltage regulators 5514, the power supplied to the SoC 5501, the power supplied to the apparatus 5500, the power consumed by the processor 5504 (or other components) of the apparatus 5500, and the like.
[0089] In some embodiments, apparatus 5500 includes one or more voltage regulator circuits, generally referred to as voltage regulator (VR) 5514. VR 5514 generates a signal at an appropriate voltage level that can be supplied to operate any suitable component of apparatus 5500. By way of example only, VR 5514 is illustrated as supplying a signal to processor 5504 of apparatus 5500. In some embodiments, VR 5514 receives one or more voltage identification (VID) signals and generates a voltage signal at an appropriate level based on the VID signal. Various types of VRs may be utilized for VR 5514. For example, VR 5514 may include a “buck” VR, a “boost” VR, a combination of a buck VR and a boost VR, a low dropout (LDO) regulator, a switching DC-DC regulator, a constant on-time controller-based DC-DC regulator, and the like. A buck VR is generally used in power delivery applications where it is necessary to convert an input voltage to an output voltage at a ratio less than 1. A boost VR is generally used in power delivery applications where it is necessary to convert an input voltage to an output voltage at a ratio greater than 1. In some embodiments, each processor core has its own VR, which is controlled by PCU 5510a / b and / or PMIC 5512. In some embodiments, each core has a network of distributed LDOs to provide efficient control for power management. The LDO may be a digital LDO, an analog LDO, or a combination of digital and analog LDOs. In some embodiments, VR 5514 includes a current tracking device that measures the current through the (one or more) power rails.
[0090] In some embodiments, VR5514 includes a digital control scheme for managing the state of a proportional-integral-derivative (PID) filter (also known as a digital type III compensator). The digital control scheme controls the integrator of the PID filter to implement a non-linear control that saturates the duty cycle, during which the proportional and derivative terms of the PID are set to 0, and the integrator and its internal state (previous value or memory) are set to a duty cycle that is the sum of adding delta D to the current nominal duty cycle. Delta D is the maximum duty cycle increment used to regulate the voltage regulator from ICCmin to ICCmax, and is a set register that can be set after the silicon process. The state machine transitions from a non-linear all-ON state (returning the output voltage Vout to the regulation window) to an open-loop duty cycle that maintains the output voltage slightly higher than the required reference voltage Vref. After a certain period in this open-loop state at the commanded duty cycle, the state machine ramps down the open-loop duty cycle value until the output voltage approaches the commanded Vref. Thus, output chatter on the output supply from VR5514 is completely eliminated (or substantially eliminated), and there is only a single undershoot transition, which can lead to a guaranteed Vmin based on the di / dt of the comparator delay and the load with available output decoupling capacitance.
[0091] In some embodiments, VR5514 includes a separate self-starting controller that operates without using fuses (or NVM) and / or trim information. The self-starting controller protects VR5514 from large inrush currents and voltage overshoots while being able to follow a variable VID (Voltage Identification) reference ramp provided by the system. In some embodiments, the self-starting controller uses a relaxation oscillator built into the controller to set the switching frequency of the buck converter. This oscillator can be initialized using either a clock or a current reference to get close to the desired operating frequency. The output of VR5514 is weakly coupled to this oscillator to set the duty cycle of the closed-loop operation. The controller is naturally biased so that the output voltage is always slightly higher than the set point, eliminating the need for trims imposed by process, voltage, and / or temperature (PVT).
[0092] In some embodiments, device 5500 includes one or more clock generator circuits, generally referred to as clock generator 5516. Clock generator 5516 generates a clock signal at an appropriate frequency level that can be supplied to any suitable component of device 5500. By way of example only, clock generator 5516 is illustrated as supplying a clock signal to processor 5504 of device 5500. In some embodiments, clock generator 5516 receives one or more Frequency Identification (FID) signals and generates a clock signal at an appropriate frequency based on the FID signals.
[0093] In some embodiments, device 5500 includes a battery 5518 that powers various components of device 5500. By way of example only, battery 5518 is shown as powering processor 5504. Although not shown in the figure, device 5500 may include a charging circuit for recharging the battery, for example, based on an alternating current (AC) power source received from an AC adapter.
[0094] In some embodiments, battery 5518 is charged to a preset voltage (e.g., 4.1V) and periodically checks the actual battery capacity or energy. The battery then determines the battery capacity or energy. If the capacity or energy is insufficient, a device within or coupled to the battery slightly increases the charging voltage until the capacity is sufficient (e.g., from 4.1V to 4.11V). This process of periodically checking and slightly increasing the charging voltage is performed until the charging voltage reaches the specification limit (e.g., 4.2V). The scheme described herein has benefits such as, for example, extending battery life, reducing the risk of insufficient energy storage, enabling burst power to be used for as long as possible, and / or enabling even higher burst power to be used.
[0095] In some embodiments, battery 5518 is a multi-battery system having a workload-dependent load sharing mechanism. This mechanism is an energy management system that operates in three modes: an energy saving mode, a balancer mode, and a turbo mode. The energy saving mode is the normal mode in which multiple batteries (collectively shown as battery 5518) supply power to their own load sets with minimum resistance dissipation. In the balance mode, the batteries are connected to each other via switches operating in the active mode such that the shared current is inversely proportional to the corresponding battery charge state. In the turbo mode, both batteries are connected in parallel via a switch (e.g., an on-switch) to supply maximum power to the processor or load. In some embodiments, battery 5518 is a hybrid battery having a high-speed charging battery and a high-energy density battery. The high-speed charging battery (FC) means a battery that can be charged faster than the high-energy density battery (HE). Since today's lithium-ion batteries can be charged faster than HE, FC may be such a battery. In some embodiments, a controller (part of battery 5518) optimizes the sequence and charging speed of the hybrid battery to maximize both the charging current and the charging speed of the battery while allowing for a longer battery life.
[0096] In some embodiments, the charging circuit (e.g., 5518) has a buck-boost converter. This buck-boost converter has DrMOS or DrGaN devices that are used in place of the half-bridge in the case of a traditional buck-boost converter. Here, various embodiments will be described with reference to DrMOS. However, the embodiments are also applicable to DrGaN. The DrMOS device enables better efficiency in power conversion due to reduced parasitic components and optimized MOSFET packaging. Since dead-time management is inherent in DrMOS, the dead-time management is more accurate than that of a traditional buck-boost converter, leading to higher conversion efficiency. A higher operating frequency enables a smaller inductor size, which in turn reduces the height of the charger having a DrMOS-based buck-boost converter. The buck-boost converters of various embodiments have a double-folded bootstrap for DrMOS devices. In some embodiments, in addition to the traditional bootstrap capacitor, a folded bootstrap capacitor that cross-couples the inductor node to two sets of DrMOS switches is added.
[0097] In some embodiments, apparatus 5500 includes a power control unit (PCU) 5510 (also referred to as a power management unit (PMU), a power management controller (PMC), a power unit (p-unit), etc.). In one example, some sections of the PCU 5510 may be implemented by one or more processing cores 5508, and these sections of the PCU 5510 are symbolically illustrated with a dotted box labeled as PCU 5510a. In one example, some other sections of the PCU 5510 may be implemented outside of the processing core 5508, and these sections of the PCU 5510 are symbolically illustrated with a dotted box labeled as PCU 5510b. The PCU 5510 may implement various power management operations related to the apparatus 5500. The PCU 5510 may include a hardware interface, a hardware circuit, a connector, a register, etc., and software components (e.g., a driver, a protocol stack) to implement various power management operations related to the apparatus 5500.
[0098] In various embodiments, the PCU or PMU 5510 is hierarchically organized to form hierarchical power management (HPM). The HPM of various embodiments enables package-level management for the platform while still accommodating islands of autonomy that may exist across the dies of the components within the package, and constructs the infrastructure. The HPM does not assume a predetermined mapping of physical partitions to domains. The HPM domains can be aligned with the functions integrated within a dielet, aligned with the dielet boundaries, aligned with one or more dielets, aligned with a companion die, or even aligned with discrete CXL devices. The HPM addresses the integration of multiple instances of the same die in a discrete form factor, mixed with proprietary or third-party functions integrated on the same die or separate dies, and further with accelerators connected via CXL (e.g., Flexbus) that may be present within the package.
[0099] HPM enables designers to meet the goals of scalability, modularity, and late binding. HPM also enables leveraging PMU functions that may already exist on other dies instead of being disabled in a flat scheme. HPM enables management of any set of functions, regardless of their integration level. HPM in various embodiments is scalable, modular, operates with symmetric multi-chip processors (MCPs), and operates with asymmetric MCPs. For example, HPM does not require the signal PM controller and package infrastructure to grow beyond reasonable scaling limits. HPM enables adding dies into the package later without changing the infrastructure of the base die. HPM addresses the need for non-converged solutions that couple dies of different process technology nodes within a single package. HPM also addresses the needs for companion die integration solutions both on and off the package.
[0100] In various embodiments, each die (or dielet) includes a power management unit (PMU) or p-unit. For example, a processor die can have a supervisor p-unit, a supervisee p-unit, or a dual-role supervisor / supervisee p-unit. In some embodiments, an I / O die has its own dual-role p-unit, such as, for example, a supervisor and / or a supervisee p-unit. The p-units within each die can be instances of a general p-unit. In one such example, all p-units have the same capabilities and circuitry, but are (dynamically or statically) configured to assume the role of supervisor, supervisee, and / or both. In some embodiments, the p-unit of a compute die is an instance of a compute p-unit, and the p-unit of an I / O die is an instance of an I / O p-unit different from the compute p-unit. Depending on the role, the p-unit obtains specific responsibilities for managing the power of the multi-chip module and / or the computing platform. Although various p-units are described with respect to the dies within a multi-chip module or a system-on-chip, the p-unit can also be part of an external device, such as, for example, an I / O device.
[0101] Here, the various p-units need not be the same. The HPM architecture can operate very different types of p-units. One common function of the p-units is that they are expected to receive and be able to understand HPM messages. In some embodiments, the p-units of the IO die may be different from those of the compute die. For example, the number of register instances of each class of registers within the IO p-unit may be different from that in the p-units of the compute die. The IO die may have the ability to be the HPM supervisor of CXL-connected devices, while the compute die may not necessarily have that ability. The IO die and the compute die may also have different firmware flows and different firmware images. These are choices that the implementation can make. The HPM architecture can choose to have one supersized firmware image and can selectively execute the flow corresponding to the type of die to which the firmware is associated. Alternatively, there can be custom firmware for each p-unit type, which makes it possible to further optimize the sizing of the firmware storage requirements for each p-unit type.
[0102] The p-units within each die can be configured to act as supervisor p-units, as supervised p-units, or to have both supervisor / supervised roles. Thus, the p-units can serve as supervisors or supervised for various domains. In various embodiments, each instance of the p-unit can autonomously manage local dedicated resources and includes a structure that aggregates data and communicates between instances to enable shared resource management by an instance configured as a shared resource supervisor. A message and wire-based infrastructure is provided that can be redundantly configured to facilitate management and flow between multiple p-units.
[0103] In some embodiments, power and thermal thresholds are communicated from the supervisor p unit to the supervised p unit. For example, the supervisor p unit knows the workload (current and future) of each die, the power measurements of each die, and other parameters (e.g., platform-level power boundaries), and determines new power limits for each die. These power limits are then communicated from the supervisor p unit to the supervised p unit via one or more interconnects and fabrics. In some embodiments, the fabric refers to a group of fabrics and interconnects that includes a first fabric, a second fabric, and a high-speed response interconnect. In some embodiments, the first fabric is used for general communication between the supervisor p unit and the supervised p unit. These general communications include changes to the voltage, frequency, and / or power state of the die planned based on a number of factors (e.g., future workload, user behavior, etc.). In some embodiments, the second fabric is used for higher-priority communication between the supervisor p unit and the supervised p unit. Examples of higher-priority communication include messages to throttle due to possible thermal runaway conditions, reliability issues, etc. In some embodiments, the high-speed response interconnect is used to communicate high-speed throttling or hard throttling of all dies. In this case, the supervisor p unit can send high-speed throttle messages to, for example, all other p units. In some embodiments, the high-speed response interconnect is a legacy interconnect whose function can be performed by the second fabric.
[0104] The HPM architectures of various embodiments enable the scalability, modularity, and later coupling of symmetric and / or asymmetric dies. Here, symmetric dies are dies of the same size, type, and / or function, and asymmetric dies are dies of different sizes, types, and / or functions. The hierarchical approach also enables leveraging PMU functions that may already exist on other dies, as opposed to being disabled in traditional flat power management schemes. HPM does not assume a predetermined mapping of physical partitions to domains. HPM domains can be aligned with functions integrated within a dielet, aligned with dielet boundaries, aligned with one or more dielets, aligned with companion dies, or even aligned with discrete CXL devices. HPM enables the management of any set of functions, regardless of their level of integration. In some embodiments, a p-unit declares a supervisor p-unit based on one or more factors. Those factors include memory size, physical constraints (e.g., the number of pinouts), and the location of sensors that determine the physical limits of the processor (e.g., temperature, power consumption, etc.).
[0105] The HPM architectures of various embodiments provide a means of scaling power management so that a single p-unit instance does not need to recognize the entire processor. This enables power management at a smaller granularity, improving response time and effectiveness. The hierarchical structure maintains a monolithic view for the user. For example, at the operating system (OS) level, the HPM architecture gives the OS a single PMU view even if its PMU is physically distributed in one or more supervisor-subordinate configurations.
[0106] In some embodiments, the HPM architecture is centralized and one supervisor controls all the supervisees. In some embodiments, the HPM architecture is decentralized and various p-units in various dies control the overall power management through peer-to-peer communication. In some embodiments, the HPM architecture is distributed and different supervisors exist in different domains. An example of a distributed architecture is a tree architecture.
[0107] In some embodiments, device 5500 includes a power management integrated circuit (PMIC) 5512, for example, to implement various power management operations related to device 5500. In some embodiments, PMIC 5512 is a reconfigurable power management IC (RPMIC) and / or an IMVP (Intel® Mobile Voltage Positioning). In one example, the PMIC is in an IC die separate from processor 5504. This can implement various power management operations related to device 5500. PMIC 5512 can include a hardware interface, hardware circuits, connectors, registers, etc., and software components (e.g., drivers, protocol stacks) to implement various power management operations related to device 5500.
[0108] In one example, device 5500 includes one or both of PCU 5510 or PMIC 5512. In one example, either PCU 5510 or PMIC 5512 may not be present within device 5500, and thus these components are illustrated using dashed lines.
[0109] The various power management operations of device 5500 can be performed by PCU 5510, by PMIC 5512, or by a combination of PCU 5510 and PMIC 5512. For example, PCU 5510 and / or PMIC 5512 can select power states (e.g., P-states) for various components of device 5500. For example, PCU 5510 and / or PMIC 5512 can select power states for various components of device 5500 (e.g., in accordance with the ACPI (Advanced Configuration and Power Interface) specification). As a mere example, PCU 5510 and / or PMIC 5512 can transition various components of device 5500 to a sleep state, an active state, an appropriate C-state (e.g., C0 state in accordance with the ACPI specification, or other appropriate C-state), etc. In one example, PCU 5510 and / or PMIC 5512 can control the voltage output by VR 5514 and / or the frequency of the clock signal output by the clock generator, for example, by outputting respective VID signals and / or FID signals. In one example, PCU 5510 and / or PMIC 5512 can control functions related to battery power usage, charging of battery 5518, and power saving operations.
[0110] The clock generator 5516 can have a phase-locked loop (PLL), a frequency-locked loop (FLL), or any suitable clock source. In some embodiments, each core of the processor 5504 has its own clock source. In this way, each core can operate at a frequency independent of the operating frequencies of the other cores. In some embodiments, the PCU 5510 and / or the PMIC 5512 perform adaptive or dynamic frequency scaling or adjustment. For example, the clock frequency of a certain processor core can be increased if the core is not operating at its maximum power consumption threshold or limit value. In some embodiments, the PCU 5510 and / or the PMIC 5512 determine the operating conditions of each core of the processor, and when the PCU 5510 and / or the PMIC 5512 determine that a core is operating below a target performance level, the clock source of that core (e.g., the PLL of that core) can adjust the frequency and / or power supply voltage of that core opportunistically without losing lock. For example, when a core is drawing less current from the power supply rail than the total current assigned to that core or the processor 5504, the PCU 5510 and / or the PMIC 5512 can temporarily increase the power draw for that core or the processor 5504 (e.g., by increasing the clock frequency and / or the power supply voltage level) so that the core or the processor 5504 can exhibit a higher performance level. In this way, the voltage and / or frequency of the processor 5504 can be temporarily increased without compromising the reliability of the product.
[0111] In one example, the PCU 5510 and / or the PMIC 5512 may perform power management operations based at least in part on receiving, for example, measurement results from the power measurement circuit 5542, the temperature measurement circuit 5540, the charge level of the battery 5518, and / or other suitable information that may be used for power management. For this purpose, the PMIC 5512 is communicatively coupled to one or more sensors to detect / detect various values / variations in one or more factors that affect the power / thermal behavior of the system / platform. Examples of such one or more factors include current, voltage droop, temperature, operating frequency, operating voltage, power consumption, inter-core communication activity, and the like. One or more of these sensors may be provided in physical proximity to (and / or in thermal contact / coupled to) one or more components or logic / IP blocks of the computing system. Further, in at least one embodiment, by directly coupling (one or more) sensors to the PCU 5510 and / or the PMIC 5512, it may be possible for the PCU 5510 and / or the PMIC 5512 to manage processor core energy based at least in part on (one or more) values detected by one or more of those sensors.
[0112] An example of the software stack of device 5500 is also shown (however, not all elements of the software stack are shown). By way of example only, processor 5504 may execute application program 5550, operating system 5552, one or more power management (PM) application programs (e.g., generally referred to as PM application 5558), and / or the like. PM application 5558 may also be executed by PCU 5510 and / or PMIC 5512. OS 5552 may also include one or more PM applications 5556a, 5556b, 5556c. OS 5552 may also include various drivers 5554a, 5554b, 5554c, some of which may be specific to power management purposes. In some embodiments, device 5500 may further include a basic input / output system (BIOS) 5520. BIOS 5520 may communicate with OS 5552 (e.g., via one or more drivers 5554), communicate with processor 5504, and so on.
[0113] For example, one or more of PM applications 5558, 5556, drivers 5554, BIOS 5520, etc. may be used to implement power management tasks such as, for example, controlling the voltage and / or frequency of various components of device 5500, controlling the wake-up state, sleep state, and / or other appropriate power states of various components of device 5500, controlling functions related to battery power usage, charging of battery 5518, power saving operations, and so on.
[0114] In some embodiments, battery 5518 is a Li metal battery with a pressure chamber that enables uniform pressure on the battery. The pressure chamber is supported by a metal plate (e.g., a pressure equalizing plate) used to apply uniform pressure to the battery. The pressure chamber may include a pressurized gas, an elastic material, a spring plate, etc. The outer skin of the pressure chamber is constrained at its edges by a (metal) film and bends freely, yet still applies uniform pressure to the plate pressing on the battery cells. The pressure chamber is used to apply uniform pressure to the battery, enabling a high energy density battery, for example, with a battery life 20% longer.
[0115] In some embodiments, battery 5518 includes a hybrid technology. For example, a mixture of a device carrying a high energy density charge (e.g., a Li-Ion battery) and a device carrying a low energy density charge (e.g., a supercapacitor) is used as a battery or energy storage device. In some embodiments, a controller (e.g., hardware, software, or a combination thereof) is used to analyze the peak power pattern, minimizing the impact on the total life of the high energy density charge-carrying device-based battery cells while maximizing the service time for a peak power shaving function. The controller may be part of battery 5518 or part of p unit 5510b.
[0116] In some embodiments, the pCode executed on the PCU5510a / b has the ability to enable additional computational and remote measurement (telemetry) resources for runtime support of the pCode. Here, the pCode refers to the firmware executed by the PCU5510a / b to manage the performance of the SoC5501. For example, the pCode may set the frequency and appropriate voltage for the processor. A portion of the pCode is accessible via the OS5552. In various embodiments, mechanisms and methods are provided for dynamically changing the Energy Performance Preference (EPP) based on workload, user behavior, and / or system conditions. There may be a well-defined interface between the OS5552 and the pCode. The interface can enable or facilitate software settings of several parameters and / or provide hints to the pCode. As an example, the EPP parameter notifies the pCode algorithm as to whether performance is more important or battery life is more important.
[0117] This support can also be done by the OS as follows, which includes machine learning support as part of the OS 5552, and by adjusting the EPP values that the OS gives hints to the hardware (e.g., various components of the SoC 5501) according to machine learning predictions, or by delivering the machine learning predictions to the pCode in a similar way as done by the Dynamic Tuning Technology (DTT) driver. In this model, the OS 5552 can be assumed to see the same set of telemetry available to the DTT. As a result of the DTT machine learning hint settings, the pCode can adjust its internal algorithms to achieve optimal power and performance results according to the activated type of machine learning predictions. As an example, the pCode can increase the bias towards energy savings either by increasing the responsibility for processor utilization changes to enable a fast response to user activities, or by reducing the responsibility for processor utilization or increasing the lost performance by reducing more power and adjusting the energy-saving optimization. This approach can facilitate saving more battery life when the types of activities that are enabled result in losing performance levels beyond what the system can afford. The pCode can include an algorithm for dynamic EPP, which can selectively choose to take two inputs, one from the OS 5552 and another from software such as DTT, to provide higher performance and / or responsiveness. As part of this method, the pCode can enable the option to adjust its reaction to the DTT for multiple different types of activities in the DTT.
[0118] In some embodiments, the pCode improves the performance of the SoC in battery mode. In some embodiments, the pCode enables a dramatically higher SoC peak power limit level (and thus higher turbo performance) in battery mode. In some embodiments, the pCode implements power throttling and is part of Intel's Dynamic Tuning Technology (DTT). In various embodiments, the peak power limit is referred to as PL4. However, the embodiments are applicable to other peak power limits as well. In some embodiments, the pCode sets the Vth threshold voltage (the voltage level at which the platform will throttle the SoC) so that the system does not experience an unexpected shutdown (or go into a black screen). In some embodiments, the pCode calculates the Psoc,pk SoC peak power limit (e.g., PL4) according to the threshold voltage (Vth). These are two dependent parameters, and if one is set, the other can be calculated. The pCode is used to optimally set one of the parameters (Vth) based on system parameters and the operation history. In some embodiments, the pCode provides a scheme for dynamically calculating the throttling level (Psoc,th) based on the available battery power (which changes slowly) to set the SoC throttling peak power (Psoc,th). In some embodiments, the pCode determines the frequency and voltage based on Psoc,th. In this case, the throttling event has little negative impact on the SoC performance. Various embodiments provide a scheme that enables the maximum performance (Pmax) framework to operate.
[0119] In some embodiments, VR5514 includes a current sensor for sensing and / or measuring the current passing through the high-side switch of VR5514. In some embodiments, the current sensor uses an amplifier that has an input capacitively coupled in feedback to sense the input offset of the amplifier, and the input offset can be compensated during measurement. In some embodiments, the amplifier with an input capacitively coupled in feedback is used to operate the amplifier in a region where the input common-mode specifications are relaxed such that the feedback loop gain and / or bandwidth is higher. In some embodiments, the amplifier with an input capacitively coupled in feedback is used to operate the sensor from the converter input voltage by using a high PSRR (power supply rejection ratio) regulator to generate a local and clean supply voltage that reduces the perturbation to the power system in the switch area. In some embodiments, a design variation is used to sample the difference between the input voltage and the controller supply and reproduce the difference in drain voltage between the power switch and the replica switch. This enables the sensor to not be exposed to the supply voltage. In some embodiments, the amplifier with an input capacitively coupled in feedback is used to compensate for the power supply network related (PDN related) variations of the input voltage during current sensing.
[0120] Some embodiments use three components to adjust the peak power of the SoC 5501 based on the state of the USB TYPE-C device 5529. Those components include an OS peak power manager (part of the OS 5552), a USB TYPE-C connector manager (part of the OS 5552), and a USB TYPE-C protocol device driver (e.g., one of the drivers 5554a, 5554b, 5554c). In some embodiments, when a USB TYPE-C power sink device is detached from the SoC 5501, the USB TYPE-C connector manager sends a synchronization request to the OS peak power manager, and when the power sink transitions the device state, the USB TYPE-C protocol device driver sends a synchronization request to the peak power manager. In some embodiments, the peak power manager receives a power budget from the CPU when the USB TYPE-C connector is attached to a power sink and is active (e.g., in a high-power device state). In some embodiments, the peak power manager returns the power budget to the CPU for performance when the USB TYPE-C connector is removed, or when the USB TYPE-C connector is attached and the power sink device is idle (in the lowest device state).
[0121] In some embodiments, logic is provided for dynamically selecting the best operating processing core for the BIOS power-up flow and sleep exit flow (e.g., S3, S4, and / or S5). The selection of the bootstrap processor (BSP) is shifted to the early power-up time in place of a fixed hardware selection at any point in time. For maximum boot performance, this logic selects the core that can be the fastest as the BSP at the early power-up time. Also, for maximum power savings, this logic selects the most power-efficient core as the BSP. The processor or switching for selecting the BSP occurs during the boot-up and power-up flows (e.g., the S3, S4, and / or S5 flows).
[0122] In some embodiments, the memory here is organized in a multi-level memory architecture and their performance is governed by a decentralized scheme. This decentralized scheme includes p-units 5510 and a memory controller. In some embodiments, this scheme dynamically balances multiple parameters such as power, heat, cost, latency, and performance for memory levels that gradually move away from the processor within the platform or device 5500 based on how the application uses memory levels far from the processor core. In some examples, the determination of the state of the far memory (FM) is decentralized. For example, the processor power management unit (p-unit), the near memory controller (NMC), and / or the far memory host controller (FMHC) make determinations about the power and / or performance state of the FM at their respective levels. Coordinating these determinations provides an optimal power and / or performance state of the FM at a given point in time. The power and / or performance state of the memory adaptively changes in accordance with the changing workload and other parameters even when the (one or more) processors are in a particular power state.
[0123] In some embodiments, a processor power state policy (e.g., a policy for C states) is implemented that provides optimal power state selection by taking into account the performance and / or responsiveness needs of threads expected to be scheduled to cores entering the idle state in order to achieve improved instructions per cycle (IPC) and performance for cores executing user-critical tasks. This scheme provides the ability to deliver responsiveness gains for important and / or user-critical threads running on a system-on-chip. A p unit 5510 coupled to a plurality of processing cores receives from the operating system 5552 hints indicating a bias to a power state or a performance state for at least one of the plurality of processing cores based on the priority of the threads in a context switch.
[0124] In some embodiments, a core ranking scheme based on processor core energy efficiency is used. This is similar to the favored core in a multi-core processor system. However, the favored core here (e.g., one of 5508) is a core with good energy efficiency that enables the SoC to use a core with the lowest V min having such a V min value can be fuse-programmed to appropriate registers or stored in NVM during a high-volume manufacturing (HVM) process. In some embodiments, an operating system (OS) scheduler uses the core ranking information to schedule specific applications to cores with the lowest V min to achieve optimal energy performance.
[0125] In some embodiments, the bootstrap flow identifies the bootstrap processor core (BSP) as the most energy-efficient core of the SoC and assigns it the lowest APIC (Advanced processor interrupt controller) ID (identification) value. The APIC ID value defines the target processor that receives interrupts delivered in logical destination mode in the local x2APIC. This is a 32-bit value initialized by hardware. The APIC ID is a term in the Intel architecture, but similar functionality in other processor architectures can also be used to identify the bootstrap processor. Initially, the BSP is core 0 5508a or any core of the multi-core system. In a set of heterogeneous cores having large cores (or complex and / or high-power applications) and small cores (less complex and / or low-power applications), the BSP can be one of the small cores. This initial BSP is used to identify the EE BSP.
[0126] In this context, microcode (e.g., pCode) or BIOS5520 (basic input / output system) reads fuses or NVM that store the V value per core. These fuses or NVM are programmed in the HVM. After reading the fuses or NVM, the microcode or BIOS5520 calculates and ranks the cores. Thus, the microcode (e.g., pCode) or BIOS5520 calculates and ranks the core APIC ID based on the efficiency near the LFM (low-frequency mode) frequency. Based on the calculated and ranked cores, the microcode or BIOS5520 transfers BSP ownership to the most efficient core (e.g., having a higher ranking number) by setting a register (e.g., IA32_APIC_BASE.BSP = 1). min
[0127] In some embodiments, the microcode or BIOS 5520 then shares the APIC IDs of the cores ranked based on efficiency with the operating system (or kernel). For example, the microcode or BIOS shares the APIC IDs of the cores ranked based on efficiency with the OS via an ACPI (Advanced Configuration and Power Interface) table such as, for example, the MADT (Multiple APIC Description Table).
[0128] Using the reordered APIC IDs, the OS efficiently services low-utilization tasks, interrupts, and DPCs on the core with the lowest Vmin. Note that the reordering of the APIC IDs also enables more efficient HW interrupt routing. In some embodiments, the OS scheduler uses the efficient cores as preferred or favored cores for thread scheduling. When the SoC supports dynamic hardware (HW) feedback, the pCode (or suitable microcode or firmware) shares the updated efficient core ranking via a shared memory or MSR (model-specific register) interface.
[0129] References to "an embodiment", "one embodiment", "some embodiments", or "other embodiments" in the specification mean that the particular mechanism, structure, or feature described in connection with those embodiments is not necessarily present in all embodiments, but is included in at least some embodiments. The various appearances of "an embodiment", "one embodiment", or "some embodiments" do not necessarily refer to the same embodiment. When the specification states that a component, mechanism, structure, or feature "may be" included, "can be", or "obtains", that particular component, mechanism, structure, or feature need not be included. When the specification or claims refer to an element with "a" or "an", it does not mean that there is only one of those elements. When the specification or claims refer to an element with "an additional", it does not exclude the possibility that there are two or more of that additional element.
[0130] Throughout the specification and in the claims, the term "connected" means a direct connection, such as an electrical, mechanical, or magnetic connection between the connected items, without an intermediate device.
[0131] The term "coupled" means a direct or indirect connection, such as a direct electrical, mechanical, or magnetic connection between the connected items, or an indirect connection through one or more passive or active intermediate devices.
[0132] The term "adjacent" generally refers here to a position where one thing is next to another (e.g., immediately next to, or close with one or more things placed between them) or in contact at the boundary (e.g., touching).
[0133] The term "circuit" or "module" can refer to one or more passive and / or active components configured to cooperate with each other to provide a desired function.
[0134] The term "signal" can refer to at least one current signal, voltage signal, magnetic signal, or data / clock signal. The meanings of "a", "an", and "the" include plural references. The meaning of "in" includes "in" and "on".
[0135] The term "analog signal" generally refers here to a continuous signal where the time-varying characteristic (variable) of the signal represents some other time-varying quantity, i.e., is similar to another time-varying signal.
[0136] The term "digital signal" is a physical signal that represents a series of discrete values (quantized discrete-time signals), such as any bitstream or a digitized (sampled and analog-to-digital converted) analog signal.
[0137] The term "scaling" generally refers to converting a design (circuit diagram and layout) from one process technology to another, and can later be reduced in layout area. In some cases, scaling can also enlarge a design from one process technology to another, and can later be increased in layout area. The term "scaling" generally also refers to reducing or enlarging the layout and devices within the same technology node. The term "scaling" can also refer to adjusting the signal frequency relative to another parameter, such as the power level (e.g., slowing down or speeding up, i.e., scaling down or scaling up, respectively).
[0138] The terms "substantially", "near", "approximately", "almost", and "about" generally refer to being within ±10% of the target value.
[0139] Unless otherwise stated, the use of ordinal adjectives such as "first", "second", "third", etc. to describe common objects merely indicates that different instances of similar objects are being referred to, and there is no intention that the objects so described must be in a given sequence, whether temporally, spatially, in terms of ranking, or in any other way.
[0140] For the purposes of this disclosure, the phrases "A and / or B" and "A or B" mean (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C).
[0141] In the specification and claims, where present, terms such as "left", "right", "front", "rear", "top", "bottom", "above", "below", and the like are used for purposes of explanation and are not necessarily used to describe permanent relative positions.
[0142] It should be noted that elements of a figure having the same reference numeral (or name) as elements of other figures can operate or function in a manner similar to that described, but are not so limited.
[0143] For the purposes of the embodiments, the transistors within the various circuits and logic blocks described herein are metal oxide semiconductor (MOS) transistors or derivatives thereof, and the MOS transistors include drain, source, gate, and bulk terminals. The transistors and / or MOS transistor derivatives also include trigate and FinFET transistors, gate all around cylindrical transistors, tunneling FETs (TFETs), square wires, or rectangular ribbon-shaped transistors, ferroelectric FETs (FeFETs), or other devices that implement transistor functionality such as carbon nanotubes or spintronic devices. The symmetric source and drain terminals of the MOSFETs are, in other words, equivalent terminals and are used interchangeably herein. On the other hand, TFET devices have asymmetric source and drain terminals. As will be appreciated by those skilled in the art, other transistors such as, for example, bipolar junction transistors (BJT PNP / NPN), BiCMOS, CMOS, etc. may also be used without departing from the scope of the disclosure.
[0144] Also, certain mechanisms, structures, functions, or features may be combined as suitable in one or more embodiments. For example, a first embodiment may be combined with a second embodiment if the specific mechanisms, structures, functions, or features associated with these two embodiments are not mutually exclusive.
[0145] While the disclosure has been described with respect to specific embodiments thereof, many modifications, changes, and variations of these embodiments will become apparent to those skilled in the art in light of the above description. The embodiments of the disclosure are intended to embrace all such modifications, changes, and variations that fall within the broad scope of the appended claims.
[0146] In addition, for simplicity of illustration and description, and to avoid obscuring the disclosure, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the presented figures. Also, the configuration may be shown in the form of a block diagram, which is for the purpose of avoiding obscuring the disclosure, and also in view of the fact that the details regarding the implementation of such a block diagram configuration highly depend on the platform on which the present disclosure will be implemented (i.e., such details should be well within the scope of those skilled in the art). When specific details (e.g., circuits) are described to describe an exemplary embodiment of the disclosure, it should be apparent to those skilled in the art that what is disclosed can be implemented without using those specific details or using modifications thereof. This description is, therefore, to be regarded as illustrative rather than restrictive.
[0147] The following examples relate to further embodiments. The specific details in these examples may be used anywhere in one or more embodiments. All of the optional features of the apparatus described herein may also be used with respect to a method or process. These examples can be combined in any combination. For example, Example 4 can be combined with Example 2.
[0148] Example 1: A machine-readable storage medium having machine-readable instructions that, when executed, cause one or more processors to read a fuse or NVM storing the minimum operating voltage (Vmin) for each processor core in a multi-core system, rank each processor core of the multi-core system from the highest rank to the lowest rank according to the Vmin for each processor core, and change an existing bootstrap processor by assigning a new bootstrap processor to the processor core having the highest rank.
[0149] Example 2: The machine-readable storage medium of Example 1, wherein an operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank.
[0150] Example 3: The machine-readable storage medium of Example 1, wherein the ranked processor cores have ranked identifiers.
[0151] Example 4: The machine-readable storage medium of Example 3, having machine-readable instructions that, when executed, cause one or more processors to share a ranked identifier of the processor cores with an operating system.
[0152] Example 5: The machine-readable storage medium of Example 4, wherein the sharing of the ranked identifier is via an Advanced Configuration and Power Interface table.
[0153] Example 6: The machine-readable storage medium of Example 5, wherein the Advanced Configuration and Power Interface table has a Multiple APIC Description table.
[0154] Example 7: The machine-readable storage medium of Example 1, wherein ranking each processor core of the multi-core system is further based on the efficiency of the processor core near a low-frequency mode frequency.
[0155] Example 8: The machine-readable storage medium of Example 1, wherein the fuse or NVM is programmed at Vmin for each processor core during mass production.
[0156] Example 9: A system-on-chip having a plurality of processor cores, a fuse or NVM storing Vmin for each processor core among the plurality of processor cores, and a power management unit coupled to the plurality of processor cores, wherein the power management unit ranks each processor core according to the Vmin for each processor core by executing firmware, and assigns a bootstrap processor to the processor core having the highest rank.
[0157] Example 10: The system-on-chip of Example 9, wherein an operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank.
[0158] Example 11: The system-on-chip of Example 9, wherein the firmware assigns ranked processor core identifiers.
[0159] Example 12: The system-on-chip of Example 11, wherein the firmware shares the ranked identifier of the processor core with an operating system.
[0160] Example 13: The system-on-chip of Example 12, wherein the firmware shares the ranked identifier via an Advanced Configuration and Power Interface table.
[0161] Example 14: The system-on-chip of Example 13, wherein the Advanced Configuration and Power Interface table has a Multiple APIC Description table.
[0162] Example 15: The system-on-chip of Example 9, wherein the firmware ranks each processor core based on the efficiency of the processor core near a low-frequency mode frequency.
[0163] Example 16: A system having a memory, a processor coupled to the memory, and a wireless interface that enables the processor to communicate with other devices, wherein the processor includes a plurality of processor cores including heterogeneous process cores, a fuse or NVM that stores a Vmin for each processor core among the plurality of processor cores, and a power management unit coupled to the plurality of processor cores, and by executing firmware, ranks each processor core with an identifier according to the Vmin for each processor core, and transfers ownership of the bootstrap processor to the processor core having the highest rank.
[0164] Example 17: The system of Example 16, wherein the operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank.
[0165] Example 18: The system of Example 16, wherein the firmware shares the ranked identifier of the processor core with the operating system.
[0166] Example 19: The system of Example 16, wherein the firmware ranks each processor core based on the efficiency of the processor core near the low-frequency mode frequency.
[0167] Example 20: The system of Example 16, wherein the firmware shares the ranked identifier via an Advanced Configuration and Power Interface table having a Multiple APIC Description table.
[0168] An abstract is provided to enable the reader to confirm the nature and gist of the technical disclosure. The abstract is submitted on the understanding that it will not be used to limit the scope or meaning of the claims. The following claims are incorporated into the detailed description, and each claim stands on its own as a separate embodiment.
Claims
1. A machine-readable storage medium having machine-readable instructions that, when executed, cause one or more processors to read a minimum operating voltage (Vmin) for each processor core in a multi-core system, rank each processor core of the multi-core system from highest rank to lowest rank according to the Vmin for each processor core, change an existing bootstrap processor by assigning a new bootstrap processor to the processor core having the highest rank, and execute a method having the above.
2. The machine-readable storage medium according to claim 1, wherein an operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank.
3. The machine-readable storage medium according to claim 1 or 2, wherein the ranked processor cores have ranked identifiers.
4. Machine-readable instructions that, when executed, cause the one or more processors to share the ranked identifier of the processor core with an operating system, are included in the machine-readable storage medium according to claim 3.
5. The machine-readable storage medium according to claim 4, wherein the sharing of the ranked identifier is via an Advanced Configuration and Power Interface table.
6. The machine-readable storage medium according to claim 5, wherein the Advanced Configuration and Power Interface table has a Multiple APIC Description table.
7. Ranking each processor core of the multi-core system is further based on the efficiency of the processor core near the low-frequency mode frequency, according to any one of claims 1 to 6.
8. The machine-readable storage medium according to any one of claims 1 to 7, wherein the Vmin for each processor core is stored in a non-volatile memory programmed with the Vmin for each processor core during mass production.
9. A plurality of processor cores, and A memory storing the minimum operating voltage (Vmin) for each of the plurality of processor cores; A power management unit coupled to the plurality of processor cores; comprising: The power management unit ranks each processor core according to the Vmin for each processor core by executing firmware; assigns a bootstrap processor to the processor core with the highest rank; a system-on-chip.
10. The system-on-chip according to claim 9, wherein the operating system schedules interrupts or services of low-utilization tasks on the processor core having the highest rank.
11. The system-on-chip according to claim 9 or 10, wherein the firmware assigns ranked processor core identifiers.
12. The system-on-chip according to claim 11, wherein the firmware shares the ranked identifier of the processor core with the operating system.
13. The system-on-chip according to claim 12, wherein the firmware shares the ranked identifier via an Advanced Configuration and Power Interface table.
14. The system-on-chip according to claim 13, wherein the Advanced Configuration and Power Interface table has a Multiple APIC Description table.
15. The system-on-chip according to any one of claims 9 to 14, wherein the firmware ranks each processor core based on the efficiency of the processor core near the low-frequency mode frequency.
16. A memory; A processor coupled to the memory; A wireless interface enabling the processor to communicate with other devices; comprising: The processor includes: A plurality of processor cores including heterogeneous process cores; One or more memories storing the minimum operating voltage (Vmin) for each of the plurality of processor cores; A power management unit coupled to the plurality of processor cores, which ranks each processor core with an identifier according to the Vmin for each processor core by executing firmware; Transfer the ownership of the bootstrap processor to the processor core with the highest rank, a power management unit, including a system.
17. The system according to claim 16, wherein the operating system schedules interrupts or services for low-utilization tasks on the processor core with the highest rank.
18. The system according to claim 16 or 17, wherein the firmware shares the ranked identifier of the processor core with the operating system.
19. The system according to any one of claims 16 to 18, wherein the firmware ranks each processor core based on the efficiency of the processor core near the low-frequency mode frequency.
20. The system according to any one of claims 16 to 19, wherein the firmware shares the ranked identifier via an Advanced Configuration and Power Interface table having a Multiple APIC Description table.
Citation Information
Patent Citations
Standby power control for low power devices
JP2008503835A
Method for booting heterogeneous system and presenting symmetric core view
JP2014225242A
Integrated circuit device, asymmetric multi-core processing module, electronic device and method of managing execution of computer program code therefor
US20140325183A1
Method and system for booting a multiprocessor computer
US6584560B1
Voltage selectors coupled to processor cores
WO2015183268A1