Power balancing among individual components in a compute node

A CPU-based power management system dynamically adjusts component frequencies using 'scaling weights' and 'power violation indices' to optimize power usage in compute nodes, addressing inefficiencies in existing systems and reducing power supply requirements.

JP7823291B2Active Publication Date: 2026-03-04INTEL CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022009550
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-03-03
Filing Date
2022-01-25
Publication Date
2026-03-04
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

Existing power management systems in compute nodes with heterogeneous components are inefficient due to slow response times, leading to over-engineered power supply units that cannot adapt quickly to varying workloads, resulting in excessive power consumption and cost.

Method used

Implementing a power balancing scheme with a CPU-based power management system that dynamically adjusts the operating frequencies of components using 'scaling weights' and 'power violation indices' to optimize PSU sizing and reduce power consumption.

Benefits of technology

The solution enables faster power balancing, up to 10 times quicker than legacy systems, allowing for more efficient use of power resources and reducing the need for oversized power supplies, thus minimizing performance impact and costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007823291000005
    Figure 0007823291000005
  • Figure 0007823291000006
    Figure 0007823291000006
  • Figure 0007823291000007
    Figure 0007823291000007
Patent Text Reader

Abstract

To provide methods for balancing power on a compute platform including one or more processing units and a plurality of variable frequency components, and a platform.SOLUTION: Methods for balancing power between discrete components, such as compute nodes, processing units (CPUs) and accelerators comprises: detecting a situation where a threshold (power supply capacity threshold) of power consumption of a compute platform is exceeded; adjusting operating frequencies of a processing unit and / or other platform components such as accelerators so as to be below the threshold; and providing power limit biasing hints (scaling weights) to platform components along with a power violation index used to adjust the operating frequencies. Optionally, a processing unit calculates the power violation index and the scaling weights and directly control the frequencies of itself and platform components.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] To solve next-generation machine learning and high-performance computing problems, the industry is working on heterogeneous computing using accelerators and processing units such as central processing units (CPUs). To achieve maximum performance, the thermal design power (TDP) of both CPUs and accelerators is increasing. Depending on the CPU-to-accelerator ratio, the overall power of the compute node increases as a result. This means that the node's power supply unit (PSU) should be sized to support the sum of the TDPs from all components on the platform. However, depending on workload behavior, not all platform components consume their TDP simultaneously. Some workloads are either CPU-centric or accelerator-centric in terms of their time distribution.

[0002] Until now, power management in compute nodes with varying degrees of functionality has been accomplished using a baseboard management controller (BMC), a node manager (NM), or a data center manager (DCM). In all of these cases, the response time for power changes / balancing is in the 100ms range, which is too slow to allow the platform to downsize the PSU. Therefore, the PSU must be designed to handle the maximum TDP of all components in the node. This is expensive and "over-engineered" in practice, since most of the time, not all components will be operating at maximum power simultaneously. [Brief explanation of the drawings]

[0003]

[0013] Many of the foregoing aspects and attendant advantages of this invention will become more readily appreciated as they become better understood by reference to the following detailed description when taken in conjunction with the accompanying drawings, wherein, unless otherwise specified, like reference numerals refer to like parts throughout the various views.

[0004] [Figure 1] FIG. 1 is a block diagram of a platform used to outline the power balancing approach disclosed herein.

[0005] [Figure 2] 1 is a flowchart illustrating an example of a power balancing implementation according to one embodiment.

[0006] [Figure 3] 1 is a schematic diagram of a platform with an exemplary configuration including two CPUs and six accelerators.

[0007] [Figure 4] FIG. 1 is a schematic diagram of a dual socket platform configured to implement a balanced power management scheme in accordance with aspects of the technology disclosed herein.

[0008] [Figure 5] 1 is a graph showing changes in power levels of the GPU, CPU, and power supply.

[0009] [Figure 6] 1 is a table showing various CPU and GPU power configurations.

[0010] [Figure 7] FIG. 1 is a diagram of a computing node that may be implemented using aspects of the embodiments described and illustrated herein. DETAILED DESCRIPTION OF THE INVENTION

[0011] Embodiments of a method and apparatus for balancing power among individual components in a computing node are described herein. In the following description, numerous specific details are set forth to provide a thorough understanding of embodiments of the invention. However, one skilled in the art will recognize that the invention can be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the invention.

[0012] References throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places throughout this specification do not necessarily all refer to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0013] For clarity, individual components in the figures herein may be referred to by their labels in the figures rather than by specific reference numbers. Additionally, reference numbers referring to a particular type of component (as opposed to a particular component) may be indicated by the reference number followed by "(typ)," meaning "typical." It will be understood that these configurations of components are typical of similar components that may be present but that, for brevity and clarity, are not shown in the figures or that are not otherwise labeled with separate reference numbers. Conversely, "(typ)" should not be interpreted to mean that the component, element, etc. is commonly used for its disclosure, function, implementation, purpose, etc.

[0014] According to aspects of the embodiments disclosed herein, methods and apparatus are provided for power balancing between individual components, such as CPUs and accelerators, in a compute node. In one aspect, the techniques enable optimizing PSU sizing to support higher TDPs of CPUs / accelerators and implementing a power balancing scheme with faster throughput. The solutions disclosed herein are not limited to CPUs and accelerators, but may be applied to any individual component on a platform, including memory, fans, FPGAs, CXL / PCIe devices, or other devices accessed via fabric or peripheral buses / interconnects, etc.

[0015] 1 shows an exemplary set of components of a platform 100 used to outline a power balancing approach. Platform 100 includes a power supply unit (PSU) 102, a power monitor sensor 104, a CPU 106 (also labeled CPU 0), and seven variable frequency components 108, 110, 112, 114, 116, 118, and 120 (also labeled components 0-6). Software 122 runs on CPU 106. In general, variable frequency components 108, 110, 112, 114, 116, 118, and 120 may comprise CPUs, accelerators, and other processing units (collectively referred to as XPUs) capable of adjusting their operating frequency to increase or reduce their power consumption, including one or more of a graphics processor unit (GPU) or general-purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an infrastructure processing unit (IPU), an artificial intelligence (AI) processor or AI inference unit, and / or other accelerators, FPGAs and / or other programmable logic (used for computational purposes), network processors, etc. In addition to the components shown, as will be recognized by those skilled in the art, platform 100 will include additional components such as memory, peripherals, storage, etc.

[0016] PSU 102 is used to provide power to the platform components. Power monitor sensor 104 is used to monitor the overall platform power consumption (or) power drawn from PSU 102 and provide that information to one CPU of the platform (CPU0 in this example). Software 122 is configured to dynamically or statically provide / adjust bias hints that enable prioritization among platform components when power reduction is required. The platform software may use either in-band or out-of-band mechanisms to communicate parameters to applicable platform components.

[0017] In general, power consumption is directly proportional to the operating frequency of platform components. A common scheme for adjusting a platform's power constraints is to have the CPU(s) and other components, such as accelerators, modulate their frequency.

[0018] [Power limit violation determination]

[0019] Under one embodiment, determining a power limit violation proceeds as follows: CPU0 continuously compares the platform's power consumption with the PSU functional limits and generates a "power violation index" comprising a number from 0 to 1, where 0 means the power limit is not violated. Based on the magnitude of the power limit violation, the "power violation index" is correspondingly increased. The "power violation index" is communicated to platform components, and the platform components are configured to alter their operating frequencies in consideration of the "power violation index" value they receive.

[0020] 2 shows a flowchart 200 illustrating an example of a power balancing implementation according to one embodiment. Communication occurs between software 122, power monitor sensor 104, CPU0 (106), and components 108 and 110. As shown in block 202, based on workload behavior, software 122 determines a scaling weight for each component. Software 122 sends scaling weights 204, 206, and 208 to CPU0, component 108, and component 110, respectively. As shown by block 210, power monitor sensor 104 senses the power consumption of the platform components (collectively) or the power drawn from PSU 102 and communicates it to CPU0, as shown by power consumption signal 212.

[0021] In block 214, CPU0 compares the platform power consumption with the PSU functional limits to generate a "power violation index." Corresponding "power violation index" signals or messages 216, 218, and 220 are sent from CPU0 to itself, component 108, and component 110, respectively.

[0022] As shown in block 222, CPU0 reduces its maximum frequency of operation using the "power violation index" and the "scaling weight" that it calculates and / or receives. (Note that in the case of a single-CPU platform, software 122 runs on CPU0.) Similarly, as shown in blocks 224 and 226, components 108 and 110 reduce their maximum frequency of operation using the "power violation index" and the "scaling weight" that they each receive.

[0023] [Power management using CPU0 vs. legacy BMC technology]

[0024] As outlined herein, power management using the embodiments disclosed herein is up to 10 times faster than legacy BMC technology. In general, the speed improvement is primarily due to two considerations: 1) the CPU's processing speed is much faster than the BMC; and 2) there is no latency associated with transferring platform control data from the BMC to the CPU. In addition, the power management scheme gives the user a choice of which components (CPU, accelerator, or other components) to throttle and to what extent. The following sections provide details on these "bias hints" for power management.

[0025] [Bias hint for prioritizing "throttle" between different components]

[0026] The software recognizes the characteristics of the workload and provides a bias hint for each component named "scaling weight," which in one embodiment is determined using the following formula: Component scaling weight = (Percentage of expected frequency degradation) / (Expected frequency degradation if all components were given equal priority) This means that if one component needs to be prioritized higher than another, it has a lower scaling weight compared to the other component.

[0027] Also, in some embodiments, the software may instruct any component to ignore the "power violation index," allowing some components to be prioritized so that the power provided to those components is not throttled.

[0028] [Frequency reduction by each component to reduce power consumption]

[0029] In some embodiments, each component uses the "power violation exponent" and the "scaling weight" to reduce its maximum frequency of operation unless it is configured to ignore the "power violation exponent." In one embodiment, the following formula is used: Frequency reduction = (maximum frequency of component when not throttled) * (power violation exponent) * (scaling weight)

[0030] [Example scenario (platform setup)]

[0031] In the following scenario, a platform 300 having an exemplary configuration including two CPUs 306 and 308 (CPU0 and CPU1) and six accelerators 310, 312, 314, 316, 318, and 320 (ACC0-5), as shown in Figure 3, is used. The maximum frequency of the CPUs is 4 GHz, while the maximum frequency of the accelerators is 2 GHz. In one example, assume that the "power violation index" determined by CPU0 is 0.2.

[0032] [Scenario 1 (all components are prioritized equally)]

[0033] Under the first scenario, all components are prioritized equally. In this case, the scaling weight of all components is 1. The accelerator is a GPU. Therefore, The maximum CPU frequency is reduced to 4Ghz*(1-0.2)=3.2GHz. The maximum GPU frequency is reduced to 2GHz*(1-0.2)=1.6GHz.

[0034] [Scenario 2 (CPU is prioritized higher than GPU)]

[0035] Under the second scenario, all power reduction must come from the accelerator. In this case, the CPU is configured to ignore the "power violation exponent." Table 1 shows the "scaling weights" that must be programmed to calculate the maximum frequency of operation for each component. [Table 1] [Table 1]

[0036] [Scenario 3 (GPU is prioritized higher than CPU)]

[0037] Under the third scenario, all power reduction must come from the CPU. To achieve this result, the accelerator is configured to ignore the "power violation exponent." Table 2 shows the "scaling weights" that are programmed to calculate the maximum frequency of operation for each component. [Table 2] [Table 2]

[0038] [Scenario 4 (CPU is prioritized higher than GPU, but not completely higher than GPU)]

[0039] Under the fourth scenario, the frequency reduction from the CPU should be 20%, and the frequency reduction from the accelerator should be 80%. Table 3 shows the “scaling weight” values ​​to be programmed to calculate the maximum frequency of operation for each component. [Table 3] [Table 3]

[0040] [Scenario 5 (All components are prioritized differently)]

[0041] Under the fifth scenario, components are prioritized at different levels. Table 4 shows the predicted frequency reduction and the "scaling weight" to be programmed to calculate the maximum frequency of operation for each component. [Table 4] [Table 4]

[0042] [Multi-socket platform]

[0043] 4 illustrates a dual-socket platform 400 configured to employ the techniques disclosed herein to implement a balanced power management scheme. The power components of platform 400 include a PSU 402, a power monitor sensor 404, voltage regulators 406 and 408, and an optional other voltage regulator (VR) 409. Voltage regulators 406 and 408 provide power to components in socket 0 and socket 1, respectively. Socket 0 includes a system-on-chip (SoC) CPU 410 (SoC0) and GPUs 412 and 414 (GPU0 and GPU1) with respective registers 413 and 415. Socket 1 includes an SoC CPU 416 (SoC1) and GPUs 418 and 420 (GPU2 and GPU3) with respective registers 419 and 421. SoC CPU 410 includes power control unit 422 (PCU0) and root ports (RP) 423 and 425, while SoC CPU 416 includes PCU 424 (PCU1) and root ports 427 and 429. Fan 426 is used to cool the SoC CPUs and GPUs in socket 0 and socket 1, as well as other platform components.

[0044] In addition to the power and socket components, platform 400 includes memory 428 coupled to one or more memory controllers on SoC0 and memory 430 coupled to one or more memory controllers on SoC1. The platform's firmware (also known as BIOS) is stored in firmware storage device 432. A network interface controller (NIC) or host fabric interface (HFI) 434 is coupled to a network or fabric 436. Storage 438 represents one or more storage devices, such as a solid-state drive (SSD), or other type of non-volatile storage device. Block 440 represents other components on platform 400, such as a BMC, zero or more peripherals, etc. In general, all or a portion of software 442 may be stored in storage 438 or loaded via network / fabric 436.

[0045] Power monitor sensor 404 is configured to sense the power drawn from PSU 402 (e.g., using a sensing resistor or other well-known mechanism) and send a corresponding analog signal 444 to voltage regulator 406. In the illustrated embodiment, voltage regulator 406 is connected to PCU 422 via a Serial Voltage Identification (SVID) interface 446 that is used to send a digital power consumption value to PCU 422 indicative of the power level being consumed by the platform. Optionally, an overcurrent protection signal may be sent from voltage regulator 406 to PCU 422 via SVID interface 446.

[0046] In one embodiment, platform BIOS / firmware is used to enable power balancing. For example, in some embodiments, platform 400 includes Unified Extensible Firmware Interface (UEFI) firmware that includes a UEFI component that enables power balancing and performs associated platform configuration operations.

[0047] In one embodiment of a multi-socket platform, one of the SoC CPUs is used to manage frequency control for components in its own socket and the other sockets. In platform 400, this is done by SoC0. PCU0 calculates the frequency limits and sends associated control signals (e.g., messages) to PCU1 via socket-to-socket interconnect 448. The control signals / messages can be used to implement two schemes. Under the first scheme, PCU0 sends both "power violation index" data and "scaling weights" to PCU1 to be implemented for each of the components on SoC1 (e.g., CPU / SoC and GPUs 2 and 3 of sockets). The frequency of memory 430 may also be adjusted. The "power violation index" data and applicable "scaling weights" are provided to each component used for power balancing in socket 1, and these components adjust their maximum operating frequency.

[0048] Under the second scheme, PCU0 sends only the "power violation index" to PCU1. The CPU of SoC1 determines the applicable "scaling weight" through the execution of its software workload.

[0049] Each of PCU0 and PCU1 is used to set the frequency of their respective SoC CPU, and optionally, the frequency of the GPU in their socket. In one embodiment, this is done by sending a SET_SLOT_POWER_LIMIT message to set a power limit value in a register on the GPU that the GPU reads to set its power limit. For example, as shown in FIG. 4, PCU1 sends a SET_SLOT_POWER_LIMIT message to set the GPU power limit for GPU2 in register 419. In one embodiment, the root ports on SoC0 and SoC1 coupled to the GPUs include a "slot capability" to which the SET_SLOT_POWER_LIMIT message is sent. The root ports then send a corresponding SET_POWER_LIMIT message to the SoC CPU in that socket, which is to be written to registers 412, 415, 419, and 421 on GPUs 0, 1, 2, and 3, over the PCIe link connecting each GPU.

[0050] Under an alternative approach, the SoC CPU of a socket provides a "power violation exponent" and a "scaling weight" to the GPU in that socket, which then calculates the change to be made in frequency.

[0051] 5 shows a graph 500 illustrating power level changes of GPU and CPU power during runtime operation of platform 400. As shown, during time frame 502, the combined power consumption of GPU power and CPU power (power of the power supply) exceeds the power supply capacity. This condition is detected, and corresponding control signals are provided to the GPU and CPU to reduce their frequency and thus return the power of the power supply to a power level below the power supply capacity.

[0052] In the preceding Tables 1-4, for simplicity, only the frequency adjustments of the CPU and accelerators are shown. In practice, as shown in Table 600 of FIG. 6, the power consumption of the platform is used to determine when power balancing operations are performed. In addition to the power consumption of the CPU and GPU, this also includes the power consumed by memory, which typically comprises several DIMMs (dual inline memory modules), the power consumption of the rest of the platform, and power supply losses. In this example, there are two CPUs per node, four GPUs per node, and 16 DIMMs per node.

[0053] Under the first power configuration 602, the two CPUs are operated at their 350 watt TDP, while the four GPUs are operated at 400 watts. The node's power consumption is 2897 watts, which is below the 3000 watt power supply limit. As a result, no throttling is performed. Under the second power configuration 604, the CPUs are operated at 150 watts and the GPUs are operated at their 500 watt TDP. The node's power consumption is 2897 watts, and no throttling is performed.

[0054] Under the third power configuration 606, the CPU operates at its TDP (350 watts) and the GPU operates at its TDP (500 watts), resulting in a node power consumption of 3365 watts, which exceeds the 3000 watt power supply limit and therefore requires power balancing.

[0055] Under the first balanced power configuration 608, GPU power is reduced to 420 watts, resulting in a node power consumption of 2991 watts. Under the second balanced power configuration 610, CPU power is reduced to 190 watts, resulting in a node power consumption of 2991 watts. Under the third balanced power configuration 612, CPU power is reduced to 308 watts and GPU power is reduced to 440 watts, resulting in a node power consumption of 2986 watts.

[0056] The SoC or CPU monitors the power consumption of the platform and periodically updates a "power violation index" for all platform components. In response to receiving a "power violation index" signal or value, the component adjusts its frequency and changes its power consumption. Optionally, the SoC or CPU of a socket may calculate maximum frequency values ​​for all or some of the power-managed components in the socket and send or write the maximum frequency values ​​to those components. This helps meet target power limit requirements with shorter response times. Additionally, depending on workload behavior, software may dynamically adjust "scaling weights" for power reduction for each component, thereby achieving maximum node performance. Dynamic adjustments to scaling parameters may be supplied by the workload as hints in the software, or they may be generated based on a learning algorithm configured to better understand workload phases and adjust scale factors. For example, such a learning algorithm may employ a machine learning (ML) algorithm, such as an ML algorithm employing reinforcement learning. Other types of ML algorithms may also be used.

[0057] The techniques described and illustrated for the dual-socket platform of FIG. 4 can be extended to a multi-socket platform having more than two sockets. In this case, SoC0 or CPU0 transmits a "power violation index" to each of the other sockets (e.g., to the PCU within each socket) via a socket-to-socket link in a manner similar to that shown in FIG. 4. As before, either SoC0 or CPU0 can determine the scaling weights, or these scaling weights can be determined by the CPU running the software workload for each socket. Information when scaling weights are determined by the socket CPU

[0058] As mentioned above, existing power management schemes, primarily BMCs at the node level, have slow response times (~100 ms) for balancing power between components. This was fine when the primary computing element in a node was the CPU and had minimal impact on performance. As more and more workloads utilize accelerators in addition to the CPU, the response time for power balancing by the BMC becomes too slow and impacts the overall performance of the node. Conversely, a power management scheme implemented within CPU0 to manage power between the CPU and accelerators with a much shorter response time (up to 10x faster) would offer a significant improvement over current schemes for board- and / or rack-level power management, while minimizing the impact on performance.

[0059] As described above, the operating frequencies of various components are adjusted to reduce power. The adjustments to the operating frequencies are made between non-zero values. Shutting down a component does not fall within the scope of adjusting its operating frequency, as, among other things, the component will no longer be operational.

[0060] [Example Platform / Compute Node]

[0061] 7 illustrates a platform comprising a computing node 700 in which aspects of the above-disclosed embodiments may be implemented. The computing node 700 includes one or more processors 710 that provide processing, operational management, and instruction execution for the computing node 700. The processor 710 may include any type of microprocessor, central processing unit (CPU), graphics processing unit (GPU), processing core, multi-core processor, or other processing hardware, or combination of processors, for providing processing for the computing node 700. The processor 710 controls the overall operation of the computing node 700 and may be or include one or more programmable general-purpose or special-purpose microprocessors, digital signal processors (DSPs), programmable controllers, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), etc., or a combination of such devices.

[0062] In one example, compute node 700 includes an interface 712 coupled to processor 710, which may represent a higher speed or high throughput interface to system components requiring a higher bandwidth connection, such as memory subsystem 720 or optional graphics interface component 740 or optional accelerator 742. Interface 712 represents interface circuitry that may be a standalone component or integrated on the processor die. If present, graphics interface 740 interfaces to a graphics component that provides a visual display to a user of compute node 700. In one example, graphics interface 740 may drive a high-definition (HD) display that provides output to the user. High resolution may refer to a display with a pixel density of approximately 100 PPI (pixels per inch) or greater and may include formats such as Full HD (e.g., 1080p), Retina display, 4K (ultra-high definition or UHD), etc. In one example, the display may include a touchscreen display. In one example, graphics interface 740 generates a display based on data stored in memory 730, or based on operations performed by processor 710, or both. In one example, graphics interface 740 generates a display based on data stored in memory 730, or based on operations performed by processor 710, or both.

[0063] In some embodiments, accelerator 742 may be a fixed-function offload engine accessible or usable by processor 710. For example, one of accelerators 742 may provide data compression functions, cryptographic services such as public key encryption (PKE), ciphers, hashing / authentication functions, decryption, or other functions or services. In some embodiments, additionally or alternatively, one of accelerators 742 provides field selection controller functionality as described herein. In some cases, accelerator 742 may be integrated into a CPU socket (e.g., a connector to a motherboard or circuit board that contains a CPU and provides an electrical interface with the CPU). For example, accelerator 742 may include programmable processing elements such as single or multi-core processors, graphics processing units, logic execution units, single or multi-level caches, functional units that can be used to independently execute programs or threads, application-specific integrated circuits (ASICs), neural network processors (NNPs), programmable control logic, and field-programmable gate arrays (FPGAs). The accelerator 742 may provide multiple neural networks, CPUs, processor cores, general-purpose graphics processing units, or graphics processing units may be made available for use by the AI ​​or ML models. For example, the AI ​​models may use or include any or combinations of reinforcement learning schemes, Q-learning schemes, deep Q-learning, or asynchronous reinforcement learning (Asynchronous Advantage Actor-Critic, A3C), combinatorial neural networks, recurrent combinatorial neural networks, or other AI or ML models. Multiple neural networks, processor cores, or graphics processing units may be made available for use by the AI ​​or ML models.

[0064] Memory subsystem 720 represents the main memory of compute node 700 and provides storage for code executed by processor 710 or data values ​​used in the execution of routines. Memory subsystem 720 may include one or more memory devices 730, such as one or more types of random access memory (RAM), such as read-only memory (ROM), flash memory, DRAM, or other memory devices, or a combination of such devices. Memory 730 stores and hosts, among other things, an operating system (OS) 732, providing a software platform for the execution of instructions within compute node 700. Additionally, applications 734 may execute on the OS 732's software platform from memory 730. Applications 734 represent programs having their own operating logic for performing the execution of one or more functions. Processes 736 represent agents or routines that provide auxiliary functionality to OS 732 or one or more applications 734, or a combination thereof. OS 732, applications 734, and processes 736 provide the software logic for providing functionality to compute node 700. In one example, memory subsystem 720 includes memory controller 722, which is a memory controller that generates commands and issues commands to memory 730. It will be appreciated that memory controller 722 may be a physical part of processor 710 or a physical part of interface 712. For example, memory controller 722 may be an integrated memory controller integrated into circuitry with processor 710.

[0065] Although not specifically shown, it will be understood that the compute node 700 may include one or more buses or bus systems between devices, such as a memory bus, a graphics bus, an interface bus, etc. A bus or other signal line may communicatively or electrically couple components, or may communicatively and electrically couple components. A bus may include a physical communication line, a point-to-point connection, a bridge, an adapter, a controller, or other circuitry, or a combination thereof. A bus may include, for example, one or more of a system bus, a Peripheral Component Interconnect (PCI) bus, a HyperTransport or Industry Standard Architecture (ISA) bus, a Small Computer System Interface (SCSI) bus, a Universal Serial Bus (USB), or an Institute of Electrical and Electronics Engineers (IEEE) Standard 1394 bus (Firewire).

[0066] In one example, compute node 700 includes interface 714, which may be coupled to interface 712. In one example, interface 714 represents an interface circuit, which may include standalone components and integrated circuits. In one example, multiple user interface and / or peripheral components are coupled to interface 714. Network interface 750 provides compute node 700 with the ability to communicate with remote devices (e.g., servers or other computing devices) over one or more networks. Network interface 750 may include an Ethernet adapter, a wireless interconnection component, a cellular network interconnection component, a Universal Serial Bus (USB), or other wired or wireless standards-based or proprietary interface. Network interface 750 may transmit data to devices within the same data center or rack or to remote devices, which may include transmitting data stored in memory. Network interface 750 may receive data from remote devices, which may include storing the received data in memory. Various embodiments may be used in connection with network interface 750, processor 710, and memory subsystem 720.

[0067] In one example, compute node 700 includes one or more IO interfaces 760. IO interface 760 may include one or more interface components through which a user interacts with compute node 700 (e.g., audio, alphanumeric, tactile / touch, or other interface connection). Peripheral interface 770 may include any hardware interface not specifically mentioned above. Peripheral generally refers to devices that connect dependently to compute node 700. A dependent connection is one in which compute node 700 provides a software or hardware platform, or both, on which operations are executed and with which a user interacts.

[0068] In one example, the compute node 700 includes a storage subsystem 780 to store data in a nonvolatile manner. In one example, in a particular system implementation, at least certain components of the storage 780 may overlap with components of the memory subsystem 720. The storage subsystem 780 includes a storage device 784, which may be or may include any conventional medium for storing large amounts of data in a nonvolatile manner, such as one or more magnetic, solid-state, or optical-based disks, or a combination thereof. The storage 784 holds code or instructions and data 786 in a persistent state (i.e., the value is retained despite interruption of power to the compute node 700). While the storage 784 may be generally considered to be “memory,” the memory 730 is typically the execution or operating memory that provides instructions to the processor 710. While the storage 784 is nonvolatile, the memory 730 may include volatile memory (i.e., the value or state of the data is indeterminate if power is interrupted to the compute node 700). In one example, storage subsystem 780 includes a controller 782 to interface with storage 784. In one example, controller 782 may be a physical part of interface 714 or processor 710, or may include circuitry or logic in both processor 710 and interface 714.

[0069] Volatile memory is memory whose state (and therefore the data stored therein) is indeterminate when power to the device is removed. Dynamic volatile memory requires refreshing of the data stored within the device to maintain its state. An example of dynamic volatile memory includes DRAM, or some variants such as synchronous DRAM (SDRAM). As described herein, the memory subsystem may be configured with a variety of memory technologies, including DDR3 (Double Data Rate Version 3, first released by JEDEC (Semiconductor Engineering Association) on June 27, 2007), DDR4 (DDR Version 4, initial specification published by JEDEC in September 2012), DDR4E (DDR Version 4), LPDDR3 (Low Power DDR Version 3, JESD209-3B, published by JEDEC in August 2013), LPDDR4 (LPDDR Version 4, JESD209-4, first published by JEDEC in August 2014), WIO2 (Wide Compatible with many memory technologies, such as Input / Output Version 2, JESD229-2, first published by JEDEC in August 2014), HBM (High Bandwidth Memory, JESD325, first published by JEDEC in October 2013), LPDDR5 (currently under consideration by JEDEC), HBM2 (HBM Version 2), etc., currently under consideration by JEDEC, or combinations of memory technologies, and technologies based on derivatives or extensions of such specifications. JEDEC standards are available at www.jedec.org.

[0070] A non-volatile memory (NVM) device is memory whose state is deterministic even when power to the device is interrupted. In one embodiment, the NVM device may comprise a block-addressable memory device such as NAND technology, or more specifically, multi-threshold level NAND flash memory (e.g., single-level cell ("SLC"), multi-level cell ("MLC"), quad-level cell ("QLC"), tri-level cell ("TLC"), or some other NAND). The NVM device may also comprise a byte-addressable write-in-place three-dimensional cross-point memory device or other byte-addressable write-in-place NVM device (also referred to as persistent memory), such as single-level or multi-level phase change memory (PCM) or switched phase change memory (PCMS), NVM devices using chalcogenide phase change materials (e.g., chalcogenide glasses), metal oxide-based, oxygen vacancy-based, and conductive bridge random access memory (CB-RAM), nanowire memory, ferroelectric random access memory (FeRAM, FRAM®), magnetoresistive random access memory (MRAM) incorporating memristor technology, spin-transfer torque (STT) MRAM, spintronic magnetic junction memory-based devices, magnetic tunnel junction (MTJ)-based devices, DW (domain wall) and SOT (spin orbit transfer)-based devices, resistive memory including thyristor-based memory devices, or any combination of the above, or other memories.

[0071] A power supply (not shown) provides power to the components of the compute node 700. More specifically, the power supply typically interfaces to one or more power sources within the compute node 700 to provide power to the components of the compute node 700. In one example, the power supply includes an AC-DC (alternating current to direct current) adapter that plugs into a wall outlet. Such an AC power source can be a renewable energy (e.g., solar-powered) power source. In one example, the power supply includes a DC power source, such as an external AC-DC converter. In one example, the power supply or power supply includes wireless charging hardware for charging via proximity to a charging field. In one example, the power source can include an internal battery, an AC supply, a motion-based power supply, a solar power supply, or a fuel cell power source.

[0072] In one example, compute node 700 may be implemented using interconnected compute threads of processors, memory, storage, network interfaces, and other components. High-speed interconnects may be used such as Ethernet (IEEE 802.3), Remote Direct Memory Access (RDMA), InfiniBand, Internet Wide Area RDMA Protocol (iWARP), Quick UDP Internet Connection (QUIC), RDMA over Converged Ethernet (RoCE), Peripheral Component Interconnect Express (PCIe), Intel® QuickPath Interconnect (QPI), Intel® Ultra Path Interconnect (UPI), Intel® On-Chip System Fabric (IOSF), Omni-Path, Compute Express Link (CXL), HyperTransport, High-Speed ​​Fabric, NVLink, Advanced Microcontroller Bus Architecture (AMBA) Interconnect, OpenCAPI, Gen-Z, Cache Coherent Interconnect for Accelerators (CCIX), 3GPP Long Term Evolution (LTE) (4G), 3GPP 5G, and variations thereof. Data may be copied or stored on virtualized storage nodes using protocols such as NVMe over Fabrics (NVMe-oF) or NVMe.

[0073] Although some embodiments have been described with reference to particular implementations, other implementations are possible in accordance with some embodiments. Additionally, the arrangement and / or order of elements or other features illustrated in the drawings and / or described herein need not be arranged in the particular manner shown and described. Many other arrangements are possible in accordance with some embodiments.

[0074] In each system shown in the figures, elements in some cases may have the same or different reference numbers, respectively, to suggest that the depicted elements may be different and / or similar. However, elements may be flexible enough to have different implementations and function with some or all of the systems shown or described herein. The various elements shown in the figures may be the same or different. It is arbitrary which is referred to as a first element and which is referred to as a second element.

[0075] In this specification and claims, the terms “coupled” and “connected,” along with their derivatives, may be used. It should be understood that these terms are not intended as synonyms for each other. Rather, in particular embodiments, “connected” may be used to indicate that two or more elements are in direct physical or electrical contact with each other. “Coupled” may mean that two or more elements are in direct physical or electrical contact. However, “coupled” may also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other. Additionally, “communicatively coupled” means that two or more elements, which may or may not be in direct contact with each other, are capable of communicating with each other. For example, if component A is connected to component B, it will in turn be connected to component C, and component A may be communicatively coupled to component C using component B as an intermediary component.

[0076] An embodiment is an implementation or example of the present invention. References herein to "an embodiment," "one embodiment," "some embodiments," or "other embodiments" mean that a particular feature, structure, or characteristic described in connection with multiple embodiments is included in at least some embodiments of the present invention, but not necessarily in all embodiments. The various appearances of "an embodiment," "one embodiment," or "some embodiments" do not necessarily all refer to the same embodiments.

[0077] Not all components, features, structures, characteristics, etc. described or illustrated herein need necessarily be included in a particular embodiment or embodiments. For example, if the specification states that a component, feature, structure, or characteristic "may," "might," "can," or "could" be included, that particular component, feature, structure, or characteristic is not required to be included. When the specification or claims refer to "a" or "an" element, this does not mean that there is only one of that element. When the specification or claims refer to "additional" elements, this does not exclude the presence of more than one of the additional element.

[0078] As noted above, various aspects of the embodiments herein may be facilitated by corresponding software and / or firmware components and applications, such as software and / or firmware executed by an embedded processor or the like. Accordingly, embodiments of the present invention may be used as or support software programs, software modules, firmware, and / or distributed software executed by some form of processor, processing core, or embedded logic or virtual machine operating on a processor or core, or otherwise implemented or realized on or within a non-transitory computer-readable or machine-readable storage medium. A non-transitory computer-readable or machine-readable storage medium includes any mechanism that stores or transmits information in a machine- (e.g., computer)-readable form. For example, a non-transitory computer-readable or machine-readable storage medium includes any mechanism that provides (i.e., stores and / or transmits) information in a form accessible by a computer or computing machine (e.g., computing device, electronic system, etc.), such as recordable / non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.). The content may be directly executable ("object" or "executable" form), source code, or differential code ("delta" or "patch" code). A non-transitory computer-readable or machine-readable storage medium may also include storage or databases from which content may be downloaded. A non-transitory computer-readable or machine-readable storage medium may also include a device or product having content stored therein at the time of sale or delivery. Thus, delivering a device that stores or provides content for download over a communications medium may be understood to provide an article of manufacture that includes a non-transitory computer-readable or machine-readable storage medium with such content as described herein.

[0079] Various components referenced above as processes, servers, or tools described herein may be means for performing the described functions. The operations and functions performed by the various components described herein may be implemented by software running on processing elements, via embedded hardware, etc., or by any combination of hardware and software. Such components may be implemented as software modules, hardware modules, dedicated hardware (e.g., application specific hardware, ASICs, DSPs, etc.), embedded controllers, hardwired circuitry, hardware logic, etc. Software content (e.g., data, instructions, configuration information, etc.) may be provided via an article of manufacture including a non-transitory computer-readable or machine-readable storage medium, which provides content representing instructions that can be executed. The content may, in turn, cause a computer to perform various functions / operations described herein.

[0080] As used herein, a list of items joined by the term "at least one" may mean any combination of the listed terms. For example, the phrase "at least one of A, B, or C" may mean "A," "B," "C," "A and B," "A and C," "B and C," or "A, B, and C."

[0081] The above description of illustrated embodiments of the invention, including what is described in the Abstract, is not intended to be exhaustive or to limit the invention to the precise form disclosed. While specific embodiments of, and examples for, the invention have been described herein for illustrative purposes, various equivalent modifications are possible within the scope of the invention, as those skilled in the art will recognize.

[0082] These variations can be made to the invention in light of the above detailed description. The terms used in the following claims should not be construed to limit the invention to the specific embodiments disclosed in the specification and drawings. Rather, the scope of the invention is to be determined entirely by the following claims, which are to be construed in accordance with established doctrines of interpretation. [Other possible items] [Item 1] 1. A method for balancing power on a computing platform including one or more processing units and a plurality of variable frequency components, the method comprising: monitoring power consumption of the computing platform; Based on the power consumption of the computing platform, the one or more processing units; the plurality of variable frequency components; adjusting the operating frequency of at least one of reducing the power consumption of the computing platform; A method comprising: [Item 2] The computing platform includes a power supply having an associated power supply capacity threshold, and the method comprises: detecting when the power consumption of the computing platform exceeds the associated power supply capacity threshold; adjusting the operating frequency of the at least one of the one or more processing units and the plurality of variable frequency components to reduce the power consumption of the computing platform, such that the power consumption is below the associated power supply capacity threshold; The method of claim 1 further comprising: [Item 3] The method of claim 1 , wherein the platform includes a power supply unit (PSU), and wherein monitoring the power is performed by a sensor that senses the level of power drawn from the PSU. [Item 4] providing power limit bias hints to at least some of the one or more processing units and the plurality of variable frequency components; calculating a power violation index according to a current power consumption and a PSU function limit of the platform; transmitting the power violation index to the one or more processing units and the at least some of the plurality of variable frequency components; adjusting the operating frequency of the at least one of the one or more processing units and the plurality of variable frequency components in response to the power limit bias hint provided to a processing unit or variable frequency component and the power violation index; The method of claim 3 further comprising: [Item 5] The method of claim 4 , wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation to implement a power prioritization scheme. [Item 6] The method of claim 1 , wherein the computing platform is a multi-socket platform including a central processing unit (CPU) per socket, each socket including multiple variable frequency components. [Item 7] employing a first CPU in a first socket to manage power consumption of the first CPU and variable frequency components in the first socket and power consumption of CPUs and variable frequency components in each socket other than the first socket in the multi-socket platform; determining, via the first CPU, power balancing to be implemented to reduce power consumption levels of the platform; transmitting a control signal or message from the first CPU to each CPU in each socket other than the first socket to perform the power balancing, the control signal or message conveying information to be employed to adjust power consumption of at least one of the CPU and one or more variable frequency components in each socket; The method of claim 6 further comprising: [Item 8] 1. A method in which each socket CPU comprises a system on a chip (SoC), the method comprising: For one or more of the plurality of sockets, 8. The method of claim 7, further comprising: sending a power limit message from the SoC to each of the variable frequency components, the power limit message being used to set a maximum frequency at which the variable frequency component operates. [Item 9] 10. The method of claim 1, wherein the variable frequency component comprises one or more of a graphics processor unit (GPU), a general-purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an artificial intelligence (AI) processor, an AI inference unit, a network processor, and a field programmable gate array (FPGA). [Item 10] a central processing unit (CPU) coupled to the memory and configured to vary the operating frequency to vary the power consumption; a plurality of variable frequency components coupled to the CPU, each of the plurality of variable frequency components configured to change an operating frequency to effect a change in power consumption; a firmware storage device having firmware stored therein and operably coupled to said CPU; A power supply unit (PSU) and a power monitor sensor configured to sense power drawn from the PSU; one or more voltage regulators coupled to the PCU and configured to provide power to the CPU and the plurality of variable frequency components; Equipped with The computing platform comprises: Detecting power consumption of the platform via the power monitor sensor; adjusting the operating frequency of the CPU and at least one of the plurality of variable frequency components to reduce the power consumption of the computing platform; configured to: Computational platform. [Item 11] The PSU has a function limit, and the platform further comprises: providing a power limit bias hint to at least a portion of the plurality of variable frequency components; calculating a power violation index according to a current power consumption of the platform and the PSU performance limit; transmitting the power violation index or data associated with the power violation index to the at least some of the plurality of variable frequency components; configured to: each of the variable frequency components configured to adjust its operating frequency in response to the power limit bias hint and the power violation index; The computing platform of claim 10. [Item 12] 12. The computing platform of claim 11, further comprising software stored on a storage device or loaded into memory on the computing platform, wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation via execution of the software on the CPU. [Item 13] The computing platform further comprises: receiving, at the CPU, data or signals indicative of or from which a power level consumed by the computing platform can be derived; calculating the power violation index in the CPU; Calculating or receiving at the CPU a power limit bias hint for the CPU; adjusting an operating frequency of the CPU in response to the power violation index and the power limit bias hint of the CPU; 12. The computing platform of claim 11 configured to: [Item 14] The computing platform further comprises: determining, via the CPU, power balancing to be implemented to reduce power consumption levels of the platform; sending a control signal or message from said CPU to itself and each of said variable frequency components, said control signal or message conveying information to be employed to adjust power consumption of at least one of said CPU and one or more of said variable frequency components; 12. The computing platform of claim 11 configured to: [Item 15] the variable frequency component includes a processing unit having one or more registers, and the computing platform further comprises: sending control signals or messages from the CPU to at least one GPU; updating, in the at least one GPU, a maximum frequency stored in a register on the GPU; 15. The computing platform of claim 14 configured to: [Item 16] A multi-socket platform having a plurality of sockets, each socket comprising: a central processing unit (CPU) coupled to the memory and configured to vary the operating frequency to vary the power consumption; a plurality of accelerators coupled to the CPU, each of the plurality of accelerators configured to change an operating frequency to change power consumption; a firmware storage device having firmware stored therein, the firmware storage device being operably coupled to the at least one socket; A power supply unit (PSU) and a power monitor sensor that senses power drawn from the PSU; one or more voltage regulators coupled to the PCU and configured to provide power to the CPUs and accelerators in the sockets; Including, The multi-socket platform comprises: using an output from the power monitor to detect a power consumption level of the multi-socket platform; one or more CPUs; One or more accelerators adjusting the operating frequency of at least one of Reducing the power consumption of the multi-socket platform; Multi-socket platform configured to: [Item 17] the PSU having a function limit, and the multi-socket platform further comprising: For each socket, calculating a power violation index according to a current power consumption of the platform and the PSU performance limit; providing power limit bias hints to the CPU and the plurality of accelerators; providing the power violation index or data associated with the power violation index to the CPU and the plurality of accelerators; configured to: the CPU and the plurality of accelerators are configured to adjust their operating frequencies in response to the power limit bias hint and the power violation index provided thereto; 17. The multi-socket platform of claim 16. [Item 18] 20. The multi-socket platform of claim 17, further comprising software stored on at least one of a storage device or loaded into memory on the multi-socket platform, wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation via execution of the software on at least one CPU. [Item 19] the threshold is a PSU function limit, and the multi-socket platform further comprises: Calculating a power violation index according to a current power consumption of the platform and the PSU functional limit, the power violation index being calculated by or received by a first CPU in a first socket; transmitting data associated with the power violation index from the first CPU to each other CPU in the socket or sockets other than the first socket; In each socket, determining, via the CPU, power balancing to be implemented for components in the socket; sending control signals or messages from the CPU to itself and at least some of the accelerators, the control signals or messages conveying information to be employed to adjust power consumption of at least one of the CPU and one or more of the accelerators; 17. The multi-socket platform of claim 16 configured to: [Item 20] 17. The multi-socket platform of claim 16, wherein the accelerator comprises one or more of a graphics processor unit (GPU), a general-purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an artificial intelligence (AI) processor, an AI inference unit, a network processor, and a field-programmable gate array (FPGA).

Claims

1. 1. A method for balancing power on a computing platform including one or more processing units and a plurality of variable frequency components, the method comprising: monitoring power consumption of the computing platform; adjusting an operating frequency of at least one of the one or more processing units and the plurality of variable frequency components based on the power consumption of the computing platform to reduce the power consumption of the computing platform; the computing platform includes a power supply unit (PSU); The step of reducing power consumption includes: calculating a power violation index according to a current power consumption of the computing platform and a PSU function limit; transmitting the power violation index to the one or more processing units and at least some of the plurality of variable frequency components; adjusting the operating frequency of the at least one of the one or more processing units and the plurality of variable frequency components in response to the power violation index and a power limit bias hint.

2. The computing platform includes a power supply having an associated power supply capacity threshold, and the method comprises: detecting when the power consumption of the computing platform exceeds the associated power supply capacity threshold; adjusting the operating frequency of the at least one of the one or more processing units and the plurality of variable frequency components to reduce the power consumption of the computing platform, such that the power consumption is below the associated power supply capacity threshold; The method of claim 1 further comprising:

3. 3. The method of claim 1 or 2, wherein the plurality of variable frequency components comprises one or more of a graphics processor unit (GPU), a general purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an artificial intelligence (AI) processor, an AI inference unit, a network processor, and a field programmable gate array (FPGA).

4. A method described in any one of claims 1 to 3, wherein the step of monitoring the power consumption is performed by a sensor that senses the power level drawn from the PSU.

5. The method of claim 1 , further comprising providing the power limit bias hint to the one or more processing units and the at least some of the plurality of variable frequency components.

6. The method of claim 5 , wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation to implement a power prioritization scheme.

7. The method of claim 1 , wherein the computing platform is a multi-socket platform including a central processing unit (CPU) per socket, each socket including multiple variable frequency components.

8. employing a first CPU in a first socket to manage power consumption of the first CPU and variable frequency components in the first socket and power consumption of CPUs and variable frequency components in each socket other than the first socket in the multi-socket platform; determining, via the first CPU, power balancing to be implemented to reduce power consumption levels of the multi-socket platform; transmitting a control signal or message from the first CPU to each CPU of each socket other than the first socket to perform the power balancing, the control signal or message conveying information to be employed to adjust power consumption of at least one of the CPU and one or more variable frequency components of each socket, the information including at least the power violation index; The method of claim 7 further comprising:

9. 1. A method in which a CPU in each socket comprises a system on a chip (SoC), the method comprising: For one or more of the plurality of sockets, 9. The method of claim 8, further comprising: transmitting the control signal or message from the SoC to each of the variable frequency components, the control signal or message being used to set a maximum frequency at which the variable frequency component operates.

10. a central processing unit (CPU) coupled to the memory and configured to vary the operating frequency to vary the power consumption; a plurality of variable frequency components coupled to the CPU, each of the plurality of variable frequency components configured to change an operating frequency to effect a change in power consumption; a firmware storage device having firmware stored therein and operably coupled to said CPU; a power supply unit (PSU); a power monitor sensor configured to sense power drawn from the PSU; one or more voltage regulators coupled to the PSU and configured to supply power to the CPU and the plurality of variable frequency components; Equipped with the PSU has a functional limit; The computing platform is Detecting power consumption of the computing platform via the power monitor sensor; adjusting the operating frequency of the CPU and at least one of the plurality of variable frequency components to reduce the power consumption of the computing platform; configured to: The computing platform further comprises: calculating a power violation index according to a current power consumption of the computing platform and the functional limits of the PSU; transmitting the power violation index or data associated with the power violation index to the at least some of the plurality of variable frequency components; configured to: each of the plurality of variable frequency components configured to adjust its operating frequency in response to the power violation index and a power limit bias hint; Computational platform.

11. The computing platform of claim 10, further configured to provide the power limit bias hint to at least some of the plurality of variable frequency components.

12. 12. The computing platform of claim 10 or 11, further comprising software stored on a storage device or loaded into memory on the computing platform, wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation via execution of the software on the CPU.

13. The computing platform further comprises: receiving, at the CPU, data or signals indicative of or from which the power level consumed by the computing platform can be derived; calculating the power violation index in the CPU; Calculating or receiving at the CPU a power limit bias hint for the CPU; adjusting an operating frequency of the CPU in response to the power violation index and the power limit bias hint for the CPU; 13. A computing platform according to any one of claims 10 to 12, configured to:

14. The computing platform further comprises: determining, via the CPU, power balancing to be implemented to reduce power consumption levels of the computing platform; transmitting a control signal or message from the CPU to itself and each of the plurality of variable frequency components, the control signal or message conveying information to be employed to adjust power consumption of at least one of the CPU and one or more of the plurality of variable frequency components, the information including at least the power violation index; 14. A computing platform according to any one of claims 11 to 13, configured to:

15. the plurality of variable frequency components include a processing unit having one or more registers, and the computing platform further comprises: sending said control signals or messages from said CPU to at least one GPU; updating, in the at least one GPU, a maximum frequency stored in a register on the GPU; The computing platform of claim 14 configured to:

16. one or more accelerators coupled to the CPU, each of the one or more accelerators configured to change its operating frequency to change its power consumption; Furthermore, The computing platform further comprises: The CPU; the one or more accelerators; adjusting the operating frequency of at least one of 16. The computing platform of claim 10, configured to reduce the power consumption of the computing platform.

17. 17. The computing platform of claim 10, wherein the one or more accelerators comprise one or more of a graphics processor unit (GPU), a general purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an artificial intelligence (AI) processor, an AI inference unit, a network processor, and a field programmable gate array (FPGA).

18. A multi-socket platform having a plurality of sockets, each socket comprising: a central processing unit (CPU) coupled to the memory and configured to vary the operating frequency to vary the power consumption; a plurality of accelerators coupled to the CPU, each of the plurality of accelerators configured to change an operating frequency to change power consumption; a firmware storage device having firmware stored therein, the firmware storage device being operably coupled to the at least one socket; a power supply unit (PSU); a power monitor sensor that senses power drawn from the PSU; one or more voltage regulators coupled to the PSU and configured to provide power to the CPUs and accelerators in the sockets; Including, The multi-socket platform comprises: detecting a power consumption level of the multi-socket platform using an output from the power monitor sensor; adjusting an operating frequency of at least one of one or more CPUs and one or more accelerators to reduce the power consumption of the multi-socket platform; configured to: The PSU has a function limit, and the multi-socket platform further comprises: calculating a power violation index according to a current power consumption of the multi-socket platform and the functional limits of the PSU; for each socket, providing the power violation index or data associated with the power violation index to the CPU and the plurality of accelerators; configured to: The multi-socket platform, wherein the CPU and the plurality of accelerators are configured to adjust their operating frequencies in response to the power violation index and a power limit bias hint.

19. The threshold is a PSU function limit, and the multi-socket platform further comprises: calculating a power violation index according to a current power consumption of the multi-socket platform and the PSU functional limits, the power violation index being calculated by or received by a first CPU in a first socket; transmitting data associated with the power violation index from the first CPU to each other CPU in the or a plurality of sockets other than the first socket; In each socket, determining, via the CPU, power balancing to be implemented for components in the socket; sending control signals or messages from the CPU to itself and at least some of the accelerators, the control signals or messages conveying information to be employed to adjust power consumption of at least one of the CPU and one or more of the accelerators, the information including at least the power violation index; 20. The multi-socket platform of claim 18 configured to:

20. 20. The multi-socket platform of claim 18 or 19, wherein the accelerator comprises one or more of a graphics processor unit (GPU), a general-purpose GPU (GP-GPU), a tensor processing unit (TPU), a data processor unit (DPU), an artificial intelligence (AI) processor, an AI inference unit, a network processor, and a field programmable gate array (FPGA).

21. A multi-socket platform as described in any one of claims 18 to 20, wherein the multi-socket platform is further configured to provide the power limit bias hint to the CPU and the multiple accelerators for each socket.

22. 22. The multi-socket platform of claim 21, further comprising software stored on at least one of a storage device or loaded into memory on the multi-socket platform, wherein the power limit bias hints include scaling weights that are dynamically adjusted during platform runtime operation via execution of the software on at least one CPU.

23. a central processing unit (CPU) coupled to the memory and configured to vary the operating frequency to vary the power consumption; a plurality of variable frequency components coupled to the CPU, each of the plurality of variable frequency components configured to change an operating frequency to effect a change in power consumption; a firmware storage device having firmware stored therein and operably coupled to said CPU; a power supply unit (PSU); a power monitor sensor configured to sense power drawn from the PSU; one or more voltage regulators coupled to the PSU and configured to supply power to the CPU and the plurality of variable frequency components; means for detecting power consumption of the computing platform via the power monitor sensor; means for adjusting the operating frequency of the CPU and at least one of the plurality of variable frequency components to reduce the power consumption of the computing platform; Equipped with The PSU has a functional limit, and further calculating a power violation index according to a current power consumption of the computing platform and the functional limits of the PSU; transmitting the power violation index or data associated with the power violation index to the at least some of the plurality of variable frequency components; and means for performing the steps of A computing platform, wherein each of the plurality of variable frequency components is configured to adjust its operating frequency in response to the power violation index and a power limit bias hint.

24. The computing platform of claim 23, further comprising means for providing the power limit bias hint to at least a portion of the plurality of variable frequency components.

25. receiving, at the CPU, data or signals indicative of or from which the power level consumed by the computing platform can be derived; calculating the power violation index in the CPU; Calculating or receiving at the CPU a power limit bias hint for the CPU; adjusting an operating frequency of the CPU in response to the power violation index and the power limit bias hint for the CPU; 25. The computing platform of claim 23 or 24, further comprising means for performing:

Citation Information

Patent Citations

  • Power management based on real time platform power sensing

    US20190041951A1

  • System, apparatus and method for optimized throttling of a processor

    WO2019212669A1