Data Center Level Power Management by Reactive Power Cap Setting
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2023-06-13
- Publication Date
- 2026-05-25
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 409,469, filed September 23, 2022, entitled "DATACENTER LEVEL POWER MANAGEMENT WITH REACTIVE POWER CAPPING," and U.S. Patent Application No. 18 / 202,712, filed May 26, 2023, entitled "DATACENTER LEVEL POWER MANAGEMENT WITH REACTIVE POWER CAPPING," the disclosures of which are incorporated herein by reference in their entireties for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure generally relates to techniques for managing power consumption in a data center. Power caps may be applied at any suitable level in a power distribution hierarchy. Power cap values may be identified and distributed. A timer may be set based on a number of factors. While the timer runs, an attempt may be made to identify a more advantageous power capping solution. If found, the more advantageous power capping solution may be implemented. However, if the timer expires and / or a more advantageous power capping solution is not found, the previously distributed power cap values may be implemented for servers in the data center. [Background technology]
[0003] background Data centers consist of a power infrastructure that provides numerous safety features according to a hierarchical power distribution. Power is supplied by a local utility and allocated according to this power distribution hierarchy to various components of the data center, including power distribution units (PDUs) (e.g., transformers, distribution panels, busways, rack PDUs, etc.) and power-consuming devices (e.g., servers, network devices, etc.). This ensures that the power consumed by all downstream devices adheres to the power limits of each upstream device. When demand peaks and the power consumption of a downstream device exceeds the power limit of an upstream device, the circuit breaker associated with the upstream device may trip, resulting in, at a minimum, a significant interruption to the downstream device, resulting in reduced processing capacity within the data center and a degraded user experience. In a worst-case scenario, the tripped breaker could trigger a cascading power failure of other devices in the same or additional data centers as workloads are redistributed in an attempt to recover from the initial outage. Summary of the Invention [Problem to be solved by the invention]
[0004] Due to the disruptive and costly nature and potentially widespread impact of exceeding these power limits, traditional systems overprovision data center components such that each power distribution device is allocated power that exceeds the maximum expected power consumption of all downstream devices. Downstream devices are allocated power caps that constrain and throttle operation at each device to ensure that the downstream devices adhere to this maximum expected power consumption. These power caps are applied to prevent or correct even momentary excursions beyond the allocated power limit. This ensures that even a worst-case scenario in which the power consumption of all downstream devices simultaneously peaks will not exceed the maximum expected power consumption of an upstream device. These techniques were intended to reduce the risk of tripping a breaker and losing computing resources. However, these approaches lead to wasteful power utilization, leaving the power allocated to upstream devices underutilized. Furthermore, power capping limits operations performed by downstream devices, which may be visible to customers and may result in a degraded user experience. As a result, it is desirable to utilize power management techniques that minimize the frequency with which downstream devices are restricted in operation while maintaining a high degree of safety in terms of avoiding power failures. [Means for solving the problem]
[0005] Quick Overview In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of some embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The illustrations and description are not intended to be limiting.
[0006] Some embodiments may include a method. The method may include a power management service identifying a plurality of components of an electric power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, nodes of the plurality of nodes of the power distribution hierarchy represent corresponding components of the plurality of components. A subset of nodes in a first level of the plurality of levels may be derived from a particular node in a second level of the plurality of levels that is higher than the first level. A set of lower-level components of the plurality of components represented by the subset of nodes in the first level may receive power distributed through a higher-level component of the plurality of components represented by the particular node in the second level. The method may include the power management service monitoring power consumption of the set of lower-level components represented by the subset of nodes in the first level. The method may include the power management service determining, based at least in part on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The method may include, in response to determining that the power consumption of a set of lower-level components has breached a budget threshold associated with the higher-level component, transmitting a power cap value to the lower-level component. In some embodiments, transmitting the power cap value causes the lower-level component to store the power cap value in memory while allowing the power consumption of each of the lower-level components to exceed the power cap value until a time period corresponding to the timing value expires.
[0007] Some embodiments may include a second method. The second method may include a power management service identifying a plurality of components of an electric power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, nodes of the plurality of nodes of the power distribution hierarchy represent corresponding components of the plurality of components. A subset of nodes in a first level of the plurality of levels may be derived from a particular node in a second level of the plurality of levels that is higher than the first level. A set of lower-level components of the plurality of components represented by the subset of nodes in the first level may receive power distributed through a higher-level component of the plurality of components represented by a particular node in the second level. The second method may include the power management service monitoring power consumption of the set of lower-level components represented by the subset of nodes in the first level. The second method may include the power management service determining, at least in part based on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The second method may include, in response to determining that the power consumption of a set of lower-level components has breached a budget threshold associated with a higher-level component, starting a timer corresponding to a timing value. In some embodiments, expiration of the timer indicates expiration of a time period corresponding to the timing value. In some embodiments, a lower-level component of the set of lower-level components stores a power cap value in memory for the time period. In some embodiments, enforcement of the power cap value at the lower-level component is postponed until expiration of the timer.
[0008] Systems, devices, and computer media are disclosed, each of which may include one or more memories that may store instructions corresponding to the methods disclosed herein. The instructions may be executed by one or more processors of the disclosed systems and devices to perform the methods disclosed herein. One or more computer programs may be configured to perform specific operations or acts corresponding to the described methods by including instructions that, when executed by a data processing device, cause the device to perform those acts. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 illustrates an example environment housing various components, according to at least one embodiment. [Figure 2] FIG. 1 is a simplified diagram of an exemplary power distribution infrastructure including various components of a data center, according to at least one embodiment. [Figure 3] 3 illustrates an example power distribution hierarchy corresponding to the arrangement of components in FIG. 2, according to at least one embodiment. [Figure 4] FIG. 1 illustrates an example power management system for managing power allocation and consumption across components of a data center, according to at least one embodiment. [Figure 5] FIG. 1 is a flow diagram illustrating an example method for managing excessive power consumption, according to at least one embodiment. [Figure 6] FIG. 1 is a flow diagram illustrating an example method for training a machine learning model to determine the likelihood that the aggregate power consumption of downstream devices will exceed a budget threshold corresponding to an upstream device, according to at least one embodiment. [Figure 7] FIG. 1 is a block diagram illustrating an example method for managing power, according to at least one embodiment. [Figure 8] FIG. 10 is a block diagram illustrating another example method for managing power, according to at least one embodiment. [Figure 9]FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 11] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 12] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 13] FIG. 1 is a block diagram illustrating an example computer system according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0010] Detailed Description In the following description, various embodiments are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Additionally, well-known features may be omitted or simplified in order not to obscure the described embodiments.
[0011] The present disclosure relates to managing a power distribution system, such as a power distribution infrastructure in a data center. More specifically, techniques are described for enabling components in a data center to safely consume power at a rate closer to the data center's maximum capacity / supply. Traditionally, at least some power capacity of a data center remains unused due to overprovisioning / overallocation of power to various power distribution units (e.g., transformers, distribution panels, busways, rack PDUs, etc.) in the power distribution infrastructure. A "power distribution unit" (also referred to as a "power distribution component") refers to any suitable device, component, or structure configured to distribute power to other devices or areas. "Power overprovisioning," also referred to as "overallocation," refers to a power distribution approach in which power is provisioned / allocated to a power distribution unit according to a worst-case power consumption scenario in which all power-consuming devices in the data center are expected to run at peak capacity. In this approach, each power distribution unit is allocated enough power to accommodate the peak consumption of all downstream devices, plus some buffer. The power allocated to the power distribution units may be referred to as "budget power" or "allocated power."
[0012] It is desirable to balance the supply and consumption of power within a data center. If the power consumed exceeds the available supply, the balance may be restored by increasing the power supply or slowing down the rate of consumption. If this balance is not maintained and a component attempts to consume more than the available supply, a circuit breaker may trip, disconnecting the component from the supply. The disconnected component may cause an interruption to operations performed by the data center components. As an example, a website hosted on a server in a data center will crash if a tripped circuit breaker disconnects the server from the data center's power source. As a result, various techniques have been adopted to avoid circuit breaker tripping.
[0013] The balance between power supply and power consumption in a data center may be managed by maintaining an equilibrium between available supply and demand and / or consumption. Because power is often statically allocated in long-term contracts with electric utilities, increasing or decreasing the power supply in a data center may sometimes be infeasible. While supply may be statically allocated, power consumption in a data center may vary, sometimes drastically. As an overly simplistic example, as the number of threads executed by a server's processor increases, the server's power consumption may increase. The ambient temperature in a data center may affect the power consumed by the data center's cooling system. The cooling system operates at a higher load, consuming more power as it operates to reduce the ambient temperature experienced in the data center. Demand caused by some components in a data center, such as uninterruptible power supplies, power distribution units, cooling systems, or busways, may be difficult to regulate. However, some components, such as servers or networking components, may be more easily constrained.
[0014] Many conventional power management techniques utilize power capping to constrain operation on power consuming devices (e.g., servers, network devices, etc.) in a data center. When power capping is utilized, a power capping limit may be used to constrain the power consumed by the device. The power capping limit is used to constrain (e.g., throttle) operation on a server to ensure that the server's power consumption does not exceed the power cap limit. The use of power capping ensures that the allocated power limit of an upstream device is not breached, leaving over-provisioned power in each upstream device unused. These techniques waste valuable power and limit the density of power consuming devices that may be utilized in a data center.
[0015] Technical effects An efficient power infrastructure within a data center is necessary to increase provider profit margins, manage scarce power resources, and make the services provided by the data center more environmentally friendly. Data centers that include components hosting multi-tenant environments (e.g., public clouds) can experience higher-than-average consumption because not all tenancies of the cloud are in use simultaneously. To improve data center efficiency and resource utilization, data center providers may increase servers and / or tenancies so that power consumption across all power-consuming devices approaches the data center's allocated power capacity. However, in some cases, reducing the gap between allocated power capacity and power consumption increases the risk of tripping circuit breakers and losing the ability to utilize computing resources. The techniques described herein minimize the frequency with which downstream devices are constrained, allowing these devices to utilize previously unused power while maintaining a high degree of safety regarding avoiding power failures.
[0016] 1 illustrates an example environment (e.g., environment 100) including various components according to at least one embodiment. Environment 100 may include a data center 102, which may include dedicated space for hosting any suitable number of servers, such as servers 104A-104P (collectively referred to as "servers 104"), and infrastructure for hosting the servers, such as network hardware, cooling systems, and storage devices. Networking hardware (not shown here) in data center 100 allows remote users to interact with the servers over a network (e.g., the Internet). Any suitable number (e.g., 10, 14, 21, 42, etc.) of servers 104 may be held in various racks, such as racks 106A-106H (collectively referred to as "racks 106"). Racks 106 may include frames or enclosures in which a corresponding set of servers is positioned and / or mounted.
[0017] Various subsets of racks 106 may be organized into groups referred to as "rows" (e.g., rows 108A-108D, collectively referred to as "rows 108"). In some implementations, rows 108 may include any suitable number of racks (e.g., 5, 8, 10, up to 10, etc.) arranged (e.g., within a threshold distance from one another). In other implementations, rows may be organizational units, and racks comprising a given row may be located in different locations (not necessarily within a threshold distance from one another). As an example, rows 108 may be located in rooms (e.g., room 110A, room 110N, etc.). A room (e.g., room 110A) may be a section of a building or a physical enclosure in which any suitable number of racks 106 are located. In other embodiments, rooms may be organizational units, and rooms may be located in different physical locations, or multiple rooms may be located in a single section of a building.
[0018] FIG. 2 illustrates a simplified diagram of an exemplary power distribution infrastructure 200 including various components (e.g., components of the data center 102 of FIG. 1 ) according to at least one embodiment. The power distribution infrastructure 200 can be connected to a utility power source (not shown), and power can be initially received by one or more uninterruptible power supplies (uninterruptible power supply (UPS) 202). In some embodiments, power may be received at the UPS from a utility company via an on-site power substation (not shown) configured to establish suitable voltage levels for distributing power throughout the data center. The UPSs 202 may each include dedicated batteries or generators to provide emergency power in the event of a failure of the input power source. The UPSs 202 may monitor the input power and provide backup power if a drop in input power is detected.
[0019] The power distribution infrastructure 200 may include any suitable number of intermediate power distribution units (PDUs) (e.g., intermediate PDUs 204) that connect to and receive power / electricity from the UPS 202. Any suitable number of intermediate PDUs 204 may be disposed between a UPS (of the UPS 202) and any suitable number of string PDUs (e.g., string PDUs 206). A power distribution unit (e.g., intermediate PDUs 204, string PDUs 206, rack PDUs 208, etc.) may be any suitable device configured to control and distribute power / electricity. Example power distribution units may include, but are not limited to, a main distribution panel, a distribution panel, a remote power panel, a bus bar, a power strip, a transformer, etc. Power may be provided from the UPS 202 to the intermediate PDUs 204. The intermediate PDUs 204 may distribute power to downstream components of the power distribution infrastructure 200 (e.g., string PDUs 206).
[0020] Power distribution infrastructure 200 may include any suitable number of string power distribution units (including string PDUs 206). A string PDU may include any suitable PDU (e.g., remote power panels, bus bars / ways, etc.) disposed between an intermediate PDU (e.g., a PDU among intermediate PDUs 204) and one or more rack PDUs (e.g., rack PDU 208A, rack PDU 208N, collectively referred to as “rack PDUs 208”). A “string PDU” refers to a PDU configured to distribute power to one or more strings of devices (e.g., string 210 including servers 212A-212D, collectively referred to as “servers 212”). As discussed above, a string (e.g., string 210) may include any suitable number of racks (e.g., racks 214A-214N, collectively referred to as “racks 214”) in which servers 212 are located.
[0021] Power distribution infrastructure 200 may include any suitable number of rack power distribution units (including rack PDUs 208). A rack PDU may include any suitable PDU positioned between a column PDU (e.g., column PDU 206) and one or more servers (e.g., servers 212A, 212B, etc.) corresponding to a rack (e.g., rack 214A, which is an example of rack 106 in FIG. 1). A "rack PDU" refers to any suitable PDU configured to distribute power to one or more servers in a rack. A rack (e.g., rack 214A) may include any suitable number of servers 212. In some embodiments, rack PDU 208 may include an intelligent PDU further configured to monitor, manage, and control consumption across multiple devices (e.g., rack PDU 208A, servers 212A, 212B, etc.).
[0022] The servers 212 (each an example of a server 104 in FIG. 1 ) may each include a power controller (power controllers 216A-216D, collectively referred to as “power controller 216”). A power controller refers to any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in that device. The power controllers 216 can individually monitor the power consumption of each server on which they operate. The power controllers 216 can each be configured to implement power capping to constrain the power consumption in their respective servers. Implementing power capping includes any suitable combination of monitoring the power consumption in the server, determining whether to constrain (e.g., constrain within a range, constrain, etc.) the power consumption in the server (e.g., based at least in part on a comparison of the server's current power consumption with a stored power capping limit), and constraining / constraining the power consumption in the server (e.g., using dynamic voltage and frequency scaling of processors and memory to throttle the server's power consumption). The enforcement of power capping may be referred to as "power capping."
[0023] The data center 102 of FIG. 1 may include various components shown in the power distribution infrastructure 200. By way of example, room 110A may include one or more busways (each an example of a row PDU 206). A busbar (also referred to as a "busway") refers to a duct of conductive material through which electrical power can be distributed (e.g., within room 110A). The busway can receive power from the power distribution units of the intermediate PDUs 204 and provide power to one or more racks (e.g., rack 106A, FIG. 1, rack 106B, FIG. 1, etc.) associated with a row (e.g., row 108A, FIG. 1). Each power infrastructure component that distributes / provides power to other components also consumes a portion of the power passing through it. This loss can be caused by heat loss due to the power flowing through that component or by power consumed directly by the component (e.g., power consumed by processors in a rack PDU).
[0024] FIG. 3 illustrates an example power distribution hierarchy 300 corresponding to the arrangement of components in FIG. 2 , according to at least one embodiment. The power distribution hierarchy 300 may represent the arrangement of any suitable number of components of an electric power system, such as the power distribution infrastructure components discussed in connection with the power distribution infrastructure 200 of FIG. 2 . The power distribution hierarchy 300 may include any suitable number of nodes organized according to any suitable number of levels. Each level may include one or more nodes. The root level (e.g., level 5) may include a single root node of the power distribution hierarchy 300. Each node of the power distribution hierarchy 300 may represent a corresponding component of the power distribution infrastructure 200. A set of one or more nodes at a given level may be derived from a particular node corresponding to a higher level of the power distribution hierarchy 300. A set of lower-level components represented by a lower-level (e.g., Level 1) node can receive power distributed through a higher-level component represented by a Level 2 node (e.g., a component upstream from the lower-level component), which in turn receives power from a higher-level component represented by a Level 3 node, which in turn receives power from a higher-level component represented by a Level 4 node. In some embodiments, all components of the power system receive power initially distributed through a component corresponding to a Level 5 (e.g., the top level of the power distribution hierarchy 300) node (e.g., node 302, the root node). Node 302 can receive power from a utility power source (e.g., a local electric utility system).
[0025] As shown, power distribution hierarchy 300 includes node 302 at level 5. In some embodiments, node 302 may represent an uninterruptible power supply, such as UPS 202 of FIG. 2. A component corresponding to node 302 may distribute / supply power to a component corresponding to node 304 at level 4. A component (e.g., a lower-level component) that receives power from a higher-level component (a component represented by a node at a higher level than the level of a node representing a lower-level component in power distribution hierarchy 300) may be considered subordinate to the higher-level component. In some embodiments, node 304 may represent a component, such as one of intermediate PDUs 204 of FIG. 2 (e.g., a power distribution panel), that is subordinate to the UPS represented by node 302. The component represented by node 304 may be configured to distribute / supply power to components represented by nodes 306 and 308, respectively.
[0026] Level 3 nodes 306 and 308 may each represent a respective component (e.g., a respective string PDU) in FIG. 2. The components (e.g., string PDU 206, bus bars) corresponding to node 306 may distribute / supply power to the components (e.g., rack PDU 208A in FIG. 2) corresponding to node 310 and the components (e.g., rack PDU 208N in FIG. 2) corresponding to node 312. The components (e.g., rack PDU 208A) corresponding to node 310 may distribute / supply power to the components (e.g., rack PDU 208A) corresponding to node 310 in FIG. 2 (representing servers 212A and 212B in FIG. 2, respectively), which may be monitored / managed by components of servers 212A and 212B, such as power controllers 216A and 216B. The components corresponding to nodes 314 and 316 may be located in the same rack.
[0027] A component corresponding to node 312 (e.g., rack PDU 208A) can distribute / supply power to a component corresponding to level 1 node 318 (e.g., server 214C in FIG. 2 including power controller 216C). Level 1 nodes 314, 316, and 318 can be located in and / or associated with the same column (e.g., column 210 in FIG. 2).
[0028] Returning to node 308 in level 3, node 308 (e.g., a different column PDU, a remote power panel) can distribute / supply power to components corresponding to node 310 (e.g., rack PDU 208A in FIG. 2 ) and components corresponding to node 312 (e.g., rack PDU 208N in FIG. 2 ). Components corresponding to node 320 (e.g., a rack PDU) can distribute / supply power to components corresponding to nodes 322 and 326 in level 1 (e.g., corresponding to respective servers including their respective power controllers). Components corresponding to nodes 324 and 326 can be located in the same rack. Components corresponding to node 322 (e.g., another rack PDU) can distribute / supply power to components corresponding to nodes 328 and 330 in level 1 (e.g., corresponding to respective servers including their respective power controllers). Components corresponding to nodes 328 and 330 can be located in the same rack.
[0029] The particular number of components (e.g., corresponding to level 1 nodes) that receive distributed power from components at a higher level (e.g., corresponding to level 2 nodes) may differ from the number shown in FIG. 3. The particular number of levels in the power distribution hierarchy may vary depending on the particular arrangement of components used in a given data center. It is contemplated that each non-root level (e.g., levels 1-4 in the example of FIG. 3) may include a different number of nodes representing a different number of components than the number of nodes shown in each non-root level in FIG. 3. The nodes may be arranged in different, yet similar, configurations to those shown in FIG. 3.
[0030] FIG. 4 illustrates an example power management system 400 (for brevity, “system 400”) for managing power allocation and consumption across components of a data center, according to at least one embodiment. Allocating power may refer to the process of allocating a budgeted amount of power (e.g., an expected amount of load / power consumption) to any suitable component. System 400 may include PDUs 402, devices 404, and a service provider computer 406. PDU 402 may be an example of any suitable power distribution unit (e.g., UPS 202, intermediate PDU 204, row PDU 206, rack PDU 208, etc.) described above in connection with FIG. 2. By way of example, PDU 402 may be an example of rack PDU 208 of FIG. 2. Devices 404 may be examples of servers and / or networking devices to which PDU 402 distributes power.
[0031] The PDUs 402, devices 404, and service provider computers 406 can communicate over one or more wired or wireless networks (e.g., network 808). In some embodiments, the network 408 may include any one or combination of many different types of networks, such as a cable network, the Internet, a wireless network, a cellular network, and other private and / or public networks.
[0032] The devices 404 may be any suitable type of computing device, such as, without limitation, a server device, a networking device, or any suitable device in a data center. In some embodiments, the PDUs 402 and the devices 404 are arranged in a power distribution hierarchy, such as the power distribution hierarchy 300 discussed in connection with FIG. 3. In some embodiments, the devices 404 may correspond to and be represented by a level 1 node of the power distribution hierarchy 300. The PDUs 402 may correspond to and be represented by a node at a higher level in the power distribution hierarchy (e.g., a level 2 node).
[0033] Each of PDU 402, device 404, and service provider computer 406 may include at least one memory (e.g., memory 410, memory 412, and memory 414, respectively) and one or more processing units (e.g., processor 416, processor 418, and processor 420, respectively). Processors 416-420 may each be implemented in hardware, computer-executable instructions, firmware, or a combination thereof, as appropriate. The computer-executable instruction or firmware implementation of processors 416-420 may include computer-executable or machine-executable instructions written in any suitable programming language to perform the various functions described.
[0034] The memories 410-414 can store program instructions that are loadable and executable on the respective processors of a given device, as well as data generated during the execution of these programs. Depending on the configuration and type of user computing device, the memories 410-414 can be volatile (such as random access memory (RAM)) and / or nonvolatile (such as read-only memory (ROM), flash memory, etc.). The PDU 402, the device 404, and the service provider computer 406 can also include additional removable and / or non-removable storage devices, including, but not limited to, magnetic storage devices, optical disks, and / or tape storage devices. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing devices. In some implementations, the memories 410-414 can individually include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM.
[0035] Referring more particularly to the contents of memories 410-414, memories 410-414 may include an operating system (e.g., operating system 422, operating system 424, and operating system 426, respectively), one or more data stores (e.g., data store 428, data store 430, and data store 432, respectively), and one or more application programs, modules, or services.
[0036] PDU 402, device 404, and service provider computer 406 may include communication connections (e.g., communication connection 434, communication connection 436, and communication connection 438, respectively) that allow PDU 402, device 404, and service provider computer 406 to communicate with each other over network 408. PDU 402, device 404, and service provider computer 406 may also include I / O devices (e.g., I / O device 440, I / O device 442, and I / O device 444, respectively), such as a keyboard, mouse, pen, voice input device, touch input device, display, speaker, printer, etc., and in some embodiments, service provider computer 406 may be one of devices 404.
[0037] Any suitable combination of devices 404 (e.g., server 104 of FIG. 1 , server 214 of FIG. 2 , etc.) may include a respective power controller (e.g., power controller 446). Power controller 446 may be an example of power controller 216 of FIG. 2. Power controller 446 may be any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in the device. Power controller 446 may individually monitor the power consumption of each device on which it operates. Power controllers 446 may each be configured to implement power capping limits to constrain power consumption in their respective devices. Enforcing a power cap limit may include monitoring power consumption at the device, determining whether to constrain (e.g., bound, constrain, etc.) power consumption at the device (e.g., based at least in part on a comparison of the device's current power consumption to a stored power cap limit), and bounding / constraining the power consumption at the device (e.g., using dynamic voltage and / or frequency scaling by processor 418 and / or memory 412 to throttle power consumption at the device). Any suitable operations associated with power cap setting may be implemented by power controller 446.
[0038] In some embodiments, power controller 446 can communicate with power controller 448 via a direct connection (e.g., via a cable) and / or via network 408. Power controller 446 can provide power consumption data indicative of the device's current power consumption (e.g., cumulative power consumption over a period of time, current power consumption rate, etc.). The power consumption data can be provided to power controller 448 at any suitable frequency, periodicity, or according to a predefined schedule or event (e.g., when a predefined consumption threshold is breached, when a change in consumption rate reaches a threshold, when one or more predefined conditions are determined to be met, when thermal attributes of the device are determined, etc.).
[0039] In some embodiments, power controller 446 can receive a power cap value (also referred to as a "power cap") from power controller 448. In some embodiments, additional data can be provided along with the power cap value. For example, an indicator may be included with the power cap value indicating whether the power cap value should be applied immediately. In some embodiments, the received power cap may be applied / enforced immediately by default. In other embodiments, the received power cap may not be applied / enforced immediately by default.
[0040] When applying / enforcing a power cap, the power controller 446 can monitor power consumption at the device. This may include utilizing a metering device (e.g., an example of an I / O device 442) or software to identify / calculate power consumption data for the device (e.g., cumulative power consumption over a period of time, a current power consumption rate, a current change in power consumption rate over a time window, etc.). As part of applying / enforcing a power cap (also referred to as “power capping”), the power controller 446 can determine whether to constrain (e.g., limit within a range, constrain, etc.) the power consumption at the device (e.g., based at least in part on comparing the device's current consumption rate with a stored power cap value). When constraining power consumption at the device, the power controller 446 can limit / constrain the power consumption at the device within a range (e.g., using dynamic voltage and frequency scaling with the processor 418 and / or memory 412 to throttle power consumption at the device). In some embodiments, the power controller 446 can execute instructions to limit / constrain the power consumption of a device when the device's current consumption data (e.g., cumulative consumption, consumption rate over a time window, etc.) approaches an enforced power cap (e.g., breaches a threshold below the power cap). When constraining / limiting (also referred to as "throttling") the power consumption of a device, the power controller 446 ensures that the device's power consumption stays below the power consumption indicated by the power cap.
[0041] In some embodiments, the power controller of the device 404 can be configured to allow the device 404 to run unconstrained (e.g., without constraints based on power consumption and power caps) until instructed to apply / enforce a power cap by the power controller 448. In some embodiments, the power controller 446 may receive a power cap value that it may or may not be instructed to enforce later, but once received, the power cap value can be stored in memory 412 without being utilized in power management at the device. Thus, in some embodiments, the power controller 446 can refrain from initiating a power cap setting operation (e.g., the comparison and determination described above) until instructed to do so (e.g., via an indication provided by the power controller 448). The indication to begin enforcing the power cap can be received along with the power cap value or may be received as a separate communication from the power controller 448.
[0042] According to power distribution hierarchy 300, each power controller (e.g., power controller 448) of PDU 402 can be configured to distribute, manage, and monitor power for any suitable number of devices 404 (e.g., a rack of servers including servers 212A and 212B, as shown in FIG. 2). In some embodiments, power controller 448 may be a computing agent or program installed on a given PDU (e.g., rack PDU 208A of FIG. 2). Power controller 448 can be configured to obtain power consumption data from devices 404 corresponding to its associated devices (e.g., devices for which a remote controller 448 operating with the given PDU is configured to distribute, manage, or monitor power). Power controller 448 can receive consumption data according to a predefined frequency, periodicity, or schedule implemented by power controller 446 and / or power controller 448 can receive consumption data in response to requesting consumption data from power controller 446. The power controller 448 can be configured to request consumption data from the power controller 446 according to a predefined frequency, periodicity, or schedule implemented by the power controller 448 .
[0043] The power controller 448 may transmit the received consumption data to a power management service 450 implemented by the service provider computer 406 at any suitable time, according to any suitable frequency, periodicity, or schedule, or as a result of one or more predefined conditions being met (e.g., changes in individual, cumulative, or collective consumption rates of the devices 404 breaching a threshold). In some embodiments, the power controller 448 may aggregate the consumption data received from the devices 404 (e.g., via each device's power controller 446) and then transmit the aggregated consumption data to the power management service 450. In some embodiments, the consumption data may be aggregated by device or across all of the devices 404.
[0044] The power controller 448 can receive one or more power cap values corresponding to the device 404 at any suitable time. These power cap values can be calculated by the power management service 450. These calculations are discussed in further detail with reference to FIG. 5. In some embodiments, the power controller 448 can receive, along with the power cap value, a timing value indicating a duration to be used for the timer. The power controller 448 can be configured to generate and start a timer having an associated duration / period corresponding to the timing value. When the timer expires to indicate that a period corresponding to the period has elapsed, the power controller 448 can transmit data to one or more devices including an indicator or other suitable data instructing the one or more devices to proceed with power cap settings using the power cap value previously stored in each device. In some embodiments, the power cap value can be provided along with the indicator. In other embodiments, the power cap value may be provided by the power controller 448 to a power controller (e.g., power controller 446) of the device immediately after receiving the power cap value from the power management service 450. Thus, in some embodiments, the power cap value is transmitted by power controller 448 to power controller 446 and stored in memory 412, but is not implemented by power controller 446 until power controller 446 receives subsequent data from power controller 448 instructing power controller 446 to begin power cap setting.
[0045] Service provider computer 406, possibly deployed in a cluster of servers or as a server farm, can implement power management service 450. In some embodiments, the functionality described in connection with service provider computer 406, including the functionality described in connection with power management service 450, is performed by one or more virtual machines implemented in a hosted computing environment (e.g., on devices 404 corresponding to any suitable number of servers 104 of FIG. 1). A hosted computing environment can include one or more rapidly provisioned and released computing resources, which may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment. Numerous example cloud computing environments are provided and described in more detail below with respect to FIGS. 10-14.
[0046] The power management service 450 can be configured to receive consumption data from the PDU 402 (e.g., from the power controller 448). As described above, the consumption data may be aggregated or cumulative for one or more devices. As a non-limiting example, an instance of consumption data received by the power management service 450 may include cumulative / aggregate consumption data for all devices associated with (e.g., managed by) a given PDU. By way of example, the PDU providing the consumption data may be a rack PDU, and the consumption data provided by that PDU may include cumulative / aggregate and / or individual consumption data values for all devices in the rack. In some embodiments, the cumulative / aggregate consumption data values may further include the power consumption of the PDU. From this consumption data, the power management service 450 can access individual and / or aggregate or cumulative power consumption data for the individual device or all of the devices with which that instance of consumption data is associated. The power management service 450 may be configured to perform any suitable operations (e.g., aggregating consumption data, calculating power cap values, calculating timing values, determining a power budget, determining whether power capping is necessary (e.g., when aggregated consumption has breached or is likely to breach the power budget, etc.) based, at least in part, on data representing the power distribution hierarchy 300 and / or any suitable configuration data specifying the placement of components within the data center (e.g., indicating which devices distribute power to which devices). By way of example, the configuration data may include data representing the power distribution hierarchy 300.
[0047] The power management service 450 can calculate aggregated / cumulative consumption data to identify aggregated / cumulative consumption data for one or more devices at higher levels of the power distribution hierarchy 300. By way of example, the power management service 450 can utilize consumption data provided by one or more rack PDUs to calculate aggregated / cumulative consumption data for column devices. For example, consumption data corresponding to devices represented by level 1 nodes in FIG. 3 that share a common set of nodes up to level 3 of the hierarchy. In this case, the common level 3 node represents a column PDU, such as a bus bar. In some embodiments, the consumption data provided to the power management service 450 can include consumption data indicating the power consumption of one or more rack PDUs (e.g., example components corresponding to level 2 of the power distribution hierarchy 300). The consumption data provided by the rack PDUs can be provided according to any suitable frequency, periodicity, or schedule, or in response to a request sent by the power management service 450. The power management service 450 may obtain consumption data associated with any suitable number of devices 404 and / or racks of devices 404 and may aggregate or calculate any suitable consumption data corresponding to any suitable number of components represented by any suitable nodes and / or levels of the power distribution hierarchy 300.
[0048] The power management service 450 can be configured to calculate one or more power caps (e.g., power cap values for one or more servers to which power is distributed by a higher-level device (e.g., in this example, a row-level device) based, at least in part, on the allocated power values (e.g., amounts of budgeted power) for that higher-level component. The higher-level component can correspond to any suitable level (e.g., levels 2-5) of the power distribution hierarchy 300 other than the lowest level (e.g., level 1). As an example, the power management service 450 can calculate power cap values for servers to which power is distributed by a given row device (e.g., bus bar). These calculations can be performed using , may be based on consumption data provided by rack PDUs to which power is distributed by the column devices. In some embodiments, the power management service 450 may store the consumption data for subsequent use. The power management service 450 may utilize historical consumption data when calculating these power cap values. In some embodiments, the power management service 450 may obtain, utilize, and / or train one or more machine learning models from the historical consumption data to identify specific power cap values for one or more devices 404 (e.g., components corresponding to level 1 of the power distribution hierarchy). These techniques are discussed in more detail with respect to FIG. 7.
[0049] The power management service 450 can calculate timing values for timers (e.g., timers that can be started and managed by the PDU 402) that can be calculated based, at least in part, on any suitable combination of the rate of change of power consumption of the higher-level components, the direction of change (increase / decrease) of power consumption of the higher-level components, the power spike tolerance of the higher-level components and / or the system 400 as a whole, or the associated enforcement times for the lower-level devices (e.g., estimated / known delays between when each of the devices 404 is instructed to enforce a power cap and when each of the devices 404 will actively enforce the cap (e.g., the first time that power consumption at that device is constrained / throttled, or at least a determination is made whether to constrain / throttle).
[0050] The power management service 450 can communicate the calculated timing values to the PDU 402 at any suitable time. In some embodiments, the power management service 450 can first determine a power cap for a given higher-level component (e.g., a column component) regardless of the consumption occurring for other components at the same level (e.g., other column components). In some embodiments, while the timer is being initialized or elapses, or at any suitable time, the power management service 450 can process consumption data corresponding to other same-level components to determine whether it is more desirable to set a power cap for one or more downstream components of those same-level components. In some embodiments, the power management service 450 can determine the power cap value using priorities associated with the devices 404 and / or workloads running on those devices. The power management service 450 can be configured to prioritize power cap settings for lower priority devices / workloads while leaving higher priority devices / workloads unconstrained. In some embodiments, power management service 450 can be configured to prioritize power capping of a set of highest-consuming devices (e.g., across a row, across multiple rows, etc.) while allowing lower-consuming devices to operate unconstrained. In some embodiments, a particular consumption of a device may remain unconstrained even if the device is included in the set of highest-consuming devices if the priority associated with that device and / or workload is high (higher than the priority associated with other devices and / or workloads). Thus, power management service 450 can be configured to prioritize the priority of a device / workload over the power consumption of that particular device.
[0051] In some embodiments, for a given column, power cap values may be initially determined by a power management service. These power cap values may be provided to power controller 448, which may then distribute the power cap values to power controller 446 for storage in device 404. Power management service 450 may process consumption data for other column devices in the same column to determine power cap values for devices corresponding to different columns. This may be advantageous because devices managed by another column may not be consuming their budgeted power, leaving some amount of power unused for devices at that column level. In some embodiments, power management service 450 may be configured to determine whether it may be more advantageous to set power caps for devices in one column while allowing at least some devices in another column to operate unconstrained. Determining the merits of a set of power cap values may be based, at least in part, on minimizing the number of devices that should have power caps set, minimizing priority values associated with devices that should have power caps set, maximizing the number of devices associated with a particular high priority value that will not have power caps set, etc. The power management service 450 can utilize a predefined protocol (e.g., a set of rules) to determine whether the implementation of a power cap value that it has already sent to one PDU is actually more advantageous / beneficial than a different power cap setting that it has identified based on processing consumption data from multiple PDUs associated with one or more other strings.
[0052] The power management service 450 can be configured to enforce the set of power cap values determined to be more advantageous to ensure that one or more circuit breakers in the data center do not trip. In some embodiments, the set of power cap values determined to be most advantageous can be selected. If the set of power caps previously provided to the power controller 448 is not determined to be more advantageous, the power management service 450 can send data to the power controller 448 to cause the power controller 448 to send an indication to the device 404 (e.g., the power controller 446) to initiate power capping based on the previously allocated power caps. Alternatively, if the power management service 450 determines that a new set of power caps is more (or most) advantageous from a power management perspective, the power management service 450 can send data to the power controller 448 to cancel a timer. In some embodiments, canceling the timer may cause the previously allocated power caps to be deleted by instruction from the power controller 448 (e.g., sent in response to canceling the timer), or they may time out by default (e.g., according to a predefined period). The power management service 450 can transmit new power caps in appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was transmitted) along with an indication that the power caps should be immediately enforced by the corresponding devices. These PDUs can transmit the power caps to the receiving devices along with instructions to immediately initiate power capping operations (including, for example, monitoring consumption relative to the power cap values, determining whether the cap is based on the power cap values and the device's current consumption, and either capping the device's operation or leaving such operation unconstrained based on the determination).
[0053] In some embodiments, if the timer expires (e.g., the period corresponding to the timing value has passed) in the original PDU and a cancellation has not been received from the power management system 450, the power controller 448 can automatically instruct the device 404 to initiate a power capping operation at the stored power cap value. This technique ensures a fail-safe in case the power management service 450 fails to instruct the PDU or cancel the timer for any reason.
[0054] The techniques described above allow more lower-level devices to operate unconstrained by minimizing the number and / or frequency at which device power consumption is capped. Furthermore, devices downstream of a given device in the hierarchy can be allowed to peak above the device's budgeted power while still ensuring that the device's maximum power capacity (a value higher than the budget value) is not exceeded. This allows the consumption of lower-level devices to operate within a buffer of power that conventional systems would leave unused. The present system and described techniques allow for more efficient utilization of the power distribution components of system 400, reducing waste while ensuring that power failures are avoided.
[0055] Figure 5 is a flow diagram illustrating an example method 500 for managing excessive power consumption, according to at least one embodiment. Method 500 may include more or fewer operations than those shown or described with respect to Figure 5. The operations may be performed in any suitable order.
[0056] Method 500 may begin at 510, where consumption data may be received and / or obtained (e.g., upon request) by PDU 504 from devices 502. Devices 502 may be examples of devices 404 of FIG. 4 (e.g., multiple servers) and may be located in a common server rack (or may be configured to receive power from at least PDU 504). PDU 504 is an example of PDU 402 of FIG. 4 (e.g., rack PDU 208A of FIG. 2). Each of devices 502 may be a device to which PDU 504 distributes power. PDU 508 may be an example of another rack PDU associated with the same column as PDU 504 and therefore receiving power from the same column PDU as PDU 504 (e.g., column PDU 206 of FIG. 2, corresponding to node 306 of FIG. 3). The consumption received from PDU 508 may correspond to the power consumption of devices 512. In some embodiments, at least some portions of PDU 508 and devices 512 correspond to a different column than the column corresponding to device 502. In some embodiments, any suitable number of instances of consumption data (referred to as “device consumption data”) received by PDU 504 may relate to a single one of devices 502. The same may be true for device consumption data received by PDU 508. In some embodiments, PDU 504 and PDU 508 may aggregate and / or perform calculations from the received device consumption data instances to generate rack consumption data corresponding to each respective PDU. For example, PDU 505 may generate rack consumption data from device consumption data provided by device 502. The rack consumption data may include device consumption data instances received by PDU 504 at 510. Similarly, rack consumption data provided by PDU 508 may include device consumption instances received from device 512. The rack consumption data generated by the PDU 504 and / or the PDU 508 may be generated based at least in part on the power consumption of the PDU 504 and / or the PDU 508, respectively.
[0057] At 514, power management service 506 (an example of power management service 450 in FIG. 4 ) can receive and / or obtain (e.g., via a request) rack consumption data from any suitable combination of PDU 504 and PDU 508 (also an example of PDU 402 in FIG. 4 ), not necessarily simultaneously. In some embodiments, power management service 506 can continuously receive / obtain rack consumption data from PDU 504 and PDU 508. PDU 504 and PDU 508 can similarly receive and / or obtain (e.g., via a request) device consumption data from devices 502 and 512, respectively (e.g., from their corresponding power controllers, each power controller being an example of power controller 446 in FIG. 4 ). PDU 508 can correspond to a PDU associated with the same column as PDU 504.
[0058] At 516, the power management service 506 can determine whether to calculate a power cap value for the device 502 based, at least in part, on maximum and / or budget power amounts associated with another PDU (e.g., a string-level device not shown in FIG. 5 , such as a bus bar) and consumption data received from the PDU 504 and / or PDU 508. In some embodiments, the power management service 506 can access these maximum and / or budget power amounts associated with a string-level PDU (e.g., string PDU 206 of FIG. 2 , not shown here) from stored data, or the maximum and / or budget power amounts can be received by the power management service 506 at any suitable time from the PDU to which the maximum and / or budget power amounts are associated (the string-level PDU). The power management service 506 can aggregate the rack consumption data from the PDUs 504 and / or 508 to determine cumulative power consumption values for the row-level devices (e.g., total power consumption from devices downstream of the row-level device within the last time window, a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, a period of time required to initiate power capping on one or more of the devices 502 and / or 512, etc.). In some embodiments, the power management service 506 can generate additional cumulative power consumption data from at least a portion of the cumulative power consumption data values and / or using historical device consumption data (e.g., cumulative or individual consumption data). For example, the power management service 506 can calculate any suitable combination of a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, etc. based, at least in part, on historical consumption data (e.g., historical device consumption data and / or historical rack consumption data corresponding to the devices 502 and / or 512).
[0059] The power management service 506 may determine that a power cap should be calculated if the cumulative power consumption (e.g., power consumption corresponding to devices 502 and 512) exceeds the budgeted power amount associated with the column-level PDU. If the cumulative power consumption does not exceed the budgeted power amount associated with the column-level PDU, the power management service 506 may determine that a power cap should not be calculated, and method 500 may end. Alternatively, the power management service may determine that a power cap should be calculated due to the cumulative power consumption of downstream devices (e.g., devices 502 and 512) exceeding the budgeted power amount associated with the column-level PDU, and may proceed to 518.
[0060] In some embodiments, the power management service 506 can additionally or alternatively determine that a power cap value should be calculated at 516 based, at least in part, on providing historical consumption data (e.g., device consumption data and / or rack consumption data corresponding to devices 502 and / or 512) as input to one or more machine learning models. The machine learning models can be trained using any suitable supervised or unsupervised machine learning algorithm to identify, from the historical consumption data provided as input, a likelihood that consumption corresponding to a device associated with the historical consumption data will exceed a budgeted amount (e.g., a budgeted amount of power allocated to a row-level device). The machine learning models can be trained using a corresponding training dataset that includes historical consumption data instances. Training of these machine learning models is discussed in more detail with respect to FIG. 6. Using the techniques described above, the power management service 506 can identify that a power cap value should be calculated based on cumulative device / rack consumption levels that exceed the budgeted power associated with the row-level devices and / or based on determining, from the output of the machine learning model, that the cumulative device / rack consumption levels are likely to exceed the budgeted power associated with the row-level devices. Determining that a device / rack consumption level is likely to exceed the row-level device's power budget can be determined based on receiving an output from a machine learning model indicating the likelihood (e.g., likely / unlikely) or by comparing an output value (e.g., a percentage, confidence value, etc.) indicating the likelihood of exceeding the row-level device's power budget with a predefined threshold. The output value indicating the likelihood of exceeding the predefined threshold allows the power management service 506 to determine that power capping is justified and that a power cap value should be calculated.
[0061] At 518, the power management service 506 can calculate power cap values for any suitable number of devices 502 and / or devices 512. By way of example, the power management service 506 can utilize device consumption data and / or rack consumption data corresponding to each of the devices 502 and 512. In some embodiments, the power management service 506 can determine the difference between the cumulative power consumption (calculated from aggregating the device and / or rack consumption data corresponding to devices 502 and 512) of a column of devices (e.g., device 502 and any of the devices 512 in the same column) and the budgeted power associated with the column-level devices. This difference can be used to identify the amount by which power consumption should be constrained at the devices associated with the column. The power management service 506 can determine power cap values for any suitable combination of devices 502 and / or 512 for devices in the same column based on identifying power cap values that, when implemented by the devices, would reduce power consumption to a value less than the budgeted power associated with the column-level devices.
[0062] In some embodiments, the power management service 506 can determine specific power cap values for any suitable combination of devices 502 and 512 corresponding to the same column based, at least in part, on consumption data and / or priority values associated with those devices and / or workloads associated with the power consumption of those devices. In some embodiments, these priority values may be provided as part of device consumption data provided by the devices whose consumption data are associated, or the priority values may be ascertained, at least in part, based on the type of device, the type of workload, etc., obtained from the consumption data or any other suitable source (e.g., from separate data accessible to the power management service 506). Using the priorities associated with each device or workload, the power management service 506 can calculate power cap values such that it prioritizes setting power caps on devices associated with lower priority workloads over setting power caps on devices with higher priority workloads. In some embodiments, the power management service 506 may calculate power cap values based, at least in part, on prioritizing setting power caps on devices consuming at higher rates (e.g., the set of highest consuming devices) over setting power caps on other devices consuming at lower rates. In some embodiments, the power management service 506 may calculate power cap values based, at least in part, on a combination of factors including the consumption value for each device and the priority associated with the device or the workload being executed by the device. The power management service 506 may identify power cap values for devices that entirely avoid capping high-consuming and / or high-priority devices or that cap those devices to a lesser extent than devices that consume less power and / or are associated with a lower priority.
[0063] At 520, the power management service 520 can calculate timing values corresponding to durations for timers that the PDU (e.g., PDU 504) is to initialize and manage. The timing values can be calculated based, at least in part, on the rate of change of power consumption from the perspective of the column-level devices, the direction of change of power consumption from the perspective of the column-level devices, or any suitable combination of the time during which power capping should be performed at each of the devices to which the power cap calculated at 518 is associated. As an example, the power management service 520 can determine a timing value that is greater than timing values for smaller increases in power consumption from the perspective of the column-level devices based on determining a relatively large increase in power consumption from the perspective of the column-level devices. Thus, a larger increase in the power consumption rate can result in a smaller timing value (corresponding to a shorter timer), while a smaller increase in the power consumption rate can result in a larger timing value (corresponding to a longer timer). Similarly, the calculated first timing value can be smaller than the calculated second timing value if the time during which power capping should be performed at each of the devices to which the power cap is associated is shorter relative to the calculation of the first timing value. Thus, the faster a device intended for power capping can implement power capping, the shorter the timing value may be.
[0064] At 522, the calculated power cap values calculated at 518 and / or the timing values calculated at 520 may be transmitted to the PDU 504. Although not shown, at 520, corresponding power cap values and / or timing values for any of the devices 512 in the same column may be transmitted.
[0065] At 524, if a timing value was provided at 522, the PDU 504 may use the timing value to initialize a timer having a duration corresponding to the timing value.
[0066] At 526, the PDU 504 may transmit the power cap value (if received at 522) to the device 502. The device 502 may store the received power cap value at 528. In some embodiments, the power cap value provided at 526 may include an indicator that enforcement should not begin, or may not provide an indicator, and the device 502 may not initiate power capping by default.
[0067] At 530, the power management service 506 can perform a high-level analysis of device / rack consumption data received from any of the devices 512 corresponding to a different column than the column to which the device 502 corresponds. As part of this process, the power management service 506 can identify unused power associated with other column-level devices. If unused power exists, the power management service 506 can calculate a new set of power cap values based, at least in part, on consumption data corresponding to at least one other column of devices. As discussed above, this can be advantageous because devices managed by another column-level device may not be consuming their budgeted power, leaving unused power for that column-level device. In some embodiments, the power management service 450 can be configured to determine whether it may be more advantageous to allow at least some of the devices 502 (and possibly some of the devices 512 corresponding to the same column as the device 502) to operate unconstrained, while setting caps on devices in another column. The benefit of each power capping technique may be calculated based, at least in part, on minimizing the number of devices that should have power caps set, minimizing the priority values associated with devices that should have power caps set, maximizing the number of devices associated with a particular high priority value that will not have power caps set, etc. The power management service 506 may utilize a predefined scheme or set of rules to determine whether implementing a power cap value that it has already determined for one column (e.g., corresponding to the power cap value transmitted at 522) is substantially more advantageous than a different power cap setting identified based on processing consumption data from multiple columns.
[0068] The power management service 506 can be configured to enforce the set of power cap values that is determined to be more (or most) favorable to ensure that one or more circuit breakers in the data center do not trip. If the set of power cap values previously provided to the PDU 504 is determined to be less favorable than the set of power cap values calculated in 530, the power management service 450 can immediately send data to the PDU 504 to cause the PDU 504 to send an indication to the device 502 to immediately begin power capping based on the previously allocated power cap. Alternatively, the power management service 506 may take no further action, allowing a timer in the PDU 504 to expire.
[0069] Alternatively, if the power management service 450 determines that a new set of power caps is more (or most) advantageous from a power management perspective, the power management service 450 may send data to the power controller 448 to cancel the timer. The previously allocated power caps may be deleted by instruction from the power controller 448 (e.g., sent in response to canceling the timer), or they may time out by default according to a predefined period. The power management service 450 can send the new power caps in appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was sent) with an indication that the power caps should be immediately enforced by the corresponding devices. These PDUs can send the power caps to the receiving devices with instructions to immediately initiate power capping operations (including, for example, monitoring consumption with respect to the power cap values, determining whether the caps are based on the power cap values and the device's current consumption, and capping the device's operation or leaving such operation unconstrained based on the determination).
[0070] If the power cap value calculated at 530 is determined to be more favorable from a power management perspective than the power cap value calculated at 518 and ultimately stored at 528 in the device 502, method 500 proceeds to 532, where the power cap value calculated at 530 may be transmitted to a PDU 508 (e.g., any of the PDUs 508 that manage the device to which the power cap value is associated). In some embodiments, the power management service 506 may provide an indication that the power cap value transmitted at 532 should be implemented immediately.
[0071] In response to receiving the power cap values and indications at 532, the PDU 508 distributing power to the devices to which the power cap values are associated can transmit the power cap values and indications to the devices to which the power cap values are associated. At 536, the receiving device among the devices 512 can perform a power capping operation without delay based on receiving the indication, which includes 1) determining whether to limit / constrain operation based on the current consumption of a given device as compared to the power cap value provided for that device, and 2) limiting / constraining or not limiting / constraining the power consumption of that device within the range.
[0072] In some embodiments, the power management service 506 can transmit data at 522 to cancel the timer and / or power cap value transmitted based, at least in part, on determining that the power cap value calculated at 530 is more favorable (or most favorable) than that calculated at 518. In some embodiments, the PDU 504 can be configured to cancel the timer at 540. In some embodiments, the PDU 504 can transmit data at 542 that causes the device 502 to delete the power cap value from memory at 544.
[0073] At 546, in the situation where a more favorable set of power cap values is not found as described above or the power management service 506 has not sent a cancellation as described above in connection with 538, the PDU 504 can determine that the timer started at 524 has expired (e.g., the duration corresponding to the timing value provided at 522 has elapsed). As a result, the PDU 504 can send an indication to the device 502 at 548 instructing the device 502 to initiate power capping based on the power cap values stored at the device 502 at 528. By receiving this indication at 548, the device 502 can initiate power capping at 550 according to the power cap values stored at 528.
[0074] In some embodiments, the PDU 504 can be configured to cancel the timer at any suitable time after the timer is started at 524 if the PDU 504 determines that the device 502 has reduced its power consumption. In some embodiments, this can cause the PDU 504 to send data to the device 502 to cause the device 502 to discard the previously stored power cap value from local memory.
[0075] In some embodiments, the power cap value implemented in any suitable device may be allowed to time out or may be replaced at any suitable time with a power cap value calculated by the power management service 506. In some embodiments, the power management service 506 may send a power cap cancellation or replacement to any suitable device via the corresponding rack PDU. If a power cap cancellation is received, the device may delete the previously stored power cap value, allowing the device to resume unconstrained operation. If a replacement power cap value is received, the device may store the new power cap value and implement the new power cap value immediately or when instructed by its rack PDU. In some embodiments, being instructed by its PDU to implement a new power cap value may cause a device (e.g., power controller 446) of the devices to replace the power cap value previously utilized to implement the power cap setting with the new power cap value that the device is instructed to implement.
[0076] Any suitable operations of method 500 may be performed continuously and at any suitable time to manage excess consumption (e.g., a situation in which the consumption of a server corresponding to a column-level device exceeds the column-level device's budgeted power) to enable more efficient use of previously unused power while avoiding power outages due to circuit breaker tripping. It should be understood that similar operations may be performed by power management service 506 with respect to any suitable level of power distribution hierarchy 300. While examples have been provided with respect to monitoring power consumption corresponding to column-level devices (e.g., devices represented by level 3 nodes of power distribution hierarchy 300), similar operations may be performed with respect to any suitable level of higher-level devices (e.g., components corresponding to any of levels 2-5 of power distribution hierarchy 300).
[0077] FIG. 6 illustrates a flow diagram illustrating an example method 600 for training a machine learning model to determine the likelihood that the total power consumption of downstream devices (e.g., device 502 of FIG. 2) will exceed a budget threshold corresponding to an upstream device (e.g., a column-level device from which power is distributed to device 502), according to at least one embodiment.
[0078] In some embodiments, the model 602 can be trained (e.g., by the power management service 506 of FIG. 5 or another device) using any suitable machine learning algorithm (e.g., supervised, unsupervised, etc.) and any suitable number of training datasets (e.g., training data 608). A supervised machine learning algorithm refers to a machine learning task that involves learning an inferred function that maps inputs to outputs based on a labeled training dataset in which example input / output pairs are known. An unsupervised machine learning algorithm refers to a set of algorithms used to analyze and cluster unlabeled datasets (e.g., unlabeled data 610). These algorithms are configured to identify patterns or groupings of data without requiring human intervention. In some embodiments, any suitable number of models 602 can be trained during the training phase 604.
[0079] The model 602 may include any suitable number of models. The model 602 may be trained from training data 608, described below, to identify the likelihood that consumption data associated with a downstream device (e.g., a device to which power is distributed via a column-level device) will breach a budget threshold (e.g., a power budget amount) associated with an upstream device (e.g., a higher-level device such as a column-level device). In some embodiments, the model 602 may be configured to determine and output one or more predicted consumption amounts likely to occur in the future, and a determination of whether the predicted amount will breach a budget threshold (e.g., a power budget amount associated with an upstream device) may be made by the power management service 450 of FIG. 4.
[0080] As a non-limiting example, at least one of the models 602 can be trained during a training phase 604 using a supervised learning algorithm and labeled data 606 to identify a likelihood value and / or one or more predicted consumption values (e.g., predicted aggregate / cumulative consumption for one or more devices). The likelihood value can be a binary indicator of the degree of likelihood, a percentage, a confidence value, etc. The likelihood value can be a binary indicator indicating whether a particular budget amount is likely or unlikely to be breached, or the likelihood value can indicate the likelihood that a particular predicted consumption value will become a reality in the future. The labeled data 606 can be any suitable portion of potential training data (e.g., training data 608) that can be used to train various models against the likelihood values. The labeled data 606 can include any suitable number of examples of historical consumption data corresponding to devices and / or racks to which power passing through a row-level device is distributed. In some embodiments, the labeled data 606 can include labels that identify known likelihood values. The labeled data 606 can be used to learn a model (e.g., an inferred function) that maps input (e.g., one or more instances of historical consumption data corresponding to one or more devices) to output (e.g., the likelihood that the consumption of one or more devices will exceed a certain amount). In some embodiments, the amounts used can be included in the training data 608 and included as inputs. In some embodiments, the model 602 can provide as output a likelihood or confidence value indicating the amount by which the cumulative consumption from the perspective of a given component (e.g., a column-level device) is expected to change and the likelihood / confidence corresponding to that amount. In some embodiments, the power management service 506 of FIG. 5 can determine that an output 612 including the amount and / or likelihood value indicates that the consumption of a downstream device associated with a column-level device is likely to breach the budgeted power of that column-level device (e.g., exceed a threshold amount).
[0081] The model 602, and the various types of models described above, may include any suitable number of models trained using unsupervised learning techniques to identify the likelihood / confidence that a downstream device associated with a column-level device will breach the power budget associated with the column-level device. Alternatively, unsupervised learning techniques may be utilized to identify the amount by which the downstream device's consumption is predicted to increase (and possibly the corresponding likelihood / confidence). An unsupervised machine learning algorithm is configured to learn patterns from untagged data. In some embodiments, the training phase 604 may utilize an unsupervised machine learning algorithm to generate one or more models. For example, the training data 608 may include unlabeled data 610 (e.g., historical consumption data instances corresponding to devices and / or racks of devices receiving power through a given column-level device). The unlabeled data 610 may be utilized with an unsupervised learning algorithm that segments entries in the unlabeled data 610 into groups. The unsupervised learning algorithm may be configured to cluster similar entries into common groups. Examples of unsupervised learning algorithms include clustering methods such as k-means clustering, DBScan, etc. In some embodiments, the unlabeled data 610 can be clustered with the labeled data 606 so that unlabeled instances in a given group can be assigned the same label as other labeled instances in that group.
[0082] In some embodiments, any suitable portion of the training data 608 can be utilized to train the model 602 during the training phase 604. For example, 70% of the labeled data 606 and / or unlabeled data 610 can be utilized to train the model 602. Once trained, or at any suitable point in time, the model 602 can be evaluated to assess its quality (e.g., the accuracy of the output 612 with respect to the label corresponding to the labeled data 606). By way of example, a portion of the examples of the labeled data 606 and / or unlabeled data 610 can be utilized as input to the model 602 to generate the output 612. By way of example, an example of the labeled data 606 may be provided as input, and the corresponding output (e.g., output 612) may be compared with a label already known to be associated with that example. If a portion of the output (e.g., a label) matches the example label, that portion of the output can be deemed accurate. Any suitable number of labeled examples can be utilized, and the number of correct labels can be compared to the total number of examples provided (and / or the total number of previously identified labels) to determine an accuracy value for a given model that quantifies the model's accuracy. For example, if 90 out of 100 input examples produce output labels that match previously known label examples, the model being evaluated can be determined to be 90% accurate.
[0083] In some embodiments, as the model 602 is utilized with subsequent inputs, subsequent outputs generated by the model 602 can be added to the corresponding inputs and used to retrain and / or update the model 602 in 616. In some embodiments, an example may not be used to retrain or update the model until a feedback procedure 614 is performed. In the feedback procedure 614, an example (e.g., an example including one or more historical consumption data instances corresponding to one or more devices and / or racks) and the corresponding output generated for that example by one of the models 602 are presented to a user, who identifies whether the generated output (e.g., quantity and / or likelihood confidence value) is correct for a given example.
[0084] The training process (eg, method 600) shown in FIG. 6 can be performed any suitable number of times, at any suitable intervals and / or according to any suitable schedule, so that the accuracy of model 602 improves over time.
[0085] In some embodiments, any suitable number and / or combination of models 602 may be used to determine the output. In some embodiments, power management service 450 may utilize any suitable combination of outputs provided by models 602 to determine whether a budget threshold for a given component (e.g., a column-level PDU) is likely to be breached (and / or likely to be breached by a certain amount). Thus, in some embodiments, power management service 450 may utilize models trained with any suitable combination of supervised and unsupervised learning algorithms.
[0086] 7 is a block diagram illustrating an example method for managing power, according to at least one embodiment. Method 700 may be performed by one or more components of system 400 of FIG. 4. By way of example, method 700 may be performed, at least in part, by power management service 406 of FIG. 4. The operations of method 700 may be performed in any suitable order. More or fewer operations than those shown in FIG. 7 may be included in method 7.
[0087] At 702, a plurality of components of an electric power system (e.g., system 400) may be identified. In some embodiments, the plurality of components may be arranged according to a power distribution hierarchy (e.g., power distribution hierarchy 300 of FIG. 3 ) comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, a node of the plurality of nodes of the power distribution hierarchy represents a corresponding component of the plurality of components. In some embodiments, a subset of nodes (e.g., nodes 314 and 316 of FIG. 3 ) in a first level (e.g., level 1) of the plurality of levels may be derived from a particular node (e.g., node 306 of FIG. 3 , corresponding to a power distribution unit) in a second level (e.g., column level) higher than the first level of the plurality of levels. In some embodiments, a set of lower-level components of the plurality of components, represented by a subset of nodes in the first level, receive power distributed through a higher-level component of the plurality of components, represented by a particular node in the second level.
[0088] At 704, the power consumption of a set of lower-level components represented by a subset of the first-level nodes may be monitored. In some embodiments, monitoring the power consumption may include receiving device and / or rack consumption data, as described above in connection with method 500 of FIG.
[0089] At 706, it may be determined, at least in part based on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. In some embodiments, the budget threshold may be deemed breached when the power consumption of the set of lower-level components exceeds the budgeted amount of power allocated to the higher-level component. In some embodiments, the budget threshold may be deemed breached when the power consumption of the set of lower-level components approaches (e.g., within a threshold) the budgeted amount of power allocated to the higher-level component. In some embodiments, determining that the budget threshold has been breached may utilize historical consumption data in the manner described above in connection with 516 of FIG. 5.
[0090] In response to determining that the power consumption of the set of lower-level components has breached the budget threshold associated with the higher-level component, a power cap value for the lower-level component may be transmitted at 708. Transmitting the power cap value enables the lower-level component to store the power cap value in memory while allowing the power consumption of each of the lower-level components to exceed the power cap value until a time period corresponding to the timing value expires.
[0091] 8 is a block diagram illustrating another example method 800 for managing power, according to at least one embodiment. Method 800 may be performed by one or more components of system 400 of FIG. 4. By way of example, method 800 may be performed, at least in part, by power management service 406 of FIG. 4. The operations of method 800 may be performed in any suitable order. Method 800 may include more or fewer operations than those shown in FIG. 8.
[0092] At 802, a plurality of components of an electric power system (e.g., system 400) can be identified. In some embodiments, the plurality of components are arranged according to a power distribution hierarchy (e.g., power distribution hierarchy 300 of FIG. 3 ) comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, a node of the plurality of nodes of the power distribution hierarchy represents a corresponding component of the plurality of components. In some embodiments, a subset of nodes (e.g., nodes 314 and 316 of FIG. 3 ) in a first level (e.g., level 1) of the plurality of levels can be derived from a particular node (e.g., node 306 of FIG. 3 , corresponding to a power distribution unit) in a second level (e.g., column level) higher than the first level of the plurality of levels. In some embodiments, a set of lower-level components of the plurality of components, represented by a subset of nodes in the first level, receive power distributed through a higher-level component of the plurality of components, represented by a particular node in the second level.
[0093] At 804, the power consumption of a set of lower-level components represented by a subset of the first-level nodes may be monitored. In some embodiments, monitoring the power consumption may include receiving device and / or rack consumption data, as described above in connection with method 500 of FIG.
[0094] At 806, based at least in part on the monitoring, it may be determined that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. In some embodiments, the budget threshold may be deemed breached when the power consumption of the set of lower-level components exceeds the budgeted amount of power allocated to the higher-level component. In some embodiments, the budget threshold may be deemed breached when the power consumption of the set of lower-level components approaches (e.g., within a threshold) the budgeted amount of power allocated to the higher-level component. In some embodiments, determining that the budget threshold has been breached may utilize historical consumption data in the manner described above in connection with 516 of FIG. 5 .
[0095] At 808, in response to determining that the power consumption of the set of lower-level components has breached the budget threshold associated with the higher-level component, a timer corresponding to a timing value may be started. In some embodiments, expiration of the timer indicates expiration of a period corresponding to the timing value. A lower-level component (e.g., a server corresponding to device 404 in FIG. 4) may store a power cap value in memory for that period. In some embodiments, enforcement of the power cap value of the lower-level component is postponed until expiration of the timer.
[0096] 9-13 illustrate numerous example environments that may be hosted by components (e.g., device 404) of a data center. The environments illustrated in FIGS. 9-13 illustrate cloud computing, multi-tenant environments. As discussed above, cloud computing and / or multi-tenant environments, as well as other environments, may benefit from utilizing the power management techniques disclosed herein. These techniques enable data centers, including components implementing the environments described in connection with FIGS. 9-13, among others, to utilize the data center's power resources more efficiently than conventional techniques that leave large amounts of power unused.
[0097] Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, cloud computing providers can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, IaaS providers can also offer various services pertaining to those infrastructure components (example services include billing software, monitoring software, logging software, load balancing software, clustering software, etc.). Accordingly, these services can be policy-driven, so IaaS users can implement policies that drive load balancing to maintain application availability and performance.
[0098] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on the VMs. The customer can then use the provider's services to perform a variety of functions, including distributing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0099] In most cases, the cloud computing model requires the participation of a cloud provider, which can be, but is not required to be, a third-party service that specializes in providing (e.g., offering, renting, selling) IaaS. An entity can also choose to deploy a private cloud and become its own provider of infrastructure services.
[0100] In some examples, IaaS deployment is the process of placing a new application, or a new version of an application, onto a prepared application server, etc. IaaS deployment may also include the process of preparing the server (e.g., installing libraries, daemons, etc.), which is often managed by the cloud provider below the hypervisor layer (e.g., server, storage, network hardware, and virtualization). Thus, the customer may be responsible for handling the deployment of the (OS), middleware, and / or application (e.g., in self-service virtual machines (e.g., that can be spun up on demand) etc.
[0101] In some examples, IaaS provisioning may refer to obtaining the computers or virtual hosts to be used and also installing the necessary libraries or services on those computers or virtual hosts. In most cases, deployment does not include provisioning, which must be performed first.
[0102] In some cases, IaaS provisioning presents two distinct challenges. First, there is the initial challenge of provisioning the initial set of infrastructure before anything is running. Second, there is the challenge of evolving the existing infrastructure once everything is provisioned (e.g., adding new services, modifying services, removing services, etc.). In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., what resources depend on what and how each works together) can be described declaratively. In some cases, once the topology is defined, workflows can be generated to create and / or manage the different components described in the configuration files.
[0103] In some examples, the infrastructure may have many interconnected elements. For example, there may be one or more virtual private clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there may also be one or more inbound / outbound traffic group rules and one or more virtual machines (VMs) provisioned to define how inbound and / or outbound traffic on the network is configured. Other infrastructure elements, such as load balancers, databases, etc., may also be provisioned. The infrastructure may evolve over time as more infrastructure elements are desired and / or added.
[0104] In some cases, continuous deployment techniques can be employed to enable deployment of infrastructure code across various virtual computing environments. Additionally, the described techniques can enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, but often many, different production environments (e.g., across various different geographic locations, sometimes across the globe). However, in some examples, the infrastructure onto which the code will be deployed must first be set up. In some cases, provisioning can be done manually, and provisioning tools may be used to provision the resources and / or deployment tools may be used to deploy the code once the infrastructure has been provisioned.
[0105] 9 is a block diagram 900 illustrating an example IaaS architecture pattern according to at least one embodiment. A service operator 902 can be communicatively coupled to a secure host tenancy 904, which can include a virtual cloud network (VCN) 906 and a secure host subnet 908. In some examples, the service operator 902 can employ one or more client computing devices, which can be portable handheld devices (e.g., iPhone®, mobile phone, iPad®, computing tablet, personal digital assistant (PDA)), or wearable devices (e.g., Google® Glass head-mounted display) running software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and supporting the Internet, email, short message service (SMS), Blackberry®, or other communication protocols. Alternatively, the client computing devices may be general-purpose personal computers, including, by way of example, personal and / or laptop computers running various versions of the Microsoft Windows operating system, the Apple Macintosh operating system, and / or the Linux operating system. The client computing devices may also be workstation computers running any of a variety of commercially available UNIX or UNIX-like operating systems, including, without limitation, various GNU / Linux operating systems such as Google Chrome OS.Alternatively or additionally, the client computing device may be any other electronic device, such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a Kinect® gesture input device), and / or a personal messaging device, that can communicate over a network with access to the VCN 906 and / or the Internet.
[0106] The VCN 906 can include a local peering gateway (LPG) 910, which can be communicatively coupled to a secure shell (SSH) VCN 912 via the LPG 910 included in the SSH VCN 912. The SSH VCN 912 can include an SSH subnet 914, which can be communicatively coupled to a control plane VCN 916 via the LPG 910 included in the control plane VCN 916. The SSH VCN 912 can also be communicatively coupled to a data plane VCN 918 via the LPG 910. The control plane VCN 916 and the data plane VCN 918 can be included in a service tenancy 919, which can be owned and / or operated by the IaaS provider.
[0107] The control plane VCN 916 may include a control plane demilitarized zone (DMZ) tier 920 that acts as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers have limited responsibility and may help mitigate breaches. Additionally, the DMZ tier 920 may include one or more load balancer (LB) subnets 922, a control plane app tier 924 that may include an app subnet 926, and a control plane data tier 928 that may include a database (DB) subnet 930 (e.g., a front-end DB subnet and / or a back-end DB subnet). The LB subnet 922 included in the control plane DMZ tier 920 may be communicatively coupled to the app subnet 926 included in the control plane app tier 924 and to an Internet gateway 934 that may be included in the control plane VCN 916, and the app subnet 926 may be communicatively coupled to the DB subnet 930, a service gateway 936, and a network address translation (NAT) gateway 938 that are included in the control plane data tier 928. The control plane VCN 916 may include a service gateway 936 and a NAT gateway 938.
[0108] The control plane VCN 916 may include a data plane mirrored app tier 940, which may include an app subnet 926. The app subnet 926 included in the data plane mirrored app tier 940 may include a virtual network interface controller (VNIC) 942 on which a compute instance 944 can run. The compute instance 944 can communicatively couple the app subnet 926 of the data plane mirrored app tier 940 to the app subnet 926, which may be included in the data plane app tier 946.
[0109] The data plane VCN 918 may include a data plane app layer 946, a data plane DMZ layer 948, and a data plane data layer 950. The data plane DMZ layer 948 may include a LB subnet 922 that can be communicatively coupled to an app subnet 926 of the data plane app layer 946 and an internet gateway 934 of the data plane VCN 918. The app subnet 926 can be communicatively coupled to a service gateway 936 of the data plane VCN 918 and a NAT gateway 938 of the data plane VCN 918. The data plane data layer 950 may also include a DB subnet 930 that can be communicatively coupled to the app subnet 926 of the data plane app layer 946.
[0110] The internet gateway 934 of the control plane VCN 916 and the internet gateway 934 of the data plane VCN 918 can be communicatively coupled to a metadata management service 952, which can be communicatively coupled to the public internet 954. The public internet 954 can be communicatively coupled to a NAT gateway 938 of the control plane VCN 916 and the NAT gateway 938 of the data plane VCN 918. The service gateway 936 of the control plane VCN 916 and the service gateway 936 of the data plane VCN 918 can be communicatively coupled to cloud services 956.
[0111] In some examples, the service gateway 936 of the control plane VCN 916 or the service gateway 936 of the data plane VCN 918 can make application programming interface (API) calls to the cloud services 956 without traversing the public internet 954. The API calls from the service gateway 936 to the cloud services 956 can be one-way. That is, the service gateway 936 can make an API call to the cloud services 956, and the cloud services 956 can send the requested data to the service gateway 936. However, the cloud services 956 cannot initiate API calls to the service gateway 936.
[0112] In some examples, secure host tenancy 904 can be directly connected to service tenancy 919, which may otherwise be isolated. Secure host subnet 908 can communicate with SSH subnet 914 through LPG 910, which may enable bidirectional communication in otherwise isolated systems. By connecting secure host subnet 908 to SSH subnet 914, secure host subnet 908 can access other entities in service tenancy 919.
[0113] The control plane VCN 916 can enable users of the service tenancy 919 to set up or otherwise provision desired resources. The desired resources provisioned in the control plane VCN 916 can be deployed or otherwise used in the data plane VCN 918. In some examples, the control plane VCN 916 can be isolated from the data plane VCN 918, and the data plane mirror app layer 940 of the control plane VCN 916 can communicate with the data plane app layer 946 of the data plane VCN 918 via a VNIC 942, which can be included in the data plane mirror app layer 940 and the data plane app layer 946.
[0114] In some examples, a user or customer of the system can make a request, for example, a request for a create, read, update, or delete (CRUD) operation, via the public internet 654, which can communicate such a request to the metadata management service 952. The metadata management service 952 can communicate the request to the control plane VCN 916 via the internet gateway 934. The request can be received by the LB subnet 922 included in the control plane DMZ tier 920. The LB subnet 922 may determine that the request is valid, and in response to this determination, the LB subnet 922 can send the request to the app subnet 926 included in the control plane app tier 924. If the request is validated and requires a call to the public internet 954, the call to the public internet 954 can be sent to the NAT gateway 938, which can make the call to the public internet 954. Metadata that may be desired to be stored by the request can be stored in the DB subnet 930.
[0115] In some examples, the data plane mirror app layer 940 may facilitate direct communication between the control plane VCN 916 and the data plane VCN 918. For example, it may be desired that a change, update, or other suitable modification to the configuration be applied to resources included in the data plane VCN 918. The control plane VCN 916 can communicate directly with the resources included in the data plane VCN 918 via VNIC 942, thereby performing the change, update, or other suitable modification to the configuration on the resources.
[0116] In some embodiments, the control plane VCN 916 and the data plane VCN 918 can be included in the service tenancy 919. In this case, a user or customer of the system need not own or operate either the control plane VCN 916 or the data plane VCN 918. Instead, an IaaS provider can own or operate the control plane VCN 916 and the data plane VCN 918, both of which can be included in the service tenancy 919. This embodiment can enable network isolation, thereby preventing users or customers from interacting with other users' or customers' resources. This embodiment can also enable users or customers of the system to store databases privately without having to rely on the public internet 954, which may not have the desired level of threat protection for storage.
[0117] In another embodiment, the LB subnet 922 included in the control plane VCN 916 can be configured to receive signals from the service gateway 936. In this embodiment, the control plane VCN 916 and the data plane VCN 918 can be configured to be called by the IaaS provider's customers without calling the public internet 954. Customers of the IaaS provider may desire this embodiment because databases used by the customers can be stored in the service tenancy 919, which can be controlled by the IaaS provider and isolated from the public internet 954.
[0118] 10 is a block diagram 1000 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1002 (e.g., service operator 902 in FIG. 9 ) can be communicatively coupled to a secure host tenancy 1004 (e.g., secure host tenancy 904 in FIG. 9 ), which can include a virtual cloud network (VCN) 1006 (e.g., VCN 906 in FIG. 9 ) and a secure host subnet 1008 (e.g., secure host subnet 908 in FIG. 9 ). VCN 1006 can include a local peering gateway (LPG) 1010 (e.g., LPG 910 in FIG. 9 ), which can be communicatively coupled to a secure shell (SSH) VCN 1012 (e.g., SSH VCN 912 in FIG. 9 ) via the LPG 910 included in the SSH VCN 1012. The SSH VCN 1012 can include an SSH subnet 1014 (e.g., SSH subnet 914 in FIG. 9 ), and the SSH VCN 1012 can be communicatively coupled to a control plane VCN 1016 (e.g., control plane VCN 916 in FIG. 9 ) via an LPG 1010 included in the control plane VCN 1016. The control plane VCN 1016 can be included in a service tenancy 1019 (e.g., service tenancy 919 in FIG. 9 ), and the data plane VCN 1018 (e.g., data plane VCN 918 in FIG. 9 ) can be included in a customer tenancy 1021, which may be owned or operated by a user or customer of the system.
[0119] The control plane VCN 1016 may include a control plane DMZ tier 1020 (e.g., control plane DMZ tier 920 in FIG. 9 ) that may include a LB subnet 1022 (e.g., LB subnet 922 in FIG. 9 ), a control plane app tier 1024 (e.g., control plane app tier 924 in FIG. 9 ) that may include an app subnet 1026 (e.g., app subnet 926 in FIG. 9 ), and a control plane data tier 1028 (e.g., control plane data tier 928 in FIG. 9 ) that may include a database (DB) subnet 1030 (e.g., similar to DB subnet 930 in FIG. 9 ). The LB subnet 1022 included in the control plane DMZ tier 1020 can be communicatively coupled to an app subnet 1026 included in the control plane app tier 1024 and to an Internet gateway 1034 (e.g., Internet gateway 934 in FIG. 9 ) that can be included in the control plane VCN 1016, and the app subnet 1026 can be communicatively coupled to a DB subnet 1030 included in the control plane data tier 1028 and to a service gateway 1036 (e.g., service gateway 936 in FIG. 9 ) and a network address translation (NAT) gateway 1038 (e.g., NAT gateway 938 in FIG. 9 ). The control plane VCN 1016 can include the service gateway 1036 and the NAT gateway 1038.
[0120] The control plane VCN 1016 may include a data plane mirror app layer 1040 (e.g., data plane mirror app layer 940 of FIG. 9 ), which may include an app subnet 1026. The app subnet 1026 included in the data plane mirror app layer 1040 may include a virtual network interface controller (VNIC) 1042 (e.g., VNIC of 942) on which a compute instance 1044 (e.g., similar to compute instance 944 of FIG. 9 ) can run. The compute instance 1044 can facilitate communication between the app subnet 1026 of the data plane mirror app layer 1040 and the app subnet 1026, which may be included in the data plane app layer 1046 (e.g., data plane app layer 946 of FIG. 9 ), via the VNIC 1042 included in the data plane mirror app layer 1040 and the VNIC 1042 included in the data plane app layer 1046.
[0121] The internet gateway 1034 included in the control plane VCN 1016 can be communicatively coupled to a metadata management service 1052 (e.g., metadata management service 952 in FIG. 9 ), which can be communicatively coupled to the public internet 1054 (e.g., public internet 954 in FIG. 9 ). The public internet 1054 can be communicatively coupled to a NAT gateway 1038 included in the control plane VCN 1016. The service gateway 1036 included in the control plane VCN 1016 can be communicatively coupled to cloud services 1056 (e.g., cloud services 956 in FIG. 9 ).
[0122] In some examples, the data plane VCN 1018 can be included in the customer tenancy 1021. In this case, the IaaS provider can provide a control plane VCN 1016 for each customer, and the IaaS provider can set up a unique compute instance 1044 for each customer, which is included in the service tenancy 1019. Each compute instance 1044 can enable communication between the control plane VCN 1016, which is included in the service tenancy 1019, and the data plane VCN 1018, which is included in the customer tenancy 1021. The compute instance 1044 can enable resources provisioned in the control plane VCN 1016, which is included in the service tenancy 1019, to be deployed to or otherwise used in the data plane VCN 1018, which is included in the customer tenancy 1021.
[0123] In another example, a customer of the IaaS provider may have a database that resides in customer tenancy 1021. In this example, control plane VCN 1016 may include a data plane mirror app tier 1040, which may include app subnet 1026. The data plane mirror app tier 1040 may reside in data plane VCN 1018, and the data plane mirror app tier 1040 may not reside in the data plane VCN 1018. That is, while the data plane mirror app tier 1040 has access to customer tenancy 1021, the data plane mirror app tier 1040 may not reside in the data plane VCN 1018 or be owned and operated by the IaaS provider's customer. The data plane mirror app tier 1040 may be configured to make calls to the data plane VCN 1018, but may not be configured to make calls to any entities included in the control plane VCN 1016. A customer may wish to deploy or otherwise use resources in the data plane VCN 1018 that have been provisioned in the control plane VCN 1016, and the data plane mirror app layer 1040 can facilitate the desired deployment or other use of the customer's resources.
[0124] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 1018. In this embodiment, the customer can determine what the data plane VCN 1018 can access, and the customer can restrict access from the data plane VCN 1018 to the public internet 1054. The IaaS provider may not be able to filter or otherwise control access from the data plane VCN 1018 to any external networks or databases. Applying filters and controls to the data plane VCN 1018 that the customer includes in the customer tenancy 1021 can help isolate the data plane VCN 1018 from other customers and from the public internet 1054.
[0125] In some embodiments, the service gateway 1036 can call cloud services 1056 to access services that may not reside on the public internet 1054, on the control plane VCN 1016, or on the data plane VCN 1018. The connection between the cloud services 1056 and the control plane VCN 1016 or the data plane VCN 1018 may not be live or continuous. The cloud services 1056 may reside on different networks owned or operated by the IaaS provider. The cloud services 1056 may be configured to receive calls from the service gateway 1036 and may not be configured to receive calls from the public internet 1054. Some cloud services 1056 may be isolated from other cloud services 1056, and the control plane VCN 1016 may be isolated from cloud services 1056 that may not be in the same region as the control plane VCN 1016. For example, the control plane VCN 1016 may be located in “Region 1,” and cloud service “Deployment 9” may be located in Region 1 and “Region 2.” If a call to deployment 9 is made by a service gateway 1036 included in a control plane VCN 1016 located in region 1, the call can be sent to deployment 9 in region 1. In this example, control plane VCN 1016 or deployment 9 in region 1 may not be communicatively coupled or otherwise in communication with deployment 9 in region 2.
[0126] 11 is a block diagram 1100 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1102 (e.g., service operator 902 in FIG. 9 ) can be communicatively coupled to a secure host tenancy 1104 (e.g., secure host tenancy 904 in FIG. 9 ), which can include a virtual cloud network (VCN) 1106 (e.g., VCN 906 in FIG. 9 ) and a secure host subnet 1108 (e.g., secure host subnet 908 in FIG. 9 ). VCN 1106 can include an LPG 1110 (e.g., LPG 910 in FIG. 9 ), which can be communicatively coupled to an SSH VCN 1112 (e.g., SSH VCN 912 in FIG. 9 ) via an LPG 1110 included in SSH VCN 1112. SSH VCN 1112 can include an SSH subnet 1114 (e.g., SSH subnet 914 in FIG. 9 ), and SSH VCN 1112 can be communicatively coupled to a control plane VCN 1116 (e.g., control plane VCN 916 in FIG. 9 ) via an LPG 1110 included in the control plane VCN 1116, and to a data plane VCN 1118 (e.g., data plane 918 in FIG. 9 ) via an LPG 1110 included in the data plane VCN 1118. The control plane VCN 1116 and the data plane VCN 1118 can be included in a service tenancy 1119 (e.g., service tenancy 919 in FIG. 9 ).
[0127] The control plane VCN 1116 may include a control plane DMZ tier 1120 (e.g., control plane DMZ tier 920 of FIG. 9 ) that may include a load balancer (LB) subnet 1122 (e.g., LB subnet 922 of FIG. 9 ), a control plane app tier 1124 (e.g., control plane app tier 924 of FIG. 9 ) that may include an app subnet 1126 (e.g., similar to app subnet 926 of FIG. 9 ), and a control plane data tier 1128 (e.g., control plane data tier 928 of FIG. 9 ) that may include a DB subnet 1130. The LB subnet 1122 included in the control plane DMZ tier 1120 can be communicatively coupled to an app subnet 1126 included in the control plane app tier 1124 and to an Internet gateway 1134 (e.g., Internet gateway 934 in FIG. 9 ) that can be included in the control plane VCN 1116, and the app subnet 1126 can be communicatively coupled to a DB subnet 1130 included in the control plane data tier 1128 and to a service gateway 1136 (e.g., service gateway in FIG. 9 ) and a network address translation (NAT) gateway 1138 (e.g., NAT gateway 938 in FIG. 9 ). The control plane VCN 1116 can include the service gateway 1136 and the NAT gateway 1138.
[0128] Data plane VCN 1118 may include a data plane app layer 1146 (e.g., data plane app layer 946 in FIG. 9 ), a data plane DMZ layer 1148 (e.g., data plane DMZ layer 948 in FIG. 9 ), and a data plane data layer 1150 (e.g., data plane data layer 950 in FIG. 9 ). Data plane DMZ layer 1148 may include LB subnet 1122, which may be communicatively coupled to trusted app subnet 1160 and untrusted app subnet 1162 of data plane app layer 1146, which are included in data plane VCN 1118, as well as to Internet gateway 1134. Trusted app subnet 1160 may be communicatively coupled to service gateway 1136, which is included in data plane VCN 1118, NAT gateway 1138, which is included in data plane VCN 1118, and DB subnet 1130, which is included in data plane data layer 1150. The untrusted app subnet 1162 can be communicatively coupled to a service gateway 1136 included in the data plane VCN 1118 and to a DB subnet 1130 included in the data plane data layer 1150. The data plane data layer 1150 can include a DB subnet 1130 that can be communicatively coupled to a service gateway 1136 included in the data plane VCN 1118.
[0129] The untrusted app subnet 1162 may include one or more primary VNICs 1164(1)-(N), which may be communicatively coupled to tenant virtual machines (VMs) 1166(1)-(N). Each tenant VM 1166(1)-(N) may be communicatively coupled to a respective app subnet 1167(1)-(N), which may be included in a respective container egress VCN 1168(1)-(N), which may be included in a respective customer tenancy 1170(1)-(N). Each secondary VNIC 1172(1)-(N) may facilitate communication between the untrusted app subnet 1162 included in the data plane VCN 1118 and the app subnet included in the container egress VCN 1168(1)-(N). Each container egress VCN 1168(1)-(N) may include a NAT gateway 1138, which may be communicatively coupled to the public internet 1154 (e.g., public internet 954 in FIG. 9 ).
[0130] The internet gateway 1134 included in the control plane VCN 1116 and the internet gateway 1134 included in the data plane VCN 1118 can be communicatively coupled to a metadata management service 1152 (e.g., metadata management system 952 of FIG. 9 ), which can be communicatively coupled to the public internet 1154. The public internet 1154 can be communicatively coupled to a NAT gateway 1138 included in the control plane VCN 1116 and the NAT gateway 1138 included in the data plane VCN 1118. The service gateway 1136 included in the control plane VCN 1116 and the service gateway 1136 included in the data plane VCN 1118 can be communicatively coupled to cloud services 1156.
[0131] In some embodiments, data plane VCN 1118 can be integrated with customer tenancies 1170. This integration can be useful or desirable for an IaaS provider's customers in some cases, such as when they may want support when running code. A customer may provide code to run that may be disruptive, may communicate with other customer resources, or may otherwise cause undesirable effects. In response, the IaaS provider can determine whether to run the code that the customer has provided to the IaaS provider.
[0132] In some examples, a customer of an IaaS provider can grant temporary network access to the IaaS provider and request a function to be attached to the data plane app layer 1146. The code that executes the function can run in VMs 1166(1)-(N), and the code may not be configured to run anywhere else on the data plane VCN 1118. Each VM 1166(1)-(N) can be connected to one customer tenancy 1170. Each container 1171(1)-(N) contained in a VM 1166(1)-(N) can be configured to run code. In this case, there can be double isolation (e.g., containers 1171(1)-(N) can execute code, and containers 1171(1)-(N) can be contained in VMs 1166(1)-(N) that are at least in the untrusted app subnet 1162). This can help prevent erroneous or otherwise unwanted code from damaging the IaaS provider's network or damaging a different customer's network. Containers 1171(1)-(N) can be communicatively coupled to customer tenancy 1170 and can be configured to send or receive data from customer tenancy 1170. Containers 1171(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 1118. When the code execution is complete, the IaaS provider can kill or otherwise discard containers 1171(I)-(N).
[0133] In some embodiments, trusted app subnet 1160 can execute code that can be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 1160 can be communicatively coupled to DB subnet 1130 and configured to perform CRUD operations on DB subnet 1130. Untrusted app subnet 1162 can be communicatively coupled to DB subnet 1130, but in this embodiment, the untrusted app subnet can be configured to perform read operations within DB subnet 1130. Containers 1171(1)-(N) that can be included in each customer's VMs 1166(1)-(N) and that can execute code from the customer need not be communicatively coupled to DB subnet 1130.
[0134] In other embodiments, the control plane VCN 1116 and the data plane VCN 1118 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 1116 and the data plane VCN 1118. However, communication may occur indirectly in at least one manner. An IaaS provider may establish an LPG 1110 that can facilitate communication between the control plane VCN 1116 and the data plane VCN 1118. In another example, the control plane VCN 1116 or the data plane VCN 1118 can make a call to a cloud service 1156 through the service gateway 1136. For example, a call from the control plane VCN 1116 to the cloud service 1156 may include a request for a service that can communicate with the data plane VCN 1118.
[0135] 12 is a block diagram 1200 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1202 (e.g., service operator 902 in FIG. 9 ) can be communicatively coupled to a secure host tenancy 1204 (e.g., secure host tenancy 904 in FIG. 9 ), which can include a virtual cloud network (VCN) 1206 (e.g., VCN 906 in FIG. 9 ) and a secure host subnet 1208 (e.g., secure host subnet 908 in FIG. 9 ). VCN 1206 can include an LPG 1210 (e.g., LPG 910 in FIG. 9 ), which can be communicatively coupled to an SSH VCN 1212 (e.g., SSH VCN 912 in FIG. 9 ) via the LPG 1210 included in the SSH VCN 1212. SSH VCN 1212 can include SSH subnet 1214 (e.g., SSH subnet 914 in FIG. 9 ), and SSH VCN 1212 can be communicatively coupled to control plane VCN 1216 (e.g., control plane VCN 916 in FIG. 9 ) via LPG 1210 included in control plane VCN 1216 and to data plane VCN 1218 (e.g., data plane 918 in FIG. 9 ) via LPG 1210 included in data plane VCN 1218. Control plane VCN 1216 and data plane VCN 1218 can be included in service tenancy 1219 (e.g., service tenancy 919 in FIG. 9 ).
[0136] The control plane VCN 1216 may include a control plane DMZ layer 1220 (e.g., control plane DMZ layer 920 of FIG. 9 ) that may include a LB subnet 1222 (e.g., LB subnet 922 of FIG. 9 ), a control plane app layer 1224 (e.g., control plane app layer 924 of FIG. 9 ) that may include an app subnet 1226 (e.g., app subnet 926 of FIG. 9 ), and a control plane data layer 1228 (e.g., control plane data layer 928 of FIG. 9 ) that may include a DB subnet 1230 (e.g., DB subnet 1130 of FIG. 11 ). LB subnet 1222 included in control plane DMZ tier 1220 can be communicatively coupled to app subnet 1226 included in control plane app tier 1224 and to an Internet gateway 1234 (e.g., Internet gateway 934 in FIG. 9 ) that can be included in control plane VCN 1216, and app subnet 1226 can be communicatively coupled to DB subnet 1230 included in control plane data tier 1228 and to service gateway 1236 (e.g., service gateway in FIG. 9 ) and network address translation (NAT) gateway 1238 (e.g., NAT gateway 938 in FIG. 9 ). Control plane VCN 1216 can include service gateway 1236 and NAT gateway 1238.
[0137] Data plane VCN 1218 may include a data plane app layer 1246 (e.g., data plane app layer 946 in FIG. 9 ), a data plane DMZ layer 1248 (e.g., data plane DMZ layer 948 in FIG. 9 ), and a data plane data layer 1250 (e.g., data plane data layer 950 in FIG. 9 ). Data plane DMZ layer 1248 may include LB subnet 1222, which may be communicatively coupled to trusted app subnet 1260 (e.g., trusted app subnet 1160 in FIG. 11 ) and untrusted app subnet 1262 (e.g., untrusted app subnet 1162 in FIG. 11 ) of data plane app layer 1246 included in data plane VCN 1218, as well as to Internet gateway 1234. Trusted app subnet 1260 can be communicatively coupled to service gateway 1236 included in data plane VCN 1218, NAT gateway 1238 included in data plane VCN 1218, and DB subnet 1230 included in data plane data layer 1250. Untrusted app subnet 1262 can be communicatively coupled to service gateway 1236 included in data plane VCN 1218 and DB subnet 1230 included in data plane data layer 1250. Data plane data layer 1250 can include DB subnet 1230 that can be communicatively coupled to service gateway 1236 included in data plane VCN 1218.
[0138] The untrusted app subnet 1262 may include primary VNICs 1264(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 1266(1)-(N) that reside within the untrusted app subnet 1262. Each tenant VM 1266(1)-(N) can execute code in a respective container 1267(1)-(N) and can be communicatively coupled to an app subnet 1226 that can be included in a data plane app tier 1246 that can be included in a container egress VCN 1268. Each secondary VNIC 1272(1)-(N) can facilitate communication between the untrusted app subnet 1262 included in the data plane VCN 1218 and the app subnet included in the container egress VCN 1268. The container egress VCN may include a NAT gateway 1238 that can be communicatively coupled to the public internet 1254 (e.g., public internet 954 in FIG. 9 ).
[0139] The internet gateway 1234 included in the control plane VCN 1216 and the internet gateway 1234 included in the data plane VCN 1218 can be communicatively coupled to a metadata management service 1252 (e.g., metadata management system 952 of FIG. 9 ), which can be communicatively coupled to the public internet 1254. The public internet 1254 can be communicatively coupled to a NAT gateway 1238 included in the control plane VCN 1216 and the NAT gateway 1238 included in the data plane VCN 1218. The service gateway 1236 included in the control plane VCN 1216 and the service gateway 1236 included in the data plane VCN 1218 can be communicatively coupled to cloud services 1256.
[0140] In some examples, the pattern illustrated by the architecture of block diagram 1200 in FIG. 12 may be considered an exception to the pattern illustrated by the architecture of block diagram 1100 in FIG. 11 and may be desirable for an IaaS provider's customers when the IaaS provider cannot communicate directly with the customers (e.g., in disconnected regions). The customers can access each of the containers 1267(1)-(N) contained in each customer's VMs 1266(1)-(N) in real time. The containers 1267(1)-(N) can be configured to call each of the secondary VNICs 1272(1)-(N) contained in the app subnet 1226 of the data plane app tier 1246, which can be contained in the container egress VCN 1268. The secondary VNICs 1272(1)-(N) can send calls to a NAT gateway 1238, which can send calls to the public Internet 1254. In this example, containers 1267(1)-(N) that a customer can access in real time can be isolated from control plane VCN 1216 and isolated from other entities included in data plane VCN 1218. Containers 1267(1)-(N) can also be isolated from resources of other customers.
[0141] In another example, a customer can use containers 1267(1)-(N) to invoke cloud service 1256. In this example, the customer can execute code in containers 1267(1)-(N) that requests a service from cloud service 1256. Containers 1267(1)-(N) can send the request to secondary VNICs 1272(1)-(N), which can send the request to a NAT gateway that can send the request to public internet 1254. Public internet 1254 can send the request to LB subnet 1222, which is included in control plane VCN 1216, via internet gateway 1234. In response to determining that the request is valid, LB subnet 1226 can send the request to app subnet 1226, which can send the request to cloud service 1256 via service gateway 1236.
[0142] It should be understood that the IaaS architectures 900, 1000, 1100, 1200 shown in the figures may have components other than those shown. Additionally, the illustrated embodiments are only some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS system may have more or fewer components than those shown in the figures, may combine two or more components, or may have a different configuration or arrangement of components.
[0143] In some embodiments, the IaaS systems described herein may include a suite of application, middleware, and database service offerings that are self-service, subscription-based, elastically scalable, reliable, highly available, and securely delivered to customers. One example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the present assignee.
[0144] 13 illustrates an example computer system 1300 upon which various embodiments may be implemented. System 1300 may be used to implement any of the computer systems described above. As shown, computer system 1300 includes a processing unit 1304 that communicates with multiple peripheral subsystems via a bus subsystem 1302. These peripheral subsystems may include a processing acceleration unit 1306, an I / O subsystem 1308, a storage subsystem 1318, and a communication subsystem 1324. Storage subsystem 1318 includes a tangible computer-readable storage medium 1322 and a system memory 1310.
[0145] Bus subsystem 1302 provides a mechanism for allowing the various components and subsystems of computer system 1300 to communicate with each other as intended. While bus subsystem 1302 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1302 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which may be implemented as a Mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0146] Processing unit 1304, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 1300. Processing unit 1304 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 1304 may be implemented as one or more independent processing units 1332 and / or 1334, with a single-core or multi-core processor included in each processing unit. In other embodiments, processing unit 1304 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0147] In various embodiments, the processing unit 1304 may execute various programs according to program code and may maintain multiple simultaneously executing programs or processes. At any given time, some or all of the program code to be executed may reside in the processor 1304 and / or in the storage subsystem 1318. Through suitable programming, the processor 1304 may provide the various functions described above. The computer system 1300 may further include a processing acceleration unit 1306, which may include a digital signal processor (DSP), a special purpose processor, etc.
[0148] The I / O subsystem 1308 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may include, for example, a motion sensing and / or gesture recognition device such as a Microsoft Kinect® motion sensor that enables a user to control and interact with an input device such as a Microsoft Xbox® 360 game controller through a natural user interface using gestures and spoken commands. User interface input devices may also include an eye gesture recognition device such as a Google Glass® blink detector that detects eye activity from a user (e.g., “blinking” during picture taking and / or menu selection) and translates the eye gesture as input to an input device (e.g., Google Glass®). Additionally, the user interface input devices may include a voice recognition sensing device that allows a user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0149] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser distance measuring devices, and eye-tracking devices. Furthermore, user interface input devices may include medical imaging input devices, such as computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, positron emission tomography (PET) scanners, and medical ultrasound scanners. User interface input devices may also include audio input devices, such as MIDI keyboards, digital musical instruments, and the like.
[0150] User interface output devices may include a display subsystem, indicator lights, or non-visual displays such as audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 1300 to a user or to another computer. For example, user interface output devices may include various display devices that visually convey text, graphics, and audio / video information, such as, but not limited to, monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, and modems.
[0151] Computer system 1300 may include a storage subsystem 1318 that provides a tangible, non-transitory, computer-readable storage medium for storing software and data structures that provide the functionality of embodiments described in this disclosure. The software may include programs, code modules, instructions, scripts, etc., that, when executed by one or more cores or processors of processing unit 1304, provide the functionality described above. Storage subsystem 1318 may also provide a repository for storing data used in accordance with the present disclosure.
[0152] 13, storage subsystem 1318 may include various components including system memory 1310, computer-readable storage medium 1322, and computer-readable storage medium reader 1320. System memory 1310 may store program instructions that are loadable and executable by processing unit 1304. System memory 1310 may also store data used during execution of the instructions and / or data generated during execution of the program instructions. A variety of different types of programs may be loaded into system memory 1310, including, but not limited to, client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), virtual machines, containers, etc.
[0153] The system memory 1310 may also store an operating system 1316. Examples of the operating system 1316 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS operating systems. In some implementations in which the computer system 1300 runs one or more virtual machines, the virtual machines, along with their guest operating systems (GOS), may be loaded into the system memory 1310 and executed by one or more processors or cores of the processing unit 1304.
[0154] The system memory 1310 may be configured differently depending on the type of computer system 1300. For example, the system memory 1310 may be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). Different types of RAM configurations can be provided, including static random access memory (SRAM), dynamic random access memory (DRAM), etc. In some implementations, the system memory 1310 may include a basic input / output system (BIOS), which contains the basic routines that help to transfer information between elements within the computer system 1300, such as during start-up.
[0155] Computer-readable storage media 1322 may represent remote, local, fixed, and / or removable storage devices, plus storage media, for temporarily and / or more permanently containing and storing computer-readable information used by computer system 1300, including instructions executable by processing unit 1304 of computer system 1300.
[0156] The computer-readable storage medium 1322 may include any suitable medium known or used in the art, including, but not limited to, storage media and communication media such as volatile and nonvolatile, removable and non-removable media, implemented in any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or other tangible computer-readable medium.
[0157] By way of example, the computer-readable storage medium 1322 may include a hard disk drive that reads from or writes to non-removable, non-volatile magnetic media, a magnetic disk drive that reads from or writes to removable, non-volatile magnetic disks, and an optical disk drive that reads from or writes to removable, non-volatile optical disks such as CD-ROMs, DVDs, and Blu-Ray® disks or other optical media. The computer-readable storage medium 1322 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD disks, digital video tapes, etc. The computer-readable storage media 1322 can also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, volatile memory-based SSDs such as magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 1300.
[0158] Machine-readable instructions executable by one or more processors or cores of the processing unit 1304 may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include physically tangible memory or storage devices, including volatile and / or non-volatile memory storage devices. Examples of non-transitory computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard drives, floppy drives, removable memory drives (e.g., USB drives), or other types of storage devices.
[0159] The communications subsystem 1324 provides an interface to other computer systems and networks. The communications subsystem 1324 serves as an interface for receiving data from the computer system 1300 and transmitting data from the computer system 1300 to other systems. For example, the communications subsystem 1324 can enable the computer system 1300 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 1324 may include a radio frequency (RF) transceiver component for accessing a wireless voice and / or data network (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution), WiFi (IEEE 802.11 family of standards, or other mobile communications technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some embodiments, the communications subsystem 1324 can provide wired network connectivity (e.g., Ethernet) in addition to or instead of a wireless interface.
[0160] In some embodiments, the communications subsystem 1324 may also receive incoming communications in the form of structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc., on behalf of one or more users who may use the computer system 1300.
[0161] By way of example, the communications subsystem 1324 may be configured to receive data feeds 1326 in real time from users of social networks and / or other communications services, such as web feeds such as Twitter® feeds, Facebook® updates, Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party sources.
[0162] Additionally, the communications subsystem 1324 may be configured to receive data in the form of continuous data streams, which may include event streams 1328 of real-time events and / or event updates 1330, which may be continuous or infinite in nature with no apparent end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0163] The communications subsystem 1324 may also be configured to output structured and / or unstructured data feeds 1326, event streams 1328, event updates 1330, etc. to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1300.
[0164] The computer system 1300 may be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head-mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or any other data processing system.
[0165] Because the nature of computers and networks is constantly changing, the description of computer system 1300 shown in the figure is intended merely as a specific example. Many other configurations are possible, having more or fewer components than the system shown in the figure. For example, customized hardware may be used, and / or particular elements may be implemented in hardware, firmware, software (including applets), or a combination. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other manners and / or methods for implementing various embodiments.
[0166] The embodiments may be implemented by using a computer program product comprising a computer program / instructions which, when executed by a processor, cause the processor to perform any of the methods described in this disclosure.
[0167] While specific embodiments have been described, various modifications, variations, alternative constructions, and equivalents are encompassed within the scope of the present disclosure. The embodiments are not limited to operating in a particular data processing environment, but can freely operate in multiple data processing environments. Furthermore, while the embodiments have been described using a particular sequence of transactions and steps, it should be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described sequence of transactions and steps. Various features and aspects of the above-described embodiments may be used individually or jointly.
[0168] Furthermore, while embodiments have been described using particular combinations of hardware and software, it should be recognized that other combinations of hardware and software are within the scope of the present disclosure. Embodiments may be implemented exclusively in hardware, exclusively in software, or using a combination thereof. The various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, when a component or service is described as being configured to perform certain operations, such configuration may be achieved, for example, by designing electronic circuitry to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations, or any combination thereof. Processes may communicate using various techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.
[0169] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and alterations may be made without departing from the broader spirit and scope as set forth in the appended claims. Accordingly, while specific disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.
[0170] In the context of describing the disclosed embodiments (particularly in the context of the claims which follow), use of the terms "a," "an," and "the" and similar referents should be construed to encompass both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" should be construed as open-ended terms (i.e., meaning "including, but not limited to"), unless otherwise noted. The term "connected" should be construed as contained within, attached to, or joined to one another, either partially or as a whole, even if there is intervening material. The recitation of ranges of values herein, unless otherwise indicated herein, is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. Unless otherwise indicated herein or clearly contradicted by context, all methods described herein can be performed in any suitable order. Any and all examples provided herein, or the use of exemplary language (e.g., "etc.") are intended merely to clarify the embodiments and do not limit the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0171] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as generally used to indicate that an item, term, etc. may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is not intended to, and should not, generally imply that some embodiments require that at least one of X, at least one of Y, or at least one of Z, respectively, be present.
[0172] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Such variations will be readily apparent to those skilled in the art, and the present disclosure may be practiced in ways other than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, unless otherwise indicated herein, any combination of the above-described elements in all possible variations thereof is encompassed by the present disclosure.
[0173] Exemplary embodiments of the present disclosure may be described in light of the following clauses. Clause 1. A method is disclosed. The method may include a power management service identifying a plurality of components of an electric power system arranged according to a power distribution hierarchy, which may include a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, a node of the plurality of nodes of the power distribution hierarchy represents a corresponding component of the plurality of components. A subset of nodes in a first level of the plurality of levels may be derived from a particular node in a second level of the plurality of levels, the second level being higher than the first level. A set of lower-level components of the plurality of components may be represented by a subset of nodes in the first level that receive power distributed through a higher-level component of the plurality of components represented by a particular node in the second level. The method may include the power management service monitoring power consumption of a set of lower-level components represented by a subset of nodes in the first level. The method may include the power management service determining, based at least in part on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The method may further include, in response to determining that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, the power management service transmitting a power cap value to the lower-level component. In some embodiments, transmitting the power cap value enables the lower-level component to store the power cap value in memory while allowing the power consumption of each of the lower-level components to exceed the power cap value until a time period corresponding to the timing value expires.
[0174] Clause 2. The method of clause 1, wherein the power cap value is stored in memory while allowing the power consumption of each of the lower-level components to exceed the power cap value, thereby allowing the set of lower-level components to continue to breach the budget threshold associated with the higher-level component for at least a portion of a period corresponding to the timing value.
[0175] Clause 3. The method of clause 1 or clause 2, further including the power management service transmitting a timing value to an intermediate component represented by a particular node at an intermediate level between the first level and the second level of the power distribution hierarchy. In some embodiments, a set of lower-level components is configured to receive power distributed from higher-level components through the intermediate component. In some embodiments, transmitting the timing value causes the intermediate component to generate a timer associated with a period corresponding to the timing value. In some embodiments, the intermediate component is configured to instruct the lower-level component, based at least in part on expiration of the timer, to trigger an action to throttle its operation based at least in part on the power cap value.
[0176] Clause 4. In any of the methods of clauses 1 to 3, determining that the power consumption of the set of lower-level components has breached a budget threshold associated with a higher-level component may further include determining that the power consumption of the set of lower-level components is likely to breach a power budget allocated to the higher-level component.
[0177] Clause 5. The method of clause 4, further including obtaining, based at least in part on historical power consumption data associated with the set of lower-level components that receive power distributed by the higher-level component, a confidence value indicating a high likelihood that the power consumption value corresponding to the set of lower-level components will breach the budget amount of power allocated to the higher-level component. The method may include determining that the confidence value exceeds a threshold.
[0178] Clause 6. Some embodiments include a system. The system may include a memory configured to store instructions and one or more processors configured to execute the instructions of the method. The method may include identifying, by a power management service, a plurality of components of an electric power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, nodes of the plurality of nodes of the power distribution hierarchy represent corresponding components of the plurality of components. In some embodiments, a subset of nodes of a first level of the plurality of levels is derived from a particular node of a second level of the plurality of levels that is higher than the first level. In some embodiments, a set of lower-level components of the plurality of components represented by a subset of nodes of the first level receive power distributed through a higher-level component of the plurality of components represented by a particular node of the second level. The method may include monitoring, by the power management service, power consumption of the set of lower-level components represented by the subset of nodes of the first level. The method may include determining, by the power management service, based at least in part on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The method may include, in response to determining that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, transmitting, by a power management service, a power cap value to the lower-level components and causing the lower-level components to store the power cap value in memory while allowing the power consumption of each of the lower-level components to exceed the power cap value until a time period corresponding to the timing value expires.
[0179] Clause 7. The system of clause 6, wherein expiration of a timer corresponding to a time period triggers lower level components to limit power consumption within a range based, at least in part, on at least one of dynamic frequency scaling or dynamic voltage scaling.
[0180] Clause 8. The system of clause 6 or clause 7, wherein the second level of the power distribution hierarchy includes two or more of the plurality of nodes representing at least a higher level component corresponding to the first higher level component and a second higher level component different from the first higher level component, and by executing the instructions, the system further performs at least: 1) determining, with respect to the second higher level component, by the power management service, that there is an unused portion of the power provisioned to the second higher level component; and 2) determining, with the power management service, a first portion of the power provisioned to the second higher level component that is receiving power through the second higher level component. determining a specific power cap value for a second lower-level component of the set of two; 3) transmitting, by a power management service, the respective power cap values for the second lower-level component, where transmitting the respective power cap values triggers the second lower-level component to throttle operation at the second lower-level component as part of enforcing the respective power cap value; and 4) transmitting, by the power management service, a cancellation signal that causes a timer corresponding to the period to be cancelled, whereby cancelling the timer causes the lower-level component to refrain from enforcing the power cap value after the period has expired.
[0181] Clause 9. The system of any of clauses 6 through 8, wherein by executing the instructions, the system further, in response to determining that the power consumption of at least the set of lower-level components has breached a budget threshold associated with the higher-level components, determines, with the power management service, a power cap value based, at least in part, on at least one of a difference between the power consumptions of the set of lower-level components, a priority value associated with a workload running on the set of lower-level components, or an individual consumption rate of the set of lower-level components.
[0182] Clause 10. The system of any of clauses 6 to 9, wherein the power cap value is determined based on monitored data values obtained at or after determining that the power consumption of a set of lower-level components has breached a budget threshold associated with a higher-level component.
[0183] Clause 11. Some embodiments include a method including a power management service identifying a plurality of components of an electric power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to respective levels of a plurality of levels. In some embodiments, nodes of the plurality of nodes of the power distribution hierarchy represent corresponding components of the plurality of components. In some embodiments, a subset of nodes in a first level of the plurality of levels is derived from a particular node in a second level of the plurality of levels that is higher than the first level. In some embodiments, a set of lower-level components of the plurality of components represented by the subset of nodes in the first level receive power distributed through a higher-level component of the plurality of components represented by a particular node in the second level. The method may include the power management service monitoring power consumption of the set of lower-level components represented by the subset of nodes in the first level. The method may include the power management service determining, based at least in part on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The method may include, in response to determining that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, starting a timer corresponding to a timing value. In some embodiments, expiration of the timer indicates expiration of a time period corresponding to the timing value, the lower-level component of the set of lower-level components stores a power cap value in memory for the time period, and enforcement of the power cap value at the lower-level component is postponed until expiration of the timer.
[0184] Clause 12. The method of clause 11 may further include, in response to determining that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, the power management service determining a timing value based, at least in part, on at least one of a rate of change of the power consumption of the set of lower-level components, a direction of change corresponding to the rate of change of the power consumption of the set of lower-level components, a tolerance for power circuitry associated with the higher-level component, or an expected time required to initiate power capping at each component in the set of lower-level components.
[0185] Clause 13. The method of clause 11 or clause 12, wherein starting a timer corresponding to the timing value further includes the power management service transmitting the timing value to an intermediate component from which the lower-level component receives power. In some embodiments, transmitting the timing value causes the intermediate component to generate a timer associated with a period corresponding to the timing value.
[0186] Clause 14. The method of clause 13, wherein the intermediate component triggers enforcement of the power cap in the lower level component when the timer expires.
[0187] Clause 15. The method of any of Clauses 11 to 14, further including: 1) obtaining, based at least in part on historical power consumption data associated with a second set of lower-level components that receive power allocated by the higher-level component, a likelihood value indicating that an aggregated power consumption corresponding to the second set of lower-level components that receive power allocated by the higher-level component is likely to exceed a budget threshold associated with the higher-level component; and 2) in response to determining that the likelihood value exceeds the likelihood threshold, the power management service transmits an additional power cap value to a second lower-level component of the second set of lower-level components, wherein by transmitting the additional power cap value, the additional power cap value is used to constrain a corresponding power consumption at the second lower-level component of the second set of lower-level components.
[0188] Clause 16. In some embodiments, a system is disclosed. The system may include a memory configured to store instructions and one or more processors configured to execute the instructions to implement the method. The method includes identifying, by a power management service, a plurality of components of an electric power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to respective levels of a plurality of levels, wherein a node of the plurality of nodes of the power distribution hierarchy represents a corresponding component of the plurality of components. In some embodiments, a subset of nodes in a first level of the plurality of levels is derived from a particular node in a second level of the plurality of levels that is higher than the first level. In some embodiments, a set of lower-level components of the plurality of components is represented by a subset of nodes in the first level that receive power distributed through a higher-level component of the plurality of components represented by a particular node in the second level. The method may include monitoring, by the power management service, power consumption of the set of lower-level components represented by the subset of nodes in the first level. The method may further include determining, by the power management service, based at least in part on the monitoring, that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component. The method may include, in response to determining that the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, starting a timer corresponding to a timing value, in some embodiments, expiration of the timer indicates expiration of a time period corresponding to the timing value, the lower-level component of the set of lower-level components stores a power cap value in memory for the time period, and enforcement of the power cap value at the lower-level component is postponed until expiration of the timer.
[0189] Clause 17. The system of clause 16, wherein by executing the instructions, the system further, in response to determining that at least the power consumption of the set of lower-level components has breached a budget threshold associated with the higher-level component, determines, by the power management service, a timing value based, at least in part, on at least one of a rate of change of the power consumption of the set of lower-level components, a direction of change corresponding to the rate of change of the power consumption of the set of lower-level components, a tolerance for power circuitry associated with the higher-level component, or an expected time required to initiate power capping at each component in the set of lower-level components.
[0190] Clause 18. The system of clause 16 or clause 17, wherein the set of lower-level components is deemed to breach the budget threshold based, at least in part, by the power management service and based, at least in part, on historical consumption data associated with the set of lower-level components, on determining that the electrical consumption of the set of lower-level components is likely to breach the budget threshold.
[0191] Clause 19. The system of any of Clauses 16 to 18, wherein the second level of the power distribution hierarchy includes two or more of the plurality of nodes representing at least a higher-level component corresponding to the first higher-level component and a second higher-level component different from the first higher-level component, and by executing the instructions, the system further performs at least: 1) determining, with respect to the second higher-level component, that there is an unused portion of the power provisioned for the second higher-level component; 2) determining, with the power management service, a specific power cap value for a second lower-level component of the second set of lower-level components receiving power through the second higher-level component; 3) transmitting, with the power management service, the specific power cap value for the second lower-level component, wherein transmitting the specific power cap value for the second lower-level component triggers execution of power consumption limits at the second lower-level component as part of enforcing the specific power cap value; and 4) transmitting, with the power management service, a cancellation signal that causes the timer to be canceled.
[0192] Clause 20. The system of any of clauses 16 to 19, wherein the power cap value is determined based on monitored data values obtained at or after determining that the power consumption of a set of lower-level components has breached a budget threshold associated with a higher-level component.
[0193] All references cited herein, including publications, patent applications, and patents, are incorporated by reference to the same extent as if each individual reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0194] While the foregoing specification has described aspects of the disclosure with reference to specific embodiments thereof, those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure can be used individually or jointly. Moreover, the embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. A method performed by a computer, The computer identifies multiple components of a power system arranged according to a power distribution hierarchy, which comprises multiple nodes organized according to each of the multiple levels. A node among the plurality of nodes in the power distribution hierarchy represents a corresponding component among the plurality of components, a subset of nodes at a first level among the plurality of levels derives from a specific node at a second level higher than the first level among the plurality of levels, and a set of lower-level components among the plurality of components represented by the subset of nodes at the first level receives power distributed through the higher-level components among the plurality of components represented by the specific node at the second level, and the method further, The computer monitors the power consumption of the set of lower-level components represented by the subset of the first-level nodes, The computer determines, at least in part, based on the monitoring, that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component, A method performed by a computer, comprising: determining that the power consumption of the set of lower-level components exceeds the budget threshold associated with the higher-level components; transmitting a power limit to the lower-level components; and causing the lower-level components to store the power limit in memory, while allowing the power consumption of each of the lower-level components to exceed the power limit until the period corresponding to the timing value expires.
2. A method performed by a computer according to claim 1, wherein the set of lower-level components can continue to exceed the budget threshold associated with the higher-level component for at least a portion of the period corresponding to the timing value, by storing the power limit in memory while allowing the power consumption of each of the lower-level components to exceed the power limit.
3. The computer further includes transmitting the timing values to an intermediate component represented by a specific node at an intermediate level between the first and second levels of the power distribution hierarchy, The set of lower-level components is configured to receive power distributed from the higher-level components through the intermediate components, By transmitting the timing value, the intermediate component generates a timer associated with the period corresponding to the timing value. The method performed by the computer according to claim 1, wherein the intermediate component is configured to instruct the lower-level component to trigger an action to throttle its operation based on the expiration of the timer, at least in part, based on the power limit.
4. A method performed by a computer according to claim 1, wherein determining that the power consumption of the set of lower-level components exceeds the budget threshold associated with the higher-level component further includes determining that the power consumption of the set of lower-level components is likely to exceed the budget amount of power allocated to the higher-level component.
5. At least in part, this involves obtaining a confidence value, based on historical power consumption data associated with the set of lower-level components that receive power distributed by the higher-level components, that indicates a high probability that the power consumption corresponding to the set of lower-level components will exceed the budget amount of power allocated to the higher-level components, The method performed by a computer according to claim 4, further comprising determining that the confidence value exceeds a threshold.
6. Memory configured to store instructions, The system comprises one or more processors, and the one or more processors execute the instructions, at least, The one or more processors identify multiple components of a power system arranged according to a power distribution hierarchy comprising multiple nodes organized according to each of the multiple levels, A node among the plurality of nodes in the power distribution hierarchy represents a corresponding component among the plurality of components, a subset of nodes at a first level among the plurality of levels derives from a specific node at a second level higher than the first level among the plurality of levels, and a set of lower-level components among the plurality of components represented by the subset of nodes at the first level receives power distributed through the higher-level components among the plurality of components represented by the specific node at the second level. The one or more processors monitor the power consumption of the set of lower-level components represented by the subset of the first-level nodes, The one or more processors, at least in part, determine, based on the monitoring, that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component. A system configured to: determine that the power consumption of the set of lower-level components exceeds the budget threshold associated with the higher-level component, transmit a power limit to the lower-level component, and store the power limit in memory, allowing the power consumption of each of the lower-level components to exceed the power limit until the period corresponding to the timing value expires.
7. The system according to claim 6, wherein the expiration of the timer corresponding to the aforementioned period triggers the lower-level components to limit power consumption within a range, at least in part, based on at least one of dynamic frequency scaling or dynamic voltage scaling.
8. The second level of the power distribution hierarchy includes at least two or more of the plurality of nodes, each representing at least one upper-level component corresponding to a first upper-level component and a second upper-level component different from the first upper-level component, and by executing the instruction, the system further includes at least, The one or more processors determine with respect to the second higher-level component that there is an unused portion of the power provisioned to the second higher-level component, The one or more processors determine a specific power limit for a second lower-level component of a second set of lower-level components that receive power through the second higher-level component, The one or more processors transmit the power limit value to the second lower-level component, Transmitting the power limit triggers the second lower-level component to throttle its operation as part of implementing the power limit. The one or more processors transmit a cancellation signal so that the timer corresponding to the period is canceled. The system according to claim 6 or 7, wherein by canceling the timer, the lower-level component is prevented from implementing the power limit after the period has expired.
9. The system according to claim 6 or 7, wherein, by executing the instruction, the system further determines, in response to determining that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level components, the one or more processors determine the power limit based, at least in part, on at least one of the following: the difference in power consumption of the set of lower-level components, a priority value associated with the workload running on the set of lower-level components, or the individual consumption rates of the set of lower-level components.
10. The system according to claim 6 or 7, wherein the power limit is determined based on monitoring data values obtained at or after it is determined that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component.
11. A method performed by a computer, The computer identifies a plurality of components of a power system arranged according to a power distribution hierarchy comprising a plurality of nodes organized according to each of a plurality of levels, wherein a node among the plurality of nodes of the power distribution hierarchy represents a corresponding component among the plurality of components, a subset of nodes of a first level among the plurality of levels derives from a specific node of a second level higher than the first level among the plurality of levels, and the set of lower-level components among the plurality of components represented by the subset of nodes of the first level receive power distributed through the higher-level components among the plurality of components represented by the specific node of the second level, and the method further, The computer monitors the power consumption of the set of lower-level components represented by the subset of the first-level nodes, The computer determines, at least in part, based on the monitoring, that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component, The process includes: determining that the power consumption of the set of lower-level components exceeds the budget threshold associated with the higher-level component, and starting a timer corresponding to a timing value; The expiration of the timer indicates the expiration of the period corresponding to the timing value. The lower-level component of the set of lower-level components stores the power limit in memory during the period. A method performed by a computer in which the enforcement of the power limit in the lower-level component is postponed until the timer expires.
12. A method performed by a computer according to claim 11, further comprising the computer determining the timing value based, at least in part, on at least one of the following: the rate of change of the power consumption of the set of lower-level components; the direction of change corresponding to the rate of change of the power consumption of the set of lower-level components; an acceptable range for the power supply circuit associated with the higher-level component; or the expected time required to initiate power limit setting in each component within the set of lower-level components.
13. Starting the timer corresponding to the aforementioned timing value means that The method performed by the computer according to claim 11, further comprising the computer transmitting the timing value to an intermediate component from which the lower-level component receives power, and the intermediate component generating the timer associated with the period corresponding to the timing value by transmitting the timing value.
14. The method performed by the computer according to claim 13, wherein when the timer expires, the intermediate component triggers the implementation of the power limit in the lower-level component.
15. Obtaining a likelihood value that, at least in part, based on historical power consumption data associated with a second set of lower-level components that receive power distributed by the higher-level components, the total power consumption corresponding to the second set of lower-level components that receive power distributed by the higher-level components is likely to exceed the budget threshold associated with the higher-level component, The computer further includes determining that the likelihood value exceeds the likelihood threshold, and transmitting an additional power limit to a second lower-level component of the second set of lower-level components, A method performed by a computer according to claim 11, wherein the additional power limit is used to restrict the corresponding power consumption in the second lower-level component of the second set of lower-level components by transmitting the additional power limit.
16. Memory configured to store instructions, The system comprises one or more processors, and the one or more processors execute the instructions, at least, The one or more processors identify multiple components of a power system arranged according to a power distribution hierarchy comprising multiple nodes organized according to each of the multiple levels, A node among the plurality of nodes in the power distribution hierarchy represents a corresponding component among the plurality of components, a subset of nodes at a first level among the plurality of levels derives from a specific node at a second level higher than the first level among the plurality of levels, and a set of lower-level components among the plurality of components represented by the subset of nodes at the first level receives power distributed through the higher-level components among the plurality of components represented by the specific node at the second level. The one or more processors monitor the power consumption of the set of lower-level components represented by the subset of the first-level nodes, The one or more processors, at least in part, determine, based on the monitoring, that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component. The system is configured to start a timer corresponding to a timing value when it is determined that the power consumption of the set of lower-level components exceeds the budget threshold associated with the higher-level component, The expiration of the timer indicates the expiration of the period corresponding to the timing value. The lower-level component of the set of lower-level components stores the power limit in memory during the period. The implementation of the power limit in the lower-level component is postponed until the timer expires, in a system.
17. The system according to claim 16, wherein, by executing the instruction, the system further determines, in response to determining that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component, the one or more processors determine the timing value based, at least in part, on at least one of the following: the rate of change of the power consumption of the set of lower-level components, the direction of change corresponding to the rate of change of the power consumption of the set of lower-level components, the tolerance for the power supply circuit associated with the higher-level component, or the expected time required to initiate power limit setting in each component within the set of lower-level components.
18. The system according to claim 16 or 17, wherein the set of lower-level components is deemed to exceed the budget threshold based on the determination, at least in part, by one or more processors and at least in part, based on historical consumption data associated with the set of lower-level components, that the power consumption of the set of lower-level components is likely to exceed the budget threshold.
19. The second level of the power distribution hierarchy includes at least two or more of the plurality of nodes, each representing at least one upper-level component corresponding to a first upper-level component and a second upper-level component different from the first upper-level component, and by executing the instruction, the system further includes at least, The one or more processors determine with respect to the second higher-level component that there is an unused portion of the power provisioned to the second higher-level component, The one or more processors determine a specific power limit for a second lower-level component of a second set of lower-level components that receive power through the second higher-level component, The one or more processors transmit the specific power limit of the second lower-level component, Transmitting the specific power limit to the second lower-level component triggers the implementation of power consumption limits in the second lower-level component as part of enforcing the specific power limit. The system according to claim 16 or 17, wherein one or more processors transmit a cancellation signal that causes the timer to be canceled.
20. The system according to claim 16 or 17, wherein the power limit is determined based on monitoring data values obtained at or after it is determined that the power consumption of the set of lower-level components has exceeded the budget threshold associated with the higher-level component.
21. A program that causes a computer to execute the method according to any one of Claims 1 to 5 and Claims 11 to 15.