Techniques for Orchestrated Load Shedding
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2023-08-11
- Publication Date
- 2026-08-05
AI Technical Summary
Traditional methods of load shedding in data centers involve manual shutdowns of hosts or devices, lacking understanding of workloads and customer impacts, leading to suboptimal power reduction strategies that can cause disruptions and inefficiencies.
A dynamically orchestrated approach using a computer system to apply curtailment actions based on response levels, considering host attributes and power consumption thresholds, with automated power capping, workload migration, and host suspension or shutdown to achieve targeted power reductions.
This approach minimizes the impact on customers and workloads while effectively reducing power consumption, avoiding power failures and optimizing resource utilization in data centers.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 423,762, entitled "Orchestrated DC-Scale Load Shedding," filed November 8, 2022; U.S. Provisional Patent Application No. 63 / 439,576, entitled "Orchestrated DC-Scale Load Shedding," filed January 18, 2023; and U.S. Patent Application No. 18 / 338,695, entitled "Techniques for Orchestrated Load Shedding," filed June 21, 2023, the contents of which are incorporated herein by reference in their entireties for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure generally relates to techniques for orchestrating the reduction of power consumption in a data center. The disclosed systems, methods, devices, and services enable a dynamically orchestrated approach to constraining power consumption in a data center, including, but not limited to, utilizing power caps, suspending hosts, migrating instances and / or hosts, and / or shutting down hosts to achieve desired power reductions. [Background technology]
[0003] background Data centers consist of a power infrastructure that provides numerous safety features according to a hierarchical power distribution. Power is supplied by a local utility and allocated according to this power distribution hierarchy to various components of the data center, including power distribution units (PDUs) (e.g., transformers, distribution panels, busways, rack PDUs, etc.) and power-consuming devices (e.g., servers, network devices, etc.). This ensures that the power consumed by all downstream devices adheres to the power limits of each upstream device. During peak demand and / or component failures, the data center may be unable to handle the demand effectively, potentially resulting in widespread power failures and, at a minimum, significant disruptions to downstream devices, resulting in reduced processing capacity within the data center and / or power failures at the data center, thereby degrading the user experience. In a worst-case scenario, a power failure in one data center can trigger cascading power failures of other devices within the same or other data centers as workloads are redistributed in an attempt to recover from the initial outage. Additionally, in some embodiments, external factors (eg, mandates by government regulations) may require that power consumption at a particular data center be reduced for a particular period of time. Summary of the Invention [Problem to be solved by the invention]
[0004] Traditional methods of managing load shedding have involved operators manually shutting down host racks or individual devices one by one to reduce power consumption. Alternatively, traditional techniques include pressing emergency shutoff switches or buttons configured to turn off power to all or most of the data center. Operators determining which components to power down typically lack understanding of the workloads or customers affected by their actions, or the extent of the impact of taking those actions. Traditional approaches such as those described herein result in suboptimal approaches with respect to the affected customers and workloads. More advanced approaches may have less impact on customers and / or devices / workloads while still allowing for the power reduction desired in a given situation. Therefore, it is desirable to improve power management techniques, particularly with respect to load shedding, to reduce power consumption to an impact and amount that is sufficient and avoids the potential hazards of traditional approaches.
[0005] Quick Overview In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of some embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The illustrations and descriptions are not intended to be limiting. [Means for solving the problem]
[0006] Some embodiments may include a method. The method may include a computer system acquiring configuration data including a plurality of response levels that specify applicability of respective sets of curtailment actions to a plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of the first set of curtailment actions to the plurality of hosts. The method may include the computer system acquiring current values of aggregate power consumptions of the plurality of hosts. The method may further include the computer system acquiring current values of aggregate power thresholds of the plurality of hosts. The method may further include selecting a first response level from the plurality of response levels based at least on a difference between the current values of the aggregate power consumptions and the current values of the aggregate power thresholds. The method may include causing application of the first set of curtailment actions to at least one host of the plurality of hosts in accordance with the selected first response level. In some embodiments, causing application of the first set of curtailment actions through performing operations to apply the first set of curtailment actions to the plurality of hosts or instructing a component of the computer system to perform such operations.
[0007] In some embodiments, the method includes the computer system executing a first set of reduction actions on the plurality of hosts according to the selected first response level.
[0008] In some embodiments, the method further includes determining whether the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold, and selecting from a plurality of response levels in response to determining that the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold. In some embodiments, the selected first response level is applicable to a subset of the plurality of hosts.
[0009] In some embodiments, the first response level specifies the applicability of a first set of abatement actions to a plurality of hosts based at least on host attributes, which may include any suitable combination of priority level, type, or category.
[0010] In some embodiments, each successive response level in the plurality of response levels specifies an increasingly severe set of mitigation actions to be applied to the plurality of hosts.
[0011] In some embodiments, a first power cap value applied to the first host by a first curtailment action according to a first response level is lower than a second power cap value applied to the first host by a different curtailment action according to a second response level.
[0012] In some embodiments, the first response level specifies application of a first curtailment action to a subset of the plurality of hosts associated with a priority level lower than the first level, and the second response level specifies application of a second curtailment action to a second subset of the plurality of hosts associated with a priority level lower than the second level. In some embodiments, the second level is higher than the first level. In some embodiments, the second subset of the plurality of hosts does not include the first subset of the plurality of hosts.
[0013] In some embodiments, the selected first response level is the response level of the plurality of response levels associated with the least severe set of reduction actions that achieves a new value of the aggregate power consumption that is less than or equal to the current value of the aggregate power threshold.
[0014] In some embodiments, selecting a first response level from a plurality of response levels based at least on a difference between the current value of the aggregate power consumption and the current value of the aggregate power threshold may include at least one of: 1) determining that a first reduction in aggregate power consumption resulting from the first response level is greater than the difference; 2) determining that a second reduction in aggregate power consumption resulting from the second response level is less than the difference; and / or 3) selecting the first response level.
[0015] In some embodiments, the current value of the aggregate power consumption breaches the current value of the aggregate power threshold due to a decrease in the aggregate power threshold. In some embodiments, the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold for a period of time.
[0016] In some embodiments, the decrease in aggregate power threshold was caused by at least one of: 1) a degradation or failure of a temperature control system for a physical environment containing multiple hosts, 2) a governmental requirement to reduce power consumption, or 3) an increase in external temperature.
[0017] In some embodiments, the method may include determining that during a first time period, a first value of the aggregate power threshold exceeds a current value of the aggregate power consumption. The method may include determining that during a current time period following the first time period, the current value of the aggregate power consumption exceeds a current value of the aggregate power threshold. In some embodiments, the current value of the aggregate power threshold during the current time period may be less than the first value of the aggregate power threshold during the first time period.
[0018] In some embodiments, the first curtailment action includes at least one of 1) enforcing a power cap on the host, 2) migrating workloads from the host, or 3) shutting down or suspending the host.
[0019] In some embodiments, the method may include, by the computer system, obtaining a new value for the aggregate power threshold. The method may include performing recovery from the first response level based on at least the new value of the aggregate power threshold.
[0020] In some embodiments, performing recovery from the second response level includes at least one of 1) quiescing power capping on the host, 2) migrating workloads back to the host, or 3) restarting a previously shut down or paused host.
[0021] In some embodiments, a first response level from the plurality of response levels is implemented in response to receiving user input at a user interface, hi some embodiments, the selected first response level is associated with a reduction that exceeds a difference between a current value of the aggregate power consumption and a current value of the aggregate power threshold.
[0022] In some embodiments, another method is disclosed. The method may include a computer system acquiring configuration data including multiple response levels that specify applicability of respective sets of curtailment actions to a plurality of hosts. In some embodiments, a first response level of the multiple response levels specifies applicability of the first set of curtailment actions to the plurality of hosts. The method may further include the computer system acquiring predicted values of aggregate power consumption of the plurality of hosts for a future time period. The method may further include the computer system acquiring predicted values of aggregate power thresholds of the plurality of hosts for a future time period. The method may further include selecting a first response level from the multiple response levels based on at least a predicted difference between the predicted values of aggregate power consumption and the predicted values of aggregate power thresholds. The method may further include identifying one or more workloads that (a) are currently executing on the plurality of hosts and (b) would be affected by applying the first set of curtailment actions to the plurality of hosts in accordance with the selected first response level. The method may further include preemptively migrating affected workloads from the plurality of hosts to one or more other hosts before the future time period.
[0023] In some embodiments, the method may further include any suitable combination of: 1) determining a predicted aggregate power threshold value based on at least the health of a temperature control system for a physical environment including a plurality of hosts, or 2) announced or anticipated governmental power supply cuts, or 3) anticipated increases in external temperature.
[0024] In some embodiments, the method may further include determining a predicted value of the aggregate power consumption based on a historical pattern of the aggregate power consumption.
[0025] In some embodiments, identifying one or more workloads that a) are currently executing on a plurality of hosts and b) would be affected by applying a first set of reduction actions to the plurality of hosts according to a selected first response level includes 1) determining that particular actions will be applied to the first host according to a selected second response level, 2) determining that the first workload is currently executing on the first host, and / or 3) determining that the first workload will be affected by applying the first set of reduction actions to the plurality of hosts according to the selected first response level.
[0026] Systems, devices, and computer media are disclosed, each of which may include one or more memories capable of storing instructions corresponding to the methods disclosed herein. The instructions may be executed by one or more processors of the disclosed systems and devices to perform the methods disclosed herein. One or more computer programs may be configured to perform specific operations or acts corresponding to the described methods by including instructions that, when executed by a data processing device, cause the device to perform those acts. [Brief explanation of the drawings]
[0027] [Figure 1] FIG. 1 illustrates an example physical environment (e.g., a data center or portion thereof) including various components, according to at least one embodiment. [Figure 2] FIG. 1 is a simplified diagram of an exemplary power distribution infrastructure including various components of a data center, according to at least one embodiment. [Figure 3] FIG. 1 illustrates an example architecture of an exemplary power orchestration system configured to orchestrate power consumption reductions applicable to various resources in a physical environment (e.g., a data center, a data center room, etc.), according to at least one embodiment. [Figure 4] FIG. 1 illustrates an example architecture of a power management service for detecting and constraining excessive power consumption, according to at least one embodiment. [Figure 5] 3 illustrates an example power distribution hierarchy corresponding to the arrangement of components in FIG. 2, according to at least one embodiment. [Figure 6] FIG. 1 is a flow diagram illustrating an example method for managing excessive power consumption, according to at least one embodiment. [Figure 7] FIG. 1 illustrates an example architecture of a VEPO orchestration service for orchestrating power consumption constraints and / or power shutdown tasks, according to at least one embodiment. [Figure 8] 1 is a table illustrating an example set of response levels and a respective set of reduction actions for each response level, according to at least one embodiment. [Figure 9] FIG. 1 illustrates an example range of components affected by one or more response levels utilized by a VEPO orchestration service to orchestrate power consumption constraints and / or power shutdown tasks, according to at least one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating an example use case in which multiple response levels are applied based on the current state of a set of hosts, according to at least one embodiment. [Figure 11] 1 is a schematic diagram of an example user interface according to at least one embodiment. [Figure 12] FIG. 1 is a flow diagram illustrating an example method for training one or more machine learning models, according to at least one embodiment. [Figure 13] FIG. 1 is a block diagram illustrating an example method for triggering the application of a corresponding set of abatement actions associated with a response level, according to at least one embodiment. [Figure 14]FIG. 1 is a block diagram illustrating an example method for preemptively migrating workloads affected by a selected response level, according to at least one embodiment. [Figure 15] FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 16] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 17] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 18] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 19] FIG. 1 is a block diagram illustrating an example computer system according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0028] Detailed Description In the following description, various embodiments are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Additionally, well-known features may be omitted or simplified so as not to obscure the described embodiments.
[0029] This disclosure relates to managing power consumption and orchestrating power consumption reductions within an environment (e.g., in hosts or power distribution units of a data center, or portions thereof). More particularly, techniques are described for enabling orchestrated load shedding to implement power consumption reductions due to current conditions and according to aggregate power thresholds.
[0030] It is desirable to balance the supply and consumption of power in a data center. If the power consumed exceeds the available supply, the balance can be restored by increasing the power supply or slowing down the rate of power consumption. If this balance is not maintained and a component attempts to consume more than the available supply, a circuit breaker may trip and disconnect the component from the supply.
[0031] The maximum aggregate power threshold for a particular environment can depend on many factors, including, but not limited to, the current power consumption values of hosts in the environment, the operational status of associated components such as temperature control systems (e.g., HVAC, chillers, etc.), and environmental conditions (e.g., ambient temperature, external temperature outside the environment, etc.). Conventional systems are not configured to manage power consumption based on these factors. When power consumption peaks and a power failure is imminent, conventional techniques primarily utilize manual intervention to shut down or suspend hosts. This leads to suboptimal load shedding strategies because operators implementing load shedding typically lack understanding of the effects of their actions. For example, when selecting devices to shut down or suspend, operators have little, if any, knowledge of customers, workloads, instances, or hosts, or their respective priorities. A power failure can cause an interruption to operations performed by components in a data center. For example, a website hosted on a server in a data center crashes if a tripped circuit breaker disconnects the server from the data center's power source. A non-optimal load shedding strategy may be inefficient and may result in more than sufficient power reduction to reduce aggregate power consumption by a desired amount. Additionally or alternatively, a non-optimal load shedding strategy may result in a broader / larger impact on the underlying hosts, instances, workloads, and corresponding customers than may be desirable.
[0032] The balance between power supply and consumption in a data center can be managed by maintaining an equilibrium between available supply and demand and / or consumption. Because power is often statically allocated in long-term contracts with electric utilities, increasing or decreasing the power supply in a data center may sometimes be infeasible. While supply may be statically allocated, power consumption in a data center can change, sometimes drastically. As an overly simplistic example, as the number of threads executed by a server's processor increases, the server's power consumption may increase. The ambient temperature within a data center can affect the power consumed by the data center's cooling system. The higher the load a cooling system operates, the more power it consumes as it operates to reduce the ambient temperature experienced within the data center. If the cooling system fails, the data center may be unable to withstand the heat generated at the current power consumption level. Demand caused by some components in a data center, such as uninterruptible power supplies, power distribution units, cooling systems, or busways, can be difficult to regulate. However, some components, such as servers, virtual machines, and / or bare metal instances, are more easily constrained. Additionally, workloads and / or instances can be migrated to other hosts and / or instances to concentrate resources on a smaller subset of hosts, thereby resulting in more idle and / or free hosts. Power reductions can then be applied to the idle / free hosts to ensure minimal impact (e.g., number of hosts, customers, instances, workloads).
[0033] Many conventional power management techniques utilize power capping to constrain operation in power-consuming devices (e.g., servers, network devices, etc.) in a data center. When power capping is utilized, a power capping limit can be used to constrain the power consumed by the device. Power capping is used to constrain (e.g., throttle) operation in a server to ensure that the server's power consumption does not exceed the power cap limit. Using power capping ensures that the allocated power limits of upstream devices are not breached, and each upstream device is provisioned according to a worst-case scenario in which each downstream device is estimated to consume its respective allocated maximum power. However, these downstream devices may often consume less power than their allocated maximum power, leaving at least a portion of the power allocated to the upstream device unused. These techniques waste valuable power and limit the density of power-consuming devices that can be utilized in a data center.
[0034] An efficient power infrastructure within a data center is necessary to increase provider profit margins, manage scarce power resources, and make the services offered by the data center more environmentally friendly. Data centers that include components hosting multi-tenant environments (e.g., public clouds) can experience higher-than-average consumption because not all tenancies of the cloud are in use simultaneously. To improve data center efficiency and resource utilization, data center providers may increase servers and / or tenancies so that power consumption across all power-consuming devices approaches the data center's allocated power capacity. However, in some cases, reducing the gap between allocated power capacity and power consumption increases the risk of tripping circuit breakers and losing the ability to utilize computing resources. The techniques described herein minimize the frequency with which downstream devices are constrained, allowing these devices to utilize previously unused power while maintaining a high degree of safety regarding avoiding power failures.
[0035] Technical effects The disclosed systems and methods provide an automated, dynamically orchestrated approach to constraining power consumption (e.g., via power capping, workload migration, etc.), suspending hosts, migrating instances and / or hosts, and / or shutting down hosts to achieve desired power reductions. In some embodiments, at least some of the functions performed toward these actions are user-selectable and / or based on user input. The described technology provides various response levels. Each response level can be associated with a set of reduction actions (referred to herein as “actions” for brevity). Each response level (referred to herein as “levels” for brevity) may provide a set of increasingly severe reduction actions to apply. If an aggregate power threshold is breached, a response level may be selected (e.g., by the system, based on user input, etc.) based on an estimated power reduction corresponding to each of the response levels. When implemented, the least severe and / or least impactful level sufficient to bring the aggregate power consumption of the data center below the aggregate power threshold may be selected (e.g., by the system, based on user input, etc.).
[0036] The estimated impact of applying the reduction action associated with the level can be determined at runtime based on the current attributes of the workload implemented by the virtual machine and / or bare metal instance, the customer with which the workload and / or affected host is associated, and the priorities corresponding to those workloads, hosts, and / or customers. Demand (e.g., corresponding to an aggregate power threshold) can change dynamically based on operational status data and environmental data corresponding to various components of the datacenter. These factors can be utilized to identify changes in aggregate power thresholds (e.g., aggregate power consumption / heat / demand thresholds that the components of the datacenter can collectively manage) as they occur in power management capabilities within the datacenter. The selection of a response level may be triggered by real-time changes in demand, via request (e.g., by user request, by a request submitted by a government agency mandating a reduction in power consumption, etc.), or by any suitable trigger, and may be based on the impact of implementing the corresponding reduction action.
[0037] By utilizing these techniques described herein, power constraints, migration tasks, and / or suspend or shutdown tasks employed can be tailored to the current state and resources (e.g., hosts, instances, workloads) within the datacenter, such that excessive curtailment is mitigated or avoided entirely. These techniques provide real-time capabilities to reduce the impact of power management response actions on customers, hosts, instances, and / or workloads, while ensuring that the risk of power failure is effectively avoided. The systems and methods described herein provide a more efficient and effective power orchestration approach than conventional systems, which can result in a more satisfying user experience. The particular action to implement may be selected (e.g., automatically by the system, through user selection, etc.) based, at least in part, on any suitable combination of: 1) maximizing overall power consumption reduction, 2) identifying / implementing the least severe response level, 3) identifying / implementing the least impactful response level, and / or 4) identifying a response level whose estimated power reduction is sufficient to reduce the current aggregate power consumption value below the current or projected aggregate power threshold while providing the least amount of excessive power reduction (e.g., reduction beyond that required to lower the current aggregate power consumption below the current aggregate power threshold). In this manner, the present technology provides a more flexible and intelligent approach to shedding loads in various situations and / or due to various triggering events.
[0038] 1 illustrates an example environment (e.g., environment 100) including various components, according to at least one embodiment. Environment 100 may be a physical environment, such as a data center (e.g., data center 102) or a portion thereof (e.g., a room in a data center, such as room 110A). Environment 100 may include a dedicated space for hosting any suitable number of servers, such as servers 104A-104P (collectively referred to as "servers 104"), and infrastructure for hosting those servers, such as networking hardware, a cooling system (also referred to as a "temperature control system"), and storage devices. Servers 104A-104P may also be referred to as "hosts." Networking hardware (not shown) in data center 100 allows remote users to interact with the servers over a network (e.g., the Internet). Any suitable number (e.g., 10, 14, 21, 42, etc.) of servers 104 may be held in various racks, such as racks 106A-106H (collectively referred to as "racks 106"). Racks 106 may include a frame or enclosure in which a corresponding set of servers is positioned and / or mounted.
[0039] Various subsets of racks 106 can be organized into groups referred to as "rows" (e.g., rows 108A-108D, collectively referred to as "rows 108"). In some implementations, rows 108 can include any suitable number of racks (e.g., 5, 8, 10, up to 10, etc.) arranged (e.g., within a threshold distance of each other). In other implementations, rows can be organizational units, and racks comprising a given row can be located in different locations (not necessarily within a threshold distance of each other). As an example, rows 108 can be located in rooms (e.g., room 110A, room 110N, etc.). A room (e.g., room 110A) can be a section of a building or a physical enclosure or physical environment in which any suitable number of racks 106 are located. In other embodiments, a room can be an organizational unit, and rooms can be located in different physical locations, or multiple rooms can be located in a single section of a building.
[0040] Various temperature control systems (e.g., temperature control systems 112A-112N, temperature control systems 114A-114N, etc.) may be configured to manage the ambient temperature of a data center or portion thereof. As a non-limiting example, temperature control systems 112A-112N may be associated with room 110A, and temperature control systems 114A-114N may be associated with room 110N. Any suitable number of temperature control systems may be associated with a data center and / or portion thereof. In some embodiments, these temperature control systems may be any suitable heating, ventilation, and air conditioning (HVAC) devices (e.g., air conditioning units), chillers (e.g., water circulators that control temperature by circulating a liquid such as water), or the like. In some embodiments, each temperature control system may be associated with a corresponding amount of heat it is capable of and / or configured to manage (e.g., the amount of heat generated by a corresponding amount of power consumption of servers 104).
[0041] FIG. 2 illustrates a simplified diagram of an exemplary power distribution infrastructure 200 including various components (e.g., components of the data center 102 of FIG. 1 ) according to at least one embodiment. The power distribution infrastructure 200 can be connected to a utility power source (not shown), and power can be initially received by one or more uninterruptible power supplies (uninterruptible power supply (UPS) 202). In some embodiments, power may be received at the UPS from a utility company via an on-site power substation (not shown) configured to establish suitable voltage levels for distributing power throughout the data center. The UPSs 202 may each include dedicated batteries or generators to provide emergency power in the event of a failure of the input power source. The UPSs 202 can monitor the input power and provide backup power if a drop in input power is detected.
[0042] The power distribution infrastructure 200 may include any suitable number of intermediate power distribution units (PDUs) (e.g., intermediate PDUs 204) that connect to and receive power / electricity from the UPS 202. Any suitable number of intermediate PDUs 204 may be disposed between a UPS (of the UPS 202) and any suitable number of string PDUs (e.g., string PDUs 206). A power distribution unit (e.g., intermediate PDUs 204, string PDUs 206, rack PDUs 208, etc.) may be any suitable device configured to control and distribute power / electricity. Example power distribution units may include, but are not limited to, a main distribution panel, a distribution panel, a remote power panel, a bus bar, a power strip, a transformer, etc. Power may be provided from the UPS 202 to the intermediate PDUs 204. The intermediate PDUs 204 may distribute power to downstream components of the power distribution infrastructure 200 (e.g., string PDUs 206).
[0043] Power distribution infrastructure 200 may include any suitable number of string power distribution units (including string PDUs 206). A string PDU may include any suitable PDU (e.g., remote power panels, bus bars / ways, etc.) disposed between an intermediate PDU (e.g., a PDU among intermediate PDUs 204) and one or more rack PDUs (e.g., rack PDU 208A, rack PDU 208N, collectively referred to as “rack PDUs 208”). A “string PDU” refers to a PDU configured to distribute power to one or more strings of devices (e.g., string 210 including servers 212A-212D, collectively referred to as “servers 212”). As discussed above, a string (e.g., string 210) may include any suitable number of racks (e.g., racks 214A-214N, collectively referred to as “racks 214”) in which servers 212 are located.
[0044] Power distribution infrastructure 200 may include any suitable number of rack power distribution units (including rack PDUs 208). A rack PDU may include any suitable PDU positioned between a column PDU (e.g., column PDU 206) and one or more servers (e.g., servers 212A, 212B, etc.) corresponding to a rack (e.g., rack 214A, which is an example of rack 106 in FIG. 1). A "rack PDU" refers to any suitable PDU configured to distribute power to one or more servers in a rack. A rack (e.g., rack 214A) may include any suitable number of servers 212. In some embodiments, rack PDU 208 may include an intelligent PDU further configured to monitor, manage, and control consumption across multiple devices (e.g., rack PDU 208A, servers 212A, 212B, etc.).
[0045] The servers 212 (each an example of a server 104 in FIG. 1 ) may each include a power controller (power controllers 216A-216D, collectively referred to as “power controller 216”). A power controller refers to any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in that device. The power controllers 216 can individually monitor the power consumption of each server on which they operate. The power controllers 216 can each be configured to implement power capping to constrain the power consumption in their respective servers. Implementing power capping includes any suitable combination of monitoring the power consumption in the server, determining whether to constrain (e.g., constrain within a range, constrain, etc.) the power consumption in the server (e.g., based at least in part on a comparison of the server's current power consumption with a stored power capping limit), and constraining / constraining the power consumption in the server (e.g., using dynamic voltage and frequency scaling of processors and memory to throttle the server's power consumption). The implementation of power capping is sometimes referred to as "power capping."
[0046] The data center 102 of FIG. 1 may include various components shown in the power distribution infrastructure 200. By way of example, room 110A may include one or more busways (each an example of a row PDU 206). A busbar (also referred to as a "busway") refers to a duct of conductive material through which electrical power can be distributed (e.g., within room 110A). The busway can receive power from the power distribution units of the intermediate PDUs 204 and provide power to one or more racks (e.g., rack 106A, FIG. 1, rack 106B, FIG. 1, etc.) associated with a row (e.g., row 108A, FIG. 1). Each power infrastructure component that distributes / provides power to other components also consumes a portion of the power passing through it. This loss can be caused by heat loss due to the power flowing through that component or by power consumed directly by the component (e.g., power consumed by processors in a rack PDU).
[0047] FIG. 3 illustrates an example architecture of an exemplary power orchestration system 300 configured to orchestrate power consumption reductions applicable to various resources of a physical environment (e.g., a data center, a data center room, etc.) according to at least one embodiment. The term “resource” may be considered to include hosts, workloads (e.g., virtual machines and / or bare metal instances executing those workloads), and / or customers associated with those hosts and / or workloads / instances. Power orchestration system 300 may be configured to monitor power consumption levels (e.g., current individual and / or aggregate power consumption values corresponding to various hosts, such as host 324, each of which is an example of server 104 in FIG. 1). Based at least in part on that monitoring, power orchestration system 300 may manage power consumption within the physical environment such that an aggregate power threshold is implemented. As used herein, an aggregate power threshold represents the maximum amount of power consumption that can be managed by components of system 300 given current conditions. The aggregate power thresholds may be dynamically adjusted as conditions change (e.g., based on government agency mandates to reduce power consumption, based on actual or predicted environmental conditions, or based on actual or predicted current power consumption values, based on the operational status of various components such as temperature control systems in the physical environment, etc.). The specific impact (e.g., scope, applicability) of changes made to implement the current aggregate power thresholds may change based on real-time conditions.Some example actions that can be taken to enforce the current aggregate power threshold can be setting a power cap (e.g., enforcing a maximum cap on power consumption in a particular device, such as one or more of the hosts 324), or otherwise allocating a budgeted amount of power to any suitable device (e.g., the host 324, the PDUs 202, 204, 206, 208 of FIG. 2, etc.), pausing or shutting down the host and / or instance (e.g., a VM or BM) and / or workloads running on the instance, migrating a workload from one instance to another, or migrating a workload from one host to another.
[0048] The power orchestration system 300 may include various components, such as those shown in FIG. 3. For example, the power orchestration system 300 may include a VEPO orchestration service 302. The VEPO orchestration service 302 may be configured to obtain various data from which impact and mitigation actions can be determined. By way of example, the VEPO orchestration service 302 may be configured to obtain any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318.
[0049] The power data 310 may include any suitable budget / allocation values for the host 324 or PDUs, including the rack PDU 320 (an example of the rack PDU 208 in FIG. 2 ), power cap values, and / or current power consumption values indicating the amount of power currently being consumed by the device. The power data 310 may be stored in a location accessible to the VEPO orchestration service 302 and / or the power data 310 may be obtained from sources and / or services configured to manage and / or obtain such data. In some embodiments, the power management service 304 may be configured to obtain the power data 310 and store such data in a location from which the VEPO orchestration service 302 can retrieve the data. In some embodiments, the power management service 304, or another component configured to obtain such data, may provide the power data 310 directly to the VEPO orchestration service 302 according to a predefined schedule, frequency, or periodicity, or via request.
[0050] The account data 312 may include any suitable attributes of a customer or any suitable attributes of hosts and / or instances associated with the customer. Some example attributes may include a category associated with the customer's account (e.g., a free tier category indicating that the customer is using the service free of charge and may have limited features or resources; a free trial category indicating that the customer is using the service free of charge for a limited period of time). The category may correspond to a priority associated with the customer, or a priority may be assigned to the customer, host, or workload in other manners. The priority may indicate a relative degree of importance, such as low priority, medium priority, and high priority, although other priority schemes are contemplated. Another example attribute may include an identifier associated with the customer, host, instance, and / or workload. The account data 312 may be stored in a location accessible to the VEPO orchestration service 302 from which the account data 312 can be retrieved, or the account data 312 may be retrieved from a service and / or source configured to manage and / or retrieve such data (e.g., an account service in a cloud computing environment, not shown). In some embodiments, account data 310 to VEPO orchestration service 302 is directly from a source, or component configured to obtain such data, according to a predefined schedule, frequency, or periodicity, or via request.
[0051] Environmental data 314 may include any suitable data associated with the physical environment or environmental data indicative of external conditions. By way of example, environmental data 314 may include an ambient temperature reading for the physical environment, an external temperature outside the physical environment, a design point indicative of an external temperature the physical environment is designed to withstand, or the like. At least a portion of environmental data 314 may be collected from one or more sensors (not shown) configured to measure particular conditions, such as the ambient temperature within the physical environment or an external temperature occurring outside the physical environment. In some embodiments, at least a portion of environmental data 314 may be provided by a weather source, such as a weather forecast, stored within storage of power orchestration system 300, or may be obtainable from an external source, such as a weather forecast website. In some embodiments, environmental data 314 may include tables and / or protocols by which curtailed capacity / capacity can be determined / identified. By way of example, the environmental data 314 may include a table or mapping indicating that a particular difference between the ambient temperature and temperatures occurring outside the physical environment (referred to as "external" temperatures) is associated with a particular reduced capacity of power consumption that a component of the system, or individual components, can monitor. The environmental data 314 may be stored in a location accessible to the VEPO orchestration service 302 from which the environmental data 314 can be retrieved, or the environmental data 314 may be obtained from a service and / or source configured to manage and / or obtain such data (e.g., a metrics service of a cloud computing environment, a power management service 304, etc.). In some embodiments, the environmental data 314 to the VEPO orchestration service 302 is directly from a source or component configured to obtain such data according to a predefined schedule, frequency, or periodicity, or via request.
[0052] Host / instance data 316 may include any suitable data associated with a host and / or instance. Host / instance data 316 may include workload metadata identifying workloads executing on a host and / or through a particular instance. In some embodiments, host / instance data 316 may identify a corresponding customer or corresponding priority of a host and / or instance and / or workload, etc. At least a portion of host / instance data 316 may initially be obtained and / or maintained by a separate service (e.g., compute service 306, an example of a compute service control plane in a cloud computing environment). Host / instance data 316 may be stored in a location accessible to VEPO orchestration service 302 and from which host / instance 316 can retrieve it, or host / instance 316 may obtain it from a service and / or source configured to manage and / or obtain such data (e.g., compute service 306, etc.). In some embodiments, host / instance data 316 may be provided directly to VEPO orchestration service 302 from a source or component configured to obtain such data according to a predefined schedule, frequency, or periodicity, or via request.
[0053] The operational data 318 may include any suitable data associated with the operational status or state corresponding to one or more devices or components of the physical environment. As a non-limiting example, the operational data 318 may include the operational status of one or more temperature control systems (e.g., temperature control system 112 of FIG. 1 ). In some embodiments, the operational data 318 may indicate which temperature control systems are operational and / or the operational capabilities of those systems. In some embodiments, the environmental data 314 may include tables and / or protocols by which reduced capacity / capacity can be determined / identified. As an example, the operational data 318 may include a table or mapping indicating that a failure (e.g., complete or partial) of a particular component (e.g., a particular temperature control system) is associated with a particular reduced capacity of power consumption that the component of the system as a whole, or an individual component of the system, can manage. At least a portion of the operational data 318 may initially be obtained and / or maintained by a separate service (e.g., power management service 304, etc.). The operational data 318 may be stored in a location accessible to the VEPO orchestration service 302 from which the operational data 318 can be retrieved, or the operational data 318 may be obtained from a service and / or source configured to manage and / or obtain such data (e.g., compute service 306, etc.). In some embodiments, the operational data 318 may be provided directly to the VEPO orchestration service 302 from a source or component configured to obtain such data according to a predefined schedule, frequency, or periodicity, or via request.
[0054] It should be understood that the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 may include current data indicating current values and / or historical data indicating corresponding past values. In some embodiments, future attributes of the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 may be predicted based, at least in part, on these past values. In some embodiments, machine learning may be utilized to predict any suitable portion of these future attributes. Example methods for predicting future values of the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 are discussed in more detail with respect to FIG. 12 . In some embodiments, the VEPO orchestration service 302 can be configured to aggregate any suitable combination of power data 310, account data 312, environment data 314, host / instance data 316, or operational data 318 data into a table, mapping, or database from which the data can be filtered or sorted to identify a subset of resources (e.g., hosts, instances, workloads, customers) for a given set of constraints (e.g., low priority workloads, free tier customers, etc.).
[0055] The VEPO orchestration service 302 may be configured to identify and / or modify aggregate power thresholds associated with the physical environment based, at least in part, on any suitable combination of: current / aggregate power consumption values (actual or predicted) associated with any suitable combination of hosts 324; received requests related to governmental entities (e.g., local power authorities) mandating / requiring specific or overall power reductions, possibly over a specific period of time (e.g., the next 24 hours); current or predicted environmental conditions (e.g., current or predicted ambient temperature, current or future external temperature, etc.); or current or predicted operational status corresponding to a temperature control system (e.g., current or predicted total / partial failure);
[0056] The VEPO orchestration service 302, through associated functionality or functionality provided by other systems and / or services, can be configured to manage actual power consumption values corresponding to hosts / instances / workloads in a physical environment. Managing actual power consumption values can include setting power caps, pausing workloads, instances, hosts, shutting down workloads / instances / hosts, migrating instances from one host to another, or migrating workloads from one instance and / or host to another, etc.
[0057] The VEPO orchestration service 302 can obtain configuration data (e.g., mappings, protocol sets, rules, etc.) corresponding to multiple response levels. Each response level can be associated with a corresponding set of curtailment actions (e.g., power capping, migration, suspend, shutdown, etc.) that can be performed on components of the physical environment (e.g., servers 104, PDUs 202, 204, 206, 208, etc.). In some embodiments, each set of curtailment actions corresponding to a particular level is associated with implementing a different possible reduction in the aggregate power consumption of the data center. In some embodiments, the multiple response levels indicate increasing severity when curtailment actions are performed to reduce the aggregate power consumption of the data center. "Severity" is intended to refer to the relative degree to which an action is disruptive. As an example, an action to set a power cap on a server can be considered less severe than an action to completely shut down the server because in the former, the server can still provide some processing power, albeit at a reduced capacity, whereas in the latter, the server does not provide processing power. An example set of response levels is discussed in more detail with respect to FIG.
[0058] The VEPO orchestration service 302 may be configured to utilize configuration data corresponding to the specification of response levels and their corresponding reduction actions to determine the impact of applying a given level of reduction action to hosts / instances / workloads (collectively, “resources”) in a physical environment. The impact of applying a given action may depend on the current state of the resources at run time. In some embodiments, identifying the impact of a given reduction action may include identifying a set of resources and / or customers to which the action, if implemented, is applicable. Thus, estimated impact is intended to refer to the scope or applicability of a given action or level. Identifying the scope and / or applicability of a given action or level may include identifying the specific resources and / or customers affected by implementing the action or level and / or any suitable attributes associated with the applicable resources and / or customers. As an overly simplistic example, VEPO orchestration service 302 may identify that if an action is taken to shut down all idle hosts, X number of idle hosts among hosts 324 will be affected, that those hosts are associated with a particular customer or number of customers, and / or that the hosts and / or customers are associated with other attributes such as category (e.g., free tier), priority (e.g., high priority), etc.
[0059] The VEPO orchestration service 302 may be configured to estimate the likely impact and / or actual power reduction for one or more of the levels based, at least in part, on configuration data specifying the levels and corresponding actions and any suitable combination of power data 310, account data 312, or host / instance data, etc. For one or more response levels or one or more reduction actions corresponding to a given level, etc., an estimated impact (e.g., what hosts, instances, workloads, customers are likely to be affected, what attributes are associated with the affected hosts, instances, workloads, and / or customers, how many hosts / instances / workloads / customers are affected, etc.) may be identified. In some embodiments, the estimated impact of a given action may be aggregated with the estimated impact for all actions at a given level to determine the estimated impact for a given level.
[0060] The VEPO orchestration service 302 may be configured to determine an estimated impact and / or actual power reduction likely to occur for one or more of the levels based, at least in part, on current, historical, or forecasted data (e.g., current, historical, or forecasted values corresponding to power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318).
[0061] In some embodiments, the VEPO orchestration service 302 may be configured to present, recommend, or automatically select a given response level from a set of possible response levels based, at least in part, on any suitable combination of estimated impact and / or estimated power reduction likely to occur if a reduction action corresponding to the given level is implemented (e.g., realized). Thus, in some embodiments, a particular level may be recommended and / or selected by the power orchestration system 300 based on any suitable combination of: 1) determining a level that, if implemented, is likely to result in a sufficient, but not excessive, reduction in power consumption in resources of the physical environment (e.g., the smallest amount of power consumption reduction sufficient to reduce current power consumption values below the current aggregate power threshold); 2) determining a level that includes the least severe set of actions; or 3) identifying the least impactful level or action (e.g., the fewest number of potentially affected resources / customers, the set of affected resources and / or customers having a priority or category indicating the lowest degree of overall importance, etc.). Determining the least severe level or action, or the level or action having the least impact, may be determined independently of the estimated power consumption reduction likely to occur by implementing that level and / or action, or determining the least severe level / action and / or the level / action having the least impact may be determined from one or more levels that, if implemented given the current conditions, are estimated to cause a reduction in power consumption sufficient to bring the power consumption value down to an aggregated value that falls within the current aggregate power threshold.
[0062] The power management service 304 may be configured to provide functionality for identifying power cap values for hosts and / or instances based, at least in part, on aggregate power thresholds, budget / allocated power thresholds associated with PDUs, and / or current individual and / or aggregate power consumption values. In some embodiments, the power management service 304 may be configured to identify power caps and / or resources to which specific values of those power caps should be applied, based, at least in part, on any suitable combination of current, historical, or predicted values of power data 310, account data 312, environmental data 314, and host / instance data 316. In some embodiments, the VEPO orchestration service 302 can trigger and utilize functionality provided by the power management service 304 to identify / change the budget / allocated power corresponding to one or more resources (e.g., hosts, instances, PDUs, etc.), identify which or how many resources are applicable, and identify estimated power reductions that are likely to occur if a given level of action is implemented. In some embodiments, the power management service 304 can perform some or all of these functions.
[0063] The VEPO orchestration service 302 and / or the power management service 304 can host a user interface 308. The user interface 308 can be configured to provide any suitable application metadata corresponding to a combination of current individual and / or power consumption values corresponding to any suitable number or type of resources (e.g., hosts, instances, workloads, etc.) or customers, any suitable attributes (e.g., priority, category, etc.) associated with those resources or customers, a predicted state (e.g., predicted aggregate power threshold change), a current state (e.g., current aggregate power threshold), an estimated impact (e.g., number, identifier, or other attribute associated with the affected host / instance / workload customer), and / or an estimated power reduction (e.g., an estimated reduction that is expected to occur if a level action is applied), and / or an identification of an action corresponding to a given level. The user interface 308 can present any suitable combination of this application metadata. The application metadata can correspond to any suitable combination of available response levels. In some embodiments, application metadata corresponding to current / projected power consumption values and / or current / projected aggregate power thresholds can be displayed. In some embodiments, the application metadata presented in the user interface 308 can correspond to an estimated power consumption reduction and / or estimated impact for any suitable number of levels. In some embodiments, the VEPO orchestration service 302 can provide a recommendation for a particular level (e.g., via the provided application metadata corresponding to that level or reduction action associated with that level), and confirmation / rejection of the recommended level can be entered via a selectable option in the user interface 308. If confirmed, the VEPO orchestration service 302 can be configured to perform operations to effectuate implementation of the confirmed level and its corresponding action.In some embodiments, user interface 308 may present application metadata for any suitable number of response levels and / or corresponding reduction actions, and one or more options may be provided in user interface 308 to enable user selection of a particular level and / or action.
[0064] Upon receiving user input indicating the selection of a particular level and / or action, or confirmation of the selection of a particular level and / or action, the VEPO orchestration service 302 can perform operations to effectuate implementation of the confirmed level and its corresponding action. Effecting or implementing the confirmed level and / or its corresponding action may include instructing any suitable component of the power orchestration system (e.g., power management service 304, compute service 306) to implement any suitable portion of the level and / or corresponding action. Instructing one or more components may include providing any suitable data to the instructed component, such as any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and operational data 318. In some embodiments, the directed component may be configured to implement an action or level according to data provided by VEPO orchestration service 302, or the directed component may obtain any suitable portion of data from where power data 310, account data 312, environment data 314, host / instance data 316, or operational data 318 are stored.
[0065] As a non-limiting example, if the action at the level at which reduction or impact is to be estimated includes power capping, the power management service 304 may be configured to estimate a power cap, estimate individual or aggregate power consumption reductions, and / or implement the estimated power cap to effect the estimated individual / aggregate power consumption reductions. In some embodiments, the power cap can be implemented directly by the power management service 304 or by instructing the compute service 306. Other actions, such as migrating instances or workloads, pausing instances, workloads, or hosts, shutting down hosts, or preventing (at least temporarily) future allocation or launch of instances and / or workloads, can be identified by the VEPO orchestration service 302 and / or the compute service 306 and implemented by the VEPO orchestration service 302 through functionality provided by the compute service 306. In some embodiments, the compute service 306 may be configured to identify potentially affected hosts / instances / workloads / customers and provide such information to the VEPO orchestration service 302 and / or the power management service 304.
[0066] Compute service 306 may be configured to communicate with a baseboard management controller (BMC) / integrated lights out manager (ILOM) 322 of rack PDU 320. ILOM is one example of a BMC used for illustrative purposes. In some embodiments, BMC / ILOM 322 may be a service processor dedicated to monitoring the physical status of a host machine, in this example, rack PDU 320. Similarly, host 324 (e.g., a host in rack 319, to which rack PDU 320 manages power according to the power distribution hierarchy discussed in connection with FIG. 5 ) may include BMC / ILOM 328, which may be a service processor dedicated to monitoring the physical status of the host, in this example, one of hosts 324. BMC / ILOM 322 and BMC / ILOM 328 may be configured to, among other things, manage and / or enforce budget power and / or power caps at the device in which BMC / ILOM 328 operates or for downstream devices. For example, BMC / ILOM 322 can receive power cap values that are utilized by BMC / ILOM 328 to throttle power consumption for a given host. BMC / ILOM 322 can distribute the power caps to the corresponding BMC / ILOM 328 with which the power caps are associated. In some embodiments, BMC / ILOM 322 can distribute power caps stored on the applicable hosts, but enforcement of these power caps is not realized until further instructions are provided by BMC / ILOM 322 to BMC / ILOM 328. BMC / ILOM 328 can enforce power caps at the host level and / or instance level and / or workload level for one or more workloads (not shown) running on the instances.
[0067] BMC / ILOM 322 can be used to instruct devices (e.g., rack PDU 320 and / or hosts 324) to suspend operation of hosts / instances / workloads or to shut them down. In some embodiments, BMC / ILOM 322 can instruct BMC / ILOM 328 to resume operation of hosts / instances / workloads. If shut down, devices (e.g., via their BMC / ILOM) can be instructed (e.g., by Compute Services 306 and / or by their upstream devices, such as PDUs) to start up from shutdown. This can include Compute Services 306 sending instructions from BMC / ILOM 322 to BMC / ILOM 328 to suspend and / or shut down hosts / instances / workloads.
[0068] In some embodiments, compute service 306 may perform any suitable operation, such as migrating a workload from one instance to another (e.g., to an instance on the same host, to another instance running on a different host, etc.), migrating an instance from one host to another (e.g., to a host in the same rack, to a host in a different rack, etc.), migrating an instance / workload back to the host that originally hosted the previously migrated instance / workload, etc. In some embodiments, compute service 306 may be configured to ensure that a host / instance / workload is not assigned to a host and / or instance that has a level and / or action applied that has been achieved or is in the process of being achieved.
[0069] Any suitable combination of devices (e.g., host 324, rack PDU 320) may include one or more power controllers (e.g., BMC / ILOM 328, BMC / ILOM 322, respectively). A power controller may be any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in the device. A power controller may individually monitor the power consumption of each device on which it operates. The power controllers may each be configured to implement power capping limits to constrain power consumption in the respective device. Implementing power capping limits may include monitoring power consumption in the device, determining whether to constrain (e.g., constrain within a range, constrain, etc.) the power consumption in the device (e.g., based at least in part on a comparison of the device's current power consumption with a stored power cap limit), and constraining / constraining the power consumption in the device (e.g., using dynamic voltage and / or frequency scaling in the device's processor and / or memory to throttle power consumption in the device). Any suitable action associated with power capping may be implemented by the power controller.
[0070] In some embodiments, a power controller (e.g., BMC / ILOM 322) can communicate with another power controller (e.g., BMC / ILOM 328) corresponding to one of the hosts 324) via a direct connection (e.g., via a cable) and / or via a network. The power controller (e.g., BMC / ILOM 328) can provide power consumption data indicating the device's current power consumption (e.g., cumulative power consumption over a period of time, current power consumption rate, etc.). The power consumption data can be provided to the power controller (e.g., BMC / ILOM 322) at any suitable frequency, periodicity, or according to a predefined schedule or event (e.g., when a predefined consumption threshold is breached, when a change in consumption rate reaches a threshold, when one or more predefined conditions are determined to be met, when thermal attributes of the device are determined, etc.).
[0071] In some embodiments, a power controller (e.g., BMC / ILOM 328) can receive a power cap value (also referred to as a "power cap") from a power controller (e.g., BMC / ILOM 322). In some embodiments, additional data can be provided along with the power cap value. By way of example, an indication of whether the power cap value should be applied immediately or not may be included with the power cap value. In some embodiments, the received power cap may be applied / enforced immediately by default. In other embodiments, the received power cap may not be applied / enforced immediately by default.
[0072] When applying / enforcing a power cap, the power controller (e.g., BMC / ILOM 328) can monitor power consumption at the device. This can include utilizing a metering device or software configured to identify / calculate power consumption data for the device (e.g., cumulative power consumption over a period of time, a current power consumption rate, a current change in power consumption rate over a time window, etc.). As part of applying / enforcing a power cap (also referred to as “power capping”), the power controller (e.g., BMC / ILOM 328) can determine whether to constrain (e.g., limit within a range, constrain, etc.) the power consumption at the device (e.g., based at least in part on comparing the device's current consumption rate with a stored power cap value). When constraining power consumption at the device, the power controller (e.g., BMC / ILOM 328) can limit / constrain the power consumption at the device (e.g., using dynamic voltage and frequency scaling on the processor and / or memory to throttle power consumption at the device). In some embodiments, the power controller (e.g., BMC / ILOM 328) can execute instructions to limit / constrain the power consumption of a device when the device's current consumption data (e.g., cumulative consumption, consumption rate over a time window, etc.) approaches an enforced power cap (e.g., breaches a threshold below the power cap). When constraining / limiting (also referred to as "throttling") power consumption at a device, the power controller (e.g., BMC / ILOM 328) ensures that the device's power consumption stays below the power consumption indicated by the power cap. The power controller (e.g., BMC / ILOM 328) can be configured to constrain power at a host generally or according to the instance and / or workload to which the power cap is applicable.
[0073] In some embodiments, the power controller of the host 324 may be configured to allow the host 324 to run unconstrained (e.g., without constraints based on power consumption and power caps) until instructed to apply / enforce a power cap by a power controller (e.g., BMC / ILOM 328). In some embodiments, the power controller (e.g., BMC / ILOM 328) may receive a power cap value that it may or may not be instructed to enforce later, but once received, the power cap value may be stored in memory without being utilized in power management at the device. Thus, in some embodiments, the power controller (e.g., BMC / ILOM 328) may not initiate a power capping operation (e.g., the comparison and determination described above) until instructed to do so (e.g., via an indication provided by the power controller 448). The indication to begin enforcing the power cap may be received along with the power cap value or may be received as a separate communication from the power controller 448.
[0074] According to power distribution hierarchy 500, each power controller (e.g., BMC / ILOM 322) of one or more PDUs may be configured to distribute, manage, and monitor power for any suitable number of devices (e.g., hosts 324 in rack 319). In some embodiments, a power controller (e.g., BMC / ILOM 322) may be a computing agent or program installed on a given PDU (e.g., rack PDU 320 of FIG. 3). A power controller (e.g., BMC / ILOM 322) may be configured to obtain power consumption data from hosts 324 in rack 319 that correspond to the devices with which it is associated (e.g., devices for which a given PDU is configured to distribute, manage, or monitor power). The power controller (e.g., BMC / ILOM 322) may receive the consumption data according to a predefined frequency, periodicity, or schedule implemented by the power controller (e.g., BMC / ILOM 328), and / or the power controller (e.g., BMC / ILOM 322) may receive the consumption data in response to requesting the consumption data from the power controller (e.g., BMC / ILOM 328). The power controller (e.g., BMC / ILOM 322) may be configured to request the consumption data from the power controller (e.g., BMC / ILOM 328) according to a predefined frequency, periodicity, or schedule implemented by the power controller (e.g., BMC / ILOM 322).
[0075] The power controller (e.g., BMC / ILOM 322) can transmit received consumption data to the power management service 402 (directly or via the compute service 306) at any suitable time, according to any suitable frequency, periodicity, or schedule, or as a result of one or more predefined conditions being met (e.g., the rate of change in the individual / cumulative / aggregate consumption of the hosts 324 breaches a threshold). In some embodiments, the power controller (e.g., BMC / ILOM 322) can aggregate power consumption data received from the hosts 324 (e.g., via each device's power controller (e.g., BMC / ILOM 328)) and then transmit the aggregated consumption data to the power management service 402. In some embodiments, the consumption data may be aggregated by host, instance type / category / priority, workload type / category / priority, or customer.
[0076] The power controller (e.g., BMC / ILOM 322) can receive one or more power cap values corresponding to the host(s) 324 at any suitable time (e.g., from the implementation manager 410 of the power management service 402). These power cap values can be calculated by the power management service 402 (e.g., via the constraint specification manager 408). In some embodiments, the power controller (e.g., BMC / ILOM 322) can receive, along with the power cap values, a timing value indicating a duration to be used for the timer. The power controller (e.g., BMC / ILOM 322) can be configured to generate and start a timer having an associated duration / period corresponding to the timing value. When the timer expires, indicating that a period corresponding to the period has elapsed, the power controller (e.g., BMC / ILOM 322) can transmit data to one or more hosts, including an indicator or other suitable data that instructs the one or more hosts to proceed with power capping using the power cap value previously stored in each device. In some embodiments, the power cap value can be provided along with the indicator. In other embodiments, the power cap value may be provided by the power controller (e.g., BMC / ILOM 322) to the device's power controller (e.g., BMC / ILOM 328) immediately upon receiving the power cap value from the power management service 402. Thus, in some embodiments, the power cap value is sent by the power controller (e.g., BMC / ILOM 322) to the power controller (e.g., BMC / ILOM 328) and stored in memory, but is not enforced by the power controller (e.g., BMC / ILOM 322) until the power controller (e.g., BMC / ILOM 328) receives subsequent data from the power controller (e.g., BMC / ILOM 322) that instructs the power controller (e.g., BMC / ILOM 328) to begin power capping.
[0077] The VEPO orchestration service 302, the power management service 304, the compute service 306, the rack PDU 320, and the host 324 (collectively referred to as the "devices of FIG. 3") can communicate over one or more wired or wireless networks (e.g., network 808). Storage devices on which the power data 310, the account data 312, the environment data 314, the host / instance data 316, and the operational data 318 are stored can also communicate with the VEPO orchestration service 302, the power management service 304, the compute service 306, the rack PDU 320, and the host 324 over one or more wired or wireless connections (e.g., over one or more networks). In some embodiments, the network can include any one or combination of many different types of networks, such as a cable network, the Internet, a wireless network, a cellular network, and other private and / or public networks.
[0078] 3 may be any suitable type of computing device, such as, without limitation, a server device, a network device, or any suitable device in a data center. In some embodiments, some of the devices (e.g., rack PDU 320 and host 324) are arranged in a power distribution hierarchy, such as power distribution hierarchy 500 discussed in connection with FIG. 5. In some embodiments, host 324 may correspond to and be represented by a level 1 node of power distribution hierarchy 500. Rack PDU 320 (the “rack leader”) may correspond to and be represented by a higher-level node (e.g., a level 2 node) of the power distribution hierarchy.
[0079] Each of the devices of Figure 3 may include at least one memory. Each of the processors may be implemented in hardware, computer-executable instructions, firmware, or a combination thereof, as appropriate. The computer-executable instructions or firmware implementation of the processor of the devices of Figure 3 may include computer-executable or machine-executable instructions written in any suitable programming language to perform the various functions described.
[0080] The memory of the devices of Figure 3 can store program instructions that are loadable and executable on each processor of a given device, as well as data generated during the execution of these programs. Depending on the configuration and type of user computing device, the memory can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). The devices of Figure 3 can also include additional removable and / or non-removable storage devices, including, but not limited to, magnetic storage devices, optical disks, and / or tape storage devices. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device. In some implementations, the memory can individually include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM.
[0081] Referring more particularly to the contents of the memory, the memory may include an operating system, one or more data stores, and one or more application programs, modules, instances, workloads, or services, such as a VEPO orchestration service 302, a power management service 304, and a compute service 306.
[0082] The devices of Figure 3 may include communication connections that allow the devices to communicate with each other over a network (not shown). The devices of Figure 3 may also include I / O devices such as a keyboard, mouse, pen, voice input device, touch input device, display, speakers, printer, etc.
[0083] 4 illustrates an example architecture 400 of a power management service 402 (an example of power management service 304 of FIG. 3 ) according to at least one embodiment. The power management service may include an input / output processing manager 404 that may be configured to receive and / or transmit any suitable data between the power management service 402 and any other suitable device of FIG. 3 . As shown, the power management service 402 may include a consumption monitoring manager 406, a constraint specification manager 408, and an enforcement manager 410, although more or fewer computing components or subroutines may likewise be utilized.
[0084] Consumption monitor manager 406 may be configured to determine current individual and / or aggregate power consumption values (e.g., corresponding to examples of host 324 in FIG. 3 and server 104 in FIG. 1). Constraint identification manager 408 may be configured to perform any suitable function corresponding to identifying one or more power cap values for a host, instance, or workload (e.g., host 324). Enforcement manager 410 may be configured to enforce and / or trigger enforcement of power caps on any suitable device, according to at least one embodiment. Allocating power may refer to the process of allocating a budgeted amount of power (e.g., expected amount of load / power consumption) to any suitable component. A power cap is an example of a budgeted amount of power.
[0085] The functionality described in connection with power management service 402 may, in some embodiments, be performed by one or more virtual machines implemented within a hosted computing environment (e.g., in hosts 324 corresponding to any suitable number of servers 104 of FIG. 1). A hosted computing environment may include one or more rapidly provisioned and released computing resources, which may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment. A number of example cloud computing environments are provided and described in more detail below with respect to FIGS. 15-18.
[0086] The power management service 450 (e.g., consumption monitor manager 406) may be configured to receive / obtain consumption data from a PDU (e.g., from a rack PDU 320 from a power controller such as BMC / ILOM 322 of FIG. 3). As described above, the consumption data may be aggregated or cumulative for one or more devices. As a non-limiting example, an instance of consumption data received by the power management service 402 may include cumulative / aggregated consumption data for all devices associated with (e.g., managed by) a given PDU. By way of example, the PDU providing consumption data may be a rack PDU (e.g., rack PDU 320), and the consumption data provided by that PDU may include cumulative / aggregated and / or individual consumption data values for all devices in the rack (e.g., hosts in hosts 324 associated with a given rack 319). In some embodiments, the cumulative / aggregated consumption data value may further include the power consumption of the PDU. From this consumption data, the power management service 402 can access individual and / or aggregate or cumulative power consumption data for the individual device or all of the devices with which that instance of consumption data is associated. The power management service 402 can be configured to perform any suitable operation (e.g., aggregating the consumption data, calculating power cap values, calculating timing values, determining budget power values, determining whether power capping is necessary (e.g., when the aggregated consumption has breached or is likely to breach the budget power, etc.) based, at least in part, on data representing the power distribution hierarchy 500 of FIG. 5 and / or any suitable configuration data specifying the placement of components within the data center (e.g., indicating which devices distribute power to which devices). By way of example, the configuration data can include data representing the power distribution hierarchy 500.
[0087] FIG. 5 illustrates an example power distribution hierarchy 500 corresponding to the arrangement of components in FIG. 2 , according to at least one embodiment. Power distribution hierarchy 500 may represent the arrangement of any suitable number of components of an electric power system, such as the power distribution infrastructure components discussed in connection with power distribution infrastructure 200 of FIG. 2 . Power distribution hierarchy 500 may include any suitable number of nodes organized according to any suitable number of levels. Each level may include one or more nodes. A root level (e.g., level 5) may include a single root node of power distribution hierarchy 500. Each node of power distribution hierarchy 500 may represent a corresponding component of power distribution infrastructure 200. A set of one or more nodes at a given level may be derived from a particular node corresponding to a higher level of power distribution hierarchy 500. A set of lower-level components represented by a lower-level (e.g., Level 1) node can receive power distributed through a higher-level component represented by a Level 2 node (e.g., a component upstream from the lower-level component), which in turn receives power from a higher-level component represented by a Level 3 node, which in turn receives power from a higher-level component represented by a Level 4 node. In some embodiments, all components of the power system receive power initially distributed through a component corresponding to a Level 5 (e.g., the top level of the power distribution hierarchy 500) node (e.g., node 502, the root node). Node 502 can receive power from a utility power source (e.g., a local electric utility system).
[0088] As shown, power distribution hierarchy 500 includes node 502 at level 5. In some embodiments, node 502 may represent an uninterruptible power supply, such as UPS 202 of FIG. 2. A component corresponding to node 502 may distribute / supply power to a component corresponding to node 504 at level 4. A component (e.g., a lower-level component) that receives power from a higher-level component (a component represented by a node at a higher level than the level of a node representing a lower-level component in power distribution hierarchy 500) may be considered subordinate to the higher-level component. In some embodiments, node 304 may represent a component, such as one of intermediate PDUs 204 of FIG. 2 (e.g., a power distribution panel), that is subordinate to the UPS represented by node 502. The component represented by node 504 may be configured to distribute / supply power to components represented by nodes 506 and 508, respectively.
[0089] Level 3 nodes 506 and 508 may each represent a respective component (e.g., a respective string PDU) in FIG. 2. The components (e.g., string PDU 206, bus bars) corresponding to node 506 may distribute / supply power to the components (e.g., rack PDU 208A in FIG. 2) corresponding to node 510 and the components (e.g., rack PDU 208N in FIG. 2) corresponding to node 512. The components (e.g., rack PDU 208A) corresponding to node 510 may distribute / supply power to the components (e.g., rack PDU 208A) corresponding to node 510 (representing servers 212A and 212B in FIG. 2, respectively) corresponding to level 1 nodes 514 and 516, which may be monitored / managed by components of servers 212A and 212B, such as power controllers 216A and 216B. The components corresponding to nodes 514 and 516 may be located in the same rack.
[0090] A component corresponding to node 512 (e.g., rack PDU 208A) can distribute / supply power to a component corresponding to level 1 node 518 (e.g., server 214C in FIG. 2 including power controller 216C). Level 1 nodes 514, 516, and 518 can be located in and / or associated with the same column (e.g., column 210 in FIG. 2).
[0091] Returning to node 508 in level 3, node 508 (e.g., a different row PDU, a remote power panel) can distribute / supply power to components corresponding to node 510 (e.g., rack PDU 208A in FIG. 2 ) and components corresponding to node 512 (e.g., rack PDU 208N in FIG. 2 ). Components corresponding to node 520 (e.g., a rack PDU) can distribute / supply power to components corresponding to nodes 522 and 526 in level 1 (e.g., corresponding to respective servers including respective power controllers). Components corresponding to nodes 524 and 526 can be located in the same rack. Components corresponding to node 522 (e.g., another rack PDU) can distribute / supply power to components corresponding to nodes 528 and 530 in level 1 (e.g., corresponding to respective servers including respective power controllers). Components corresponding to nodes 528 and 530 can be located in the same rack.
[0092] The particular number of components (e.g., corresponding to level 1 nodes) that receive distributed power from components at a higher level (e.g., corresponding to level 2 nodes) may differ from the number shown in FIG. 5. The particular number of levels in the power distribution hierarchy may vary depending on the particular arrangement of components used in a given data center. It is contemplated that each non-root level (e.g., levels 1-4 in the example of FIG. 5) may include a different number of nodes representing a different number of components than the number of nodes shown in each non-root level in FIG. 5. The nodes may be arranged in different, yet similar, configurations to those shown in FIG. 5.
[0093] Returning to FIG. 4 , the power management service 402 (e.g., the constraint identification manager 408) can calculate aggregated / cumulative consumption data to identify aggregated / cumulative consumption data for one or more devices at higher levels of the power distribution hierarchy 500. By way of example, the power management service 402 (e.g., the constraint identification manager 408) can utilize consumption data provided by one or more rack PDUs to calculate aggregated / cumulative consumption data for column devices. For example, consumption data corresponding to devices represented by the level 1 nodes of FIG. 5 that share a common set of nodes up to level 3 of the hierarchy. In this case, the common level 3 node represents a column PDU, such as a bus bar. In some embodiments, the consumption data provided to the power management service 402 can include consumption data indicating the power consumption of one or more rack PDUs (e.g., example components corresponding to level 2 of the power distribution hierarchy 500). The consumption data provided by the rack PDUs can be provided according to any suitable frequency, periodicity, or schedule, or in response to a request sent by the power management service 402. The power management service 402 may obtain consumption data associated with any suitable number of devices in FIG. 3 , such as a host 324 and / or a rack of devices corresponding to one or more racks (e.g., rack 319 ), and may aggregate or calculate any suitable consumption data corresponding to any suitable number of components represented by any suitable node and / or level of the power distribution hierarchy 500 .
[0094] The power management service 402 (e.g., constraint identification manager 408) may be configured to calculate one or more power caps (e.g., power cap values for one or more servers to which power is distributed by a higher-level device (e.g., a row-level device in this example) based, at least in part, on the allocated power values (e.g., amounts of budgeted power) for that higher-level component. The higher-level component may correspond to any suitable level (e.g., levels 2-5) of the power distribution hierarchy 500 other than the lowest level (e.g., level 1). As an example, the power management service 402 may calculate a power cap value for a server to which power is distributed by a given row device (e.g., bus bar). These calculations may be based on consumption data provided by rack PDUs to which power is distributed by the column devices. In some embodiments, the power management service 402 may store the consumption data for subsequent use. The power management service 402 may utilize historical consumption data when calculating these power cap values. In some embodiments, the power management service 402 may obtain, utilize, and / or train one or more machine learning models from the historical consumption data to identify specific power cap values for one or more hosts 324 (e.g., components corresponding to level 1 of the power distribution hierarchy). These techniques are discussed in more detail with respect to FIG. 12.
[0095] The power management service 402 (e.g., the enforcement manager 410) can calculate timing values for timers (e.g., timers that may be initiated and managed by a PDU, such as the rack PDU 320). The timing values can be calculated based, at least in part, on any suitable combination of the rate of change of the power consumption of the higher-level components, the direction (increase / decrease) of the change in power consumption of the higher-level components, the power spike tolerance of the higher-level components and / or the system 300 as a whole, or the associated enforcement time for the lower-level devices (e.g., an estimated / known delay between when each of the devices (e.g., the host 324) is instructed to enforce a power cap and when each of the devices (e.g., the host 324) will actively enforce the cap (e.g., the first time that the power consumption at that device is constrained / throttled, or at least a determination is made as to whether to constrain / throttle).
[0096] The power management service 402 (e.g., implementation manager 410) can communicate the calculated timing values to the PDU 402 at any suitable time. In some embodiments, the power management service 450 can first determine a power cap for a given higher-level component (e.g., a column component) regardless of the consumption occurring for other components at the same level (e.g., other column components). In some embodiments, while the timer is being initialized or elapses, or at any suitable time, the power management service 402 can process consumption data corresponding to other same-level components to determine whether it is more desirable to set a power cap for one or more downstream components of those same-level components. In some embodiments, the power management service 402 can utilize priorities associated with devices and / or workloads running on those devices to determine power cap values (e.g., priorities identified based at least in part on the host / instance data 316 of FIG. 3 ). The power management service 402 can be configured to prioritize power cap settings for lower priority devices / workloads while leaving higher priority devices / workloads unconstrained. In some embodiments, the power management service 402 may be configured to prioritize power capping of a set of highest consuming devices (e.g., across a row, across multiple rows, etc.) while allowing lower consuming devices to operate unconstrained. In some embodiments, a particular consumption of a device may remain unconstrained even if the device is included in the set of highest consuming devices if the priority associated with that device and / or workload is high (higher than the priority associated with other devices and / or workloads). Thus, the power management service 402 may be configured to prioritize the priority of a device / workload over the power consumption of that particular device.
[0097] In some embodiments, for a given string, power cap values may be initially determined by the power management service 402 (e.g., by the constraint identification manager). These power cap values may be provided to a power controller (e.g., BMC / ILOM 320 or another agent or component of the rack PDU 320) of the rack PDU, which may then distribute the power cap values to a power controller (e.g., BMC / ILOM 328) of the host 324 for storage. The power management service 402 (e.g., constraint identification manager 408) may process consumption data for other string devices in the same string to determine power cap values for devices corresponding to a different string. This may be advantageous because devices managed by another string may not be consuming their budgeted power, leaving some amount of power unused for devices at that string level. In some embodiments, the power management service 402 (e.g., constraint identification manager 408) may be configured to determine whether it may be more advantageous to set power caps for devices in one string while allowing at least some devices in another string to operate unconstrained. Determining the benefit of a set of power cap values may be based, at least in part, on minimizing the estimated impact (e.g., the number of devices that should be power capped), minimizing the priority values associated with devices that should be power capped, and maximizing the number of devices associated with a particular high priority value that will not be power capped. The power management service 402 may utilize a predefined protocol (e.g., a set of rules) to determine whether implementing a power cap value that it has already sent to one PDU is substantially more advantageous / beneficial than a different power cap setting that it has identified based on processing consumption data from multiple PDUs associated with one or more other strings.
[0098] The power management service 402 (e.g., the enforcement manager 410) may be configured to perform enforcement of a set of power cap values determined to be most advantageous (based on a predefined set of rules) to ensure that one or more circuit breakers in the data center do not trip and / or that the aggregate power threshold is enforced. If the current / aggregate power consumption level exceeds the current aggregate power threshold, the power management service 402 may be configured to perform or trigger the execution of an action to bring the current / aggregate power consumption level below the current aggregate power threshold.
[0099] In some embodiments, power cap-based monitoring, identification, and / or constraints can be determined by the power management service 402 directly based on power data 310, account data 312, environmental data 314, host / instance data 316, operational data 318, etc.
[0100] In some embodiments, the power management service 402 may be configured to identify and / or enforce power caps based, at least in part, on instructions and / or constraints or other applicable data provided by the VEPO orchestration service 302. As a non-limiting example, the VEPO orchestration service 302 may filter aggregated data of any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318 based on a reduction action associated with a level (e.g., "set a 10% power cap on all low priority hosts") to determine the estimated impact of implementing the action, including the number and / or subset of resources (e.g., hosts, instances, workloads, customers) to which the action will be applied if implemented. As an example, the VEPO orchestration service 302 can identify a number and / or identifiers corresponding to all low-priority hosts from any suitable combination of aggregated data: power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the identifiers and / or instructions (e.g., set a 10% power cap) can be provided to the power management service 402, received through the input / output processing module 404, and implemented by the enforcement manager 410. As another non-limiting example, any suitable portion of metadata associated with a curtailment action (e.g., "set a 10% power cap on all low-priority hosts") can be provided to the power management service 402, and the constraint identification manager 408 can be configured to identify the particular hosts / instances / workloads from the metadata associated with the curtailment action. Stated another way, any suitable combination of the VEPO orchestration service 302 and / or the power management service 402 (e.g., the power management service 304) can be configured to identify the estimated impact of a given action.In some embodiments, the power management service 402 (e.g., the constraint identification manager 408) may be configured to identify estimated power reduction values corresponding to individual or aggregate power consumption of one or more devices (e.g., hosts 324) taking into account the hosts / instances / workloads / customers expected to be affected if the action (or level) is implemented. If the set of power caps previously provided to the power controller (e.g., BMC / ILOM 322 and / or BMC / ILOM 322) is determined to be most advantageous, the power management service 402 can send data to the power controller to cause the power controller to send an indication to the device to initiate power capping based on the previously allocated power caps. Alternatively, if the power management service 402 determines that a new set of power caps is more (or most) advantageous from a power management perspective, the power management service 402 can send data to the power controller (e.g., BMC / ILOM 322) to cancel a timer. In some embodiments, canceling the timer may cause the previously allocated power caps to be deleted by an instruction from the power controller (e.g., sent in response to canceling the timer), or the power caps may time out by default (e.g., according to a predefined period). The power management service 402 can send the new power caps to the appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was sent) along with an indication that the power caps should be immediately enforced by the corresponding device. These PDUs can transmit the power caps to the receiving device along with instructions to immediately initiate power capping operations (including, for example, monitoring consumption relative to the power cap value, determining whether the cap is based on the power cap value and the device's current consumption, and either capping the device's operation or leaving such operation unconstrained based on the determination).
[0101] In some embodiments, if the timer expires (e.g., the period corresponding to the timing value has passed) on the original PDU and a cancellation has not been received from the power management system 402, the power controller (e.g., BMC / ILOM 322) can automatically instruct the device (e.g., via BMC / ILOM 328) to initiate a power capping operation at the stored power cap value. This technique ensures a fail-safe in case the power management service 402 fails to instruct the PDU or cancel the timer for any reason.
[0102] The techniques described above allow more lower-level devices to operate unconstrained by minimizing the number and / or frequency at which device power consumption is capped. Furthermore, devices downstream of a given device in the hierarchy can be allowed to peak above the device's budgeted power while still ensuring that the device's maximum power capacity (a value higher than the budget value) is not exceeded. This allows the consumption of lower-level devices to operate within a buffer of power that conventional systems would leave unused. The present system and described techniques allow for more efficient utilization of the power distribution components of system 400, reducing waste while ensuring that power failures are avoided.
[0103] In some embodiments, a set of power cap values determined to be most advantageous (e.g., based at least in part on any suitable combination of maximum expected reduction and / or minimum estimated impact, etc.) may be selected by the power management service 402, or a particular set of power cap values (previously identified by the power management service 402 and provided to the VEPO orchestration service 302) may be provided from and / or implemented by the power management service 402. In some embodiments, the particular power cap values identified may be identified at least in part based on maximizing power consumption reductions, identifying sufficient power cap-related reductions, individually or collectively, along with other actions corresponding to reduction actions or levels associated with lowering the device's aggregate power consumption value below the current aggregate power threshold.
[0104] 6 is a flow diagram illustrating an example method 600 for managing excessive power consumption, according to at least one embodiment. Method 600 may include more or fewer operations than those shown or described with respect to FIG. 6. The operations may be performed in any suitable order. Any suitable portion of the operations described with respect to power management service 606 may additionally or alternatively be performed by VEPO orchestration service 302 and / or compute service 306 of FIG. 3.
[0105] Method 600 may begin at 610, where consumption data may be received and / or obtained (e.g., upon request) by PDU 604 from devices 602. Devices 602 may be examples of hosts 324 of FIG. 3 (e.g., multiple servers) and may be located within a common server rack (e.g., rack 319). PDU 604 is an example of rack PDU 208A of FIG. 2 and / or rack PDU 320 of FIG. 3. Each of devices 602 may be a device to which PDU 604 distributes power. PDU 608 may be an example of another rack PDU associated with the same column as PDU 604 and therefore receiving power from the same column PDU as PDU 604 (e.g., column PDU 206 of FIG. 2, corresponding to node 506 of FIG. 5). The consumption received from PDU 608 may correspond to the power consumption of devices 612. In some embodiments, PDU 608 and at least some portions of devices 612 correspond to a different column than the column corresponding to devices 602. In some embodiments, any suitable number of instances of consumption data (referred to as “device consumption data”) received by PDU 604 may relate to a single one of devices 602. The same may be true for device consumption data received by PDU 608. In some embodiments, PDU 604 and PDU 608 can aggregate and / or perform calculations from the received device consumption data instances to generate rack consumption data corresponding to each respective PDU. For example, PDU 605 can generate rack consumption data from device consumption data provided by device 602. The rack consumption data may include device consumption data instances received by PDU 604 at 610. Similarly, the rack consumption data provided by PDU 608 may include device consumption instances received from device 612. The rack consumption data generated by PDU 604 and / or PDU 608 can be generated, at least in part, based on the power consumption of PDU 604 and / or PDU 608, respectively.
[0106] At 614, power management service 606 (an example of power management service 402 in FIG. 4 , power management service 304 in FIG. 3 ) can receive and / or obtain (e.g., via a request) rack consumption data from any suitable combination of PDU 604 and PDU 608 (which are also examples of PDU 208 in FIG. 2 ), not necessarily simultaneously. In some embodiments, power management service 606 can continuously receive / obtain rack consumption data from PDU 604 and PDU 608. PDU 604 and PDU 608 can similarly receive and / or obtain (e.g., via a request) device consumption data from devices 602 and 612, respectively (e.g., from their corresponding power controllers, each power controller being an example of BMC / ILOM 328 in FIG. 3 ). PDU 608 can correspond to a PDU associated with the same column as PDU 604.
[0107] At 616, the power management service 606 may determine whether to calculate a power cap value for the device 602 based, at least in part, on maximum and / or budget power amounts associated with another PDU (e.g., a string-level device not shown in FIG. 5 , such as a bus bar) and consumption data received from PDU 604 and / or PDU 608. In some embodiments, the power management service 606 may access these maximum and / or budget power amounts associated with a string-level PDU (e.g., string PDU 206 of FIG. 2 , not shown here) from stored data, or the maximum and / or budget power amounts may be received by the power management service 606 at any suitable time from the PDU to which the maximum and / or budget power amounts are associated (the string-level PDU). The power management service 606 can aggregate the rack consumption data from PDU 604 and / or PDU 608 to determine cumulative power consumption values for the row-level devices (e.g., total power consumption from devices downstream of the row-level device within the last time window, a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, a period of time required to initiate power capping on one or more of devices 602 and / or 612, etc.). In some embodiments, the power management service 606 can generate additional cumulative power consumption data from at least a portion of the cumulative power consumption data values and / or using historical device consumption data (e.g., cumulative or individual consumption data). For example, the power management service 606 can calculate any suitable combination of a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, etc. based, at least in part, on historical consumption data (e.g., historical device consumption data and / or historical rack consumption data corresponding to devices 602 and / or 612).
[0108] Alternatively, in some embodiments, the power management service 606 may receive 616 from the VEPO orchestration service 402 any suitable data and / or instructions identifying any suitable portion of the curtailment action (e.g., setting power caps only on low-priority hosts), and / or any suitable constraints, and / or any suitable devices to which the constraints apply. The power management service 606 may identify a subset of hosts for which power caps should be calculated (e.g., in this example, from only low-priority host candidates). For example, an action to set power caps only on low-priority hosts may result in the power management service 606 identifying, from all low-priority hosts (e.g., subset of hosts 324), the number of power caps for those subset of low-priority hosts.
[0109] The power management service 606 may determine that a power cap should be calculated if the cumulative power consumption (e.g., the power consumption corresponding to devices 602 and 612) exceeds the budgeted power amount associated with the column-level PDU. If the cumulative power consumption does not exceed the budgeted power amount associated with the column-level PDU, the power management service 606 may determine that a power cap should not be calculated, and method 600 may end. Alternatively, the power management service may determine that a power cap should be calculated due to the cumulative power consumption of downstream devices (e.g., devices 602 and 612) exceeding the budgeted power amount associated with the column-level PDU, and may proceed to 618.
[0110] In some embodiments, the power management service 606 can additionally or alternatively determine that a power cap value should be calculated at 616 based, at least in part, on providing historical consumption data (e.g., device consumption data and / or rack consumption data corresponding to devices 602 and / or 612) as input to one or more machine learning models. The machine learning models can be trained using any suitable supervised or unsupervised machine learning algorithm to identify, from the historical consumption data provided as input, a likelihood that consumption corresponding to a device associated with the historical consumption data will exceed a budgeted amount (e.g., a budgeted amount of power allocated to a row-level device). The machine learning models can be trained using a corresponding training dataset that includes historical consumption data instances. Training of these machine learning models is discussed in more detail with respect to FIG. 12 . Using the techniques described above, the power management service 506 can identify that a power cap value should be calculated based on cumulative device / rack consumption levels that exceed the budgeted power associated with the row-level devices and / or based on determining, from the output of the machine learning model, that the cumulative device / rack consumption levels are likely to exceed the budgeted power associated with the row-level devices. Determining that a device / rack consumption level is likely to exceed the row-level device's power budget can be determined based on receiving an output from a machine learning model indicating the likelihood (e.g., likely / unlikely) or by comparing an output value (e.g., a percentage, confidence value, etc.) indicating the likelihood of exceeding the row-level device's power budget with a predefined threshold. The output value indicating the likelihood of exceeding the predefined threshold can enable the power management service 606 to determine that power capping is warranted and that a power cap value should be calculated.
[0111] At 618, the power management service 606 can calculate power cap values for any suitable number of devices 602 and / or devices 612. By way of example, the power management service 606 can utilize device consumption data and / or rack consumption data corresponding to each of the devices 602 and 612. In some embodiments, the power management service 606 can determine the difference between the cumulative power consumption (calculated from aggregating the device and / or rack consumption data corresponding to devices 602 and 612) of a column of devices (e.g., device 602 and any of the devices 612 in the same column) and the budgeted power associated with the column-level devices. This difference can be used to identify the amount by which power consumption should be constrained at the devices associated with the column. The power management service 606 can determine power cap values for any suitable combination of devices 602 and / or 612 for devices in the same column based on identifying power cap values that, when implemented by the devices, would reduce power consumption to a value less than the budgeted power associated with the column-level devices.
[0112] In some embodiments, the power management service 606 can determine specific power cap values for any suitable combination of devices 602 and 612 corresponding to the same column based, at least in part, on consumption data and / or priority values associated with those devices and / or workloads associated with the power consumption of those devices. In some embodiments, these priority values may be provided as part of device consumption data provided by the devices whose consumption data are associated, or the priority values may be ascertained, at least in part, based on the type of device, the type of workload, etc., obtained from the consumption data or any other suitable source (e.g., from separate data accessible to the power management service 606). Using the priorities associated with each device or workload, the power management service 606 can calculate power cap values such that it prioritizes setting power caps on devices associated with lower priority workloads over setting power caps on devices with higher priority workloads. In some embodiments, the power management service 606 may calculate power cap values based, at least in part, on prioritizing setting power caps on devices consuming at higher rates (e.g., the set of highest consuming devices) over setting power caps on other devices consuming at lower rates. In some embodiments, the power management service 606 may calculate power cap values based, at least in part, on a combination of factors including the consumption value for each device and the priority associated with the device or the workload being executed by the device. The power management service 606 may identify power cap values for devices that entirely avoid capping high-consuming and / or high-priority devices or that cap those devices to a lesser extent than devices that consume less power and / or are associated with a lower priority.
[0113] At 620, the power management service 620 can calculate timing values corresponding to durations for timers that the PDU (e.g., PDU 604) should initialize and manage. The timing values can be calculated based, at least in part, on the rate of change of power consumption from the perspective of the column-level devices, the direction of change of power consumption from the perspective of the column-level devices, or any suitable combination of the time during which power capping should be performed at each of the devices to which the power cap calculated at 618 is associated. As an example, the power management service 620 can determine a timing value that is greater than timing values for smaller increases in power consumption from the perspective of the column-level devices based on determining a relatively large increase in power consumption from the perspective of the column-level devices. Thus, a larger increase in the power consumption rate can result in a smaller timing value (corresponding to a shorter timer), while a smaller increase in the power consumption rate can result in a larger timing value (corresponding to a longer timer). Similarly, the calculated first timing value can be smaller than the calculated second timing value if the time during which power capping should be performed at each of the devices to which the power cap is associated is shorter relative to the calculation of the first timing value. Thus, the faster a device for which power capping is intended can implement the power cap, the shorter the timing value may be. In some embodiments, timers are not calculated for power caps determined via data provided or directed by the VEPO orchestration service 302.
[0114] At 622, the calculated power cap value calculated at 618 and / or the timing value calculated at 620 may be transmitted to the PDU 604. Although not shown, corresponding power cap values and / or timing values for any of the devices 612 in the same row may be transmitted at 620. Although not shown, in some embodiments, the transmission at 622 may be based at least in part on instructions received from the VEPO orchestration service 302 to enforce / implement the power cap value calculated at 618.
[0115] At 624, if a timing value was provided at 622, the PDU 604 may use the timing value to initialize a timer having a duration corresponding to the timing value. As noted above, in instances where the power management service 606 is directed by the VEPO orchestration service 302, the timing value need not be provided at 622 or stored at 624.
[0116] At 626, the PDU 604 may transmit the power cap value (if received at 622) to the device 602. The device 602 may store the received power cap value at 628. In some embodiments, the power cap value provided at 626 may include an indicator that enforcement should not begin, or may not provide an indicator, and the device 602 may not initiate power capping by default.
[0117] At 630, the power management service 606 can perform a high-level analysis of device / rack consumption data received from any of the devices 612 corresponding to a different row than the row to which the device 602 corresponds. As part of this process, the power management service 606 can identify unused power associated with other row-level devices. If unused power exists, the power management service 606 can calculate a new set of power cap values based, at least in part, on consumption data corresponding to at least one other row of devices. In some embodiments, any suitable power cap values identified by the power management service 606 can be provided to the VEPO orchestration service 302 at any suitable time. As discussed above, the new set of power cap values can be advantageous because devices managed by another row-level device may not be consuming their budgeted power, leaving unused power for that row-level device. In some embodiments, the power management service 606 may be configured to determine whether it may be more advantageous to allow at least some of the devices 602 (and possibly some of the devices 612 corresponding to the same column as device 602) to operate unconstrained while setting caps on devices in another column. The benefits of each power capping approach may be calculated based, at least in part, on minimizing the number of devices for which power caps should be set, minimizing the priority values associated with devices for which power caps should be set, maximizing the number of devices associated with a particular high priority value for which no power caps should be set, etc. The power management service 606 may utilize a predefined scheme or set of rules to determine whether implementing a power cap value it has already determined for one column (e.g., corresponding to the power cap value transmitted in 622) is effectively more advantageous than a different power cap setting identified based on processing consumption data from multiple columns.In some embodiments, each set of power cap values identified by the power management service 606 can be beneficial, and the most advantageous set of power caps can be selected by the power management service 606 (e.g., according to a set of rules, according to user input selecting a particular response level, etc.).
[0118] The power management service 606 may be configured to execute the implementation of the set of power cap values (or the set of power cap values instructed by the VEPO orchestration service 302) that is determined to be more (or most) favorable to ensure that one or more circuit breakers in the data center do not trip and / or that aggregate power thresholds are implemented. If the set of power cap values previously provided to the PDU 604 is determined to be less favorable than the set of power cap values calculated in 630, the power management service 606 may immediately send data to the PDU 604 to cause the PDU 604 to send an indication to the device 602 to immediately begin power capping based on the previously allocated power cap. Alternatively, the power management service 606 may take no further action, allowing a timer in the PDU 604 to expire. Any suitable action discussed as being executed by the power management service 606 based on a determination or identification made by the power management service 606 may alternatively be triggered by direction of the VEPO orchestration service 302.
[0119] If the power management service 606 determines that the new set of power caps is more (or most) advantageous from a power management perspective, the power management service 606 may send data to the power controller (e.g., the BMC / ILOM 322 in FIG. 3 ) to cancel the timer. The previously allocated power caps may be deleted by instruction from the power controller (e.g., the BMC / ILOM 322) (e.g., sent in response to canceling the timer), or the power caps may time out by default according to a predefined period. The power management service 606 may send the new power caps in appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was sent) with an indication that the power caps should be immediately enforced by the corresponding devices. These PDUs may send the power caps to the receiving devices with instructions to immediately initiate power capping operations (including, for example, monitoring consumption with respect to the power cap values, determining whether the caps are based on the power cap values and the device's current consumption, and capping the device's operation or leaving such operation unconstrained based on the determination).
[0120] If it is determined that the power cap value calculated at 630 is more favorable from a power management perspective than the power cap value calculated at 618 and ultimately stored at the device 602 at 628, the method 600 proceeds to 632, where the power cap value calculated at 630 may be transmitted to a PDU 608 (e.g., any of the PDUs 608 that manage the device to which the power cap value is associated). In some embodiments, the power management service 606 may provide an indication that the power cap value transmitted at 632 should be implemented immediately. The operation described at 632 may be triggered by the power management service 606 based on a determination / identification made by the power management service 606 or in response to receiving an instruction from the VEPO orchestration service 302.
[0121] In response to receiving the power cap values and indications at 632, the PDU 608 that distributes power to the devices to which the power cap values are associated can transmit the power cap values and indications to the devices to which the power cap values are associated. At 636, the receiving device among the devices 612 can perform a power capping operation without delay based on receiving the indication, which includes 1) determining whether to limit / constrain operation based on the current consumption of a given device as compared to the power cap value provided for that device, and 2) limiting / constraining or not limiting / constraining the power consumption of that device within the range.
[0122] In some embodiments, the power management service 606 can transmit data at 622 canceling the timer and / or power cap value transmitted based, at least in part, on determining that the power cap value calculated at 630 is more favorable (or most favorable) than that calculated at 618. In some embodiments, the PDU 604 can be configured to cancel the timer at 640. In some embodiments, the PDU 604 can transmit data at 642 that causes the device 602 to delete the power cap value from memory at 644.
[0123] At 646, in the situation where a more favorable set of power cap values is not found as described above or the power management service 606 has not sent a cancellation as described above in connection with 638, the PDU 604 may determine that the timer started at 624 has expired (e.g., the duration corresponding to the timing value provided at 622 has elapsed). As a result, the PDU 604 may send an indication at 648 to the device 602 instructing the device 602 to initiate power capping based on the power cap values stored at 628 in the device 602. Upon receiving this indication at 648, the device 602 may initiate power capping at 650 according to the power cap values stored at 628. In some embodiments, the VEPO orchestration service 302 may be configured to control the operations of 646-650, and the operations of 646-650 may not be performed by the power management service 606 unless directed to do so by the VEPO orchestration service 302.
[0124] In some embodiments, the PDU 604 may be configured to cancel the timer at any suitable time after the timer is started at 624 if the PDU 604 determines that the device 602 has reduced its power consumption, which in some embodiments may cause the PDU 604 to send data to the device 602 to cause the device 602 to discard the previously stored power cap value from local memory.
[0125] In some embodiments, the power cap value implemented in any suitable device may be allowed to time out or may be replaced at any suitable time with a power cap value calculated by the power management service 606. In some embodiments, the power management service 606 may send a power cap cancellation or replacement to any suitable device via the corresponding rack PDU. If a power cap cancellation is received, the device may delete the previously stored power cap value, allowing the device to resume unconstrained operation. If a replacement power cap value is received, the device may store the new power cap value and implement the new power cap value immediately or when instructed by its rack PDU. In some embodiments, being instructed by its PDU to implement a new power cap value may cause a device (e.g., power controller 446) of the devices to replace the power cap value previously utilized to implement the power cap setting with the new power cap value that the device is instructed to implement.
[0126] Any suitable operations of method 600 may be performed continuously or at any suitable time to manage excess consumption (e.g., a situation in which the consumption of a server corresponding to a column-level device exceeds the column-level device's budgeted power) to enable more efficient use of previously unused power while avoiding power outages due to circuit breaker tripping. It should be understood that similar operations may be performed by power management service 606 with respect to any suitable level of power distribution hierarchy 500. While examples have been provided with respect to monitoring power consumption corresponding to column-level devices (e.g., devices represented by level 3 nodes of power distribution hierarchy 500), similar operations may be performed with respect to any suitable level of higher-level devices (e.g., components corresponding to any of levels 2-5 of power distribution hierarchy 500). As discussed above, any suitable function / operation of power management service 606 may be triggered, at least in part, based on a determination / determination performed by power management service 606, or any suitable function / operation of power management service 606 may be triggered via direction by VEPO orchestration service 302.
[0127] 7 illustrates an example architecture of a VEPO orchestration service 702 (an example of VEPO orchestration service 302 of FIG. 3) configured to orchestrate management of host power consumption relative to aggregate power thresholds, according to at least one embodiment. VEPO orchestration service 702 may include an input / output processing manager 704 that may be configured to receive and / or transmit any suitable data between VEPO orchestration service 702 and any other suitable device of FIG. 3. As shown, power management service 402 may include a demand manager 710, an impact identification manager 706, and an implementation manager 708, although more or fewer computing components or subroutines may likewise be utilized.
[0128] Demand manager 710 may be configured to monitor and / or modify aggregate power thresholds associated with a physical environment (e.g., data center 102, room 110A, column 108A of FIG. 1, etc.). Monitoring may include monitoring any suitable combination of power data 310, environmental data 314, host / instance data 316, and / or operational data 318 of FIG. 3. The modification of the aggregate power threshold may be performed, and the amount by which the aggregate power threshold may be modified may be determined, at least in part, based on a number of triggering events. These triggering events may include, but are not limited to, 1) receiving a request from a governmental agency (e.g., a local power authority) requesting a power reduction indefinitely or for a period of time; 2) determining current and / or predicted environmental conditions (e.g., current ambient or external temperature, predicted ambient or external temperature); 3) determining current or predicted aggregate power consumption corresponding to devices of the physical environment; and 4) determining the current and / or predicted operational status of one or more temperature control units of the physical environment. The amount by which the aggregate power threshold can or should be changed can be determined, at least in part, based on a number of factors, including, but not limited to, the difference between ambient temperature and external temperature, the amount of power consumption attributable to or associated with a particular component (e.g., a particular temperature control unit), an amount or percentage specified in a request received from a government agency, or the difference between a current or predicted aggregate power consumption value and a current or predicted aggregate power threshold.
[0129] The impact identification manager 706 may be configured to generate (directly or by invoking functionality of another service or process, such as the compute service 306 and / or the power management service 304) an estimated impact corresponding to the response level and / or one or more instances of a reduction action corresponding to the response level. The VEPO orchestration service 702 may obtain configuration data 712 that specifies a set of multiple VEPO levels and their corresponding one or more reduction actions.
[0130] 8 is a table 800 illustrating an example set of response levels and a respective set of reduction actions for each response level, according to at least one embodiment. The data in table 800 can be provided in different ways, such as via a mapping or embedded within the program code of the VEPO orchestration service 702. Each response level in the set of response levels can be associated with a corresponding set of reduction actions that can be performed on multiple resources (e.g., hosts, instances, workloads, etc.) of the data center, with each corresponding set of reduction actions associated with implementing a different reduction in the aggregate power consumption of the data center. In some embodiments, the multiple response levels indicate increasing severity when the reduction actions are performed to reduce the aggregate power consumption of the data center.
[0131] By way of example, table 800 includes four VEPO response levels, although any suitable number of response levels may be utilized. VEPO response level 1 may be associated with one curtailment action. Specifically, the curtailment action specifies that all free and / or idle hosts and hypervisors that can be fully evacuated (e.g., via migration) should be shut down (e.g., powered off).
[0132] In some embodiments, VEPO Response Level 1 may include stopping or preventing any launch or assignment of instances and / or workloads to the affected hosts. In some embodiments, VEPO Response Level 1 may include mitigation actions including migrating instances and / or workloads from one host to another, or from one instance to another (in the case of workload migration).
[0133] VEPO Response Level 2 may be associated with multiple actions, such as any / all of the mitigation actions described above in connection with VEPO Response Level 1, as well as one or more additional mitigation actions, such as an additional mitigation action corresponding to powering off hypervisors hosting instances and / or workloads associated with Category 1 customers (e.g., free tier customers, free trial customers, etc.).
[0134] VEPO Response Level 3 may be associated with multiple actions, such as any / all of the reduction actions associated with VEPO Response Levels 1 and / or 2, as well as one or more additional reduction actions. By way of example, VEPO Response Level 3 may include an additional reduction action corresponding to powering off hypervisors (e.g., after evacuating higher priority VMs) hosting Category 2 bare metal instances (e.g., Pay-As-You-Go and / or Enterprise BMs) and Category 3 virtual machines (e.g., VMs associated with the Pay-As-You-Go and Enterprise categories). In some embodiments, the additional reduction action may specify excluding from consideration BMs and / or VMs associated with Priority Level 1 customers and / or cloud services (e.g., critical customer and cloud provider services).
[0135] VEPO Response Level 4 may be associated with multiple actions, such as any / all of the mitigation actions associated with VEPO Response Levels 1, 2, and / or 3, as well as one or more additional mitigation actions. By way of example, VEPO Response Level 4 may include an additional mitigation action corresponding to powering off all compute hosts in a customer enclave except those that remain in a critical state and / or are deemed critical for recovery.
[0136] The number of response levels and / or specific reduction actions illustrated in FIG. 8 are not intended to limit the scope of the present disclosure. Any suitable number of response levels may be employed, each corresponding to any suitable number of reduction actions. The reduction actions may include or exclude hosts, instances, workloads, and / or customers based on any suitable attributes associated with them, such as based on priority, category, status, and / or future recovery needs. In some embodiments, the hosts, instances, and workloads considered as candidates for applying these reduction actions may be based on limited ranges, as described in more detail with respect to FIG. 9.
[0137] 8 include some degree of overlapping reduction actions, but this is not a requirement. In some embodiments, the reduction actions at each level may include a lesser or greater degree of overlap, including no overlap, resulting in completely unique actions associated with each level.
[0138] 9 is a diagram 900 illustrating example ranges for components affected by one or more response levels utilized by a VEPO orchestration service (e.g., VEPO orchestration service 702 of FIG. 7 ) to orchestrate power consumption constraints, migration, suspension, and / or power shutdown tasks, according to at least one embodiment. The diagram shows a user enclave 902 and a cloud service enclave 904. The user enclave 902 may include compute instances running workloads corresponding to services 906 (or any suitable workloads) operating within a customer overlay (e.g., overlay 908). The instances and / or workloads corresponding to services 906 may be associated with users 910.
[0139] In contrast, cloud service enclave 904 may include compute instances running workloads corresponding to services 912 (e.g., control and / or data plane services such as those discussed in connection with the service tenancies of Figures 15-18), and / or resources such as object storage and virtual cloud networking.
[0140] In some embodiments, a host 324 may host any suitable combination of instances and / or workloads corresponding to services 906 associated with a user (e.g., a customer) and / or services 912 and / or resources associated with a cloud provider, but only hosts, instances, and / or workloads corresponding to customer enclaves (e.g., compute resources) may be considered as candidates to which curtailment actions may be applied. In some embodiments, application of curtailment actions to services and / or resources associated with a cloud provider may be completely avoided (at least temporarily) because those services and / or resources are needed for later recovery.
[0141] Returning now to FIG. 7 , the functionality described in connection with power management service 402 may, in some embodiments, be performed by one or more virtual machines implemented in a hosted computing environment (e.g., in hosts 324 corresponding to any suitable number of servers 104 of FIG. 1 ). A hosted computing environment may include one or more rapidly provisioned and released computing resources, which may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment. A number of example cloud computing environments are provided and described in more detail below with respect to FIGS. 15-18 .
[0142] The VEPO orchestration service 702, through associated functionality or functionality provided by other systems and / or services, can be configured to manage power consumption values corresponding to hosts / instances / workloads in the physical environment. Managing power consumption values can include setting power caps, pausing workloads, instances, hosts, shutting down workloads / instances / hosts, migrating instances from one host to another, or migrating workloads from one instance and / or host to another, etc.
[0143] Impact identification manager 706 may be configured to utilize configuration data 712 (e.g., table 800 of FIG. 8 ) corresponding to the specification of multiple response levels and their corresponding reduction actions to determine the estimated impact of applying a given level of reduction action to hosts / instances / workloads (collectively, “resources”) in a physical environment. The estimated impact of applying a given action and / or level may depend on the current state of the resources at the time of execution. In some embodiments, determining the estimated impact of a given reduction action may include identifying a set of resources and / or customers to which the action, if implemented, is applicable. Thus, estimated impact is intended to refer to the scope or applicability of a given action or level. Determining the scope and / or applicability of a given action or level may include identifying the specific resources and / or customers affected by implementing the action or level and / or identifying any suitable attributes associated with the applicable resources and / or customers. As an overly simplistic example, impact identification manager 706 may identify that if an action to shut down all idle hosts is taken, X number of idle hosts among hosts 342 will be affected, that the hosts are associated with a particular customer or number of customers, and / or that the hosts and / or customers are associated with other attributes such as category (e.g., free tier) and priority (e.g., high priority). The curtailment actions may provide increasingly more severe actions and / or increasingly broader impacts. As an example of increasingly more severe actions, a lower level may specify an action to set a power cap for a particular host associated with a particular attribute, while a higher level may specify shutting down that host completely.An example of increasingly broader impact may include a lower level affecting a small number of hosts (e.g., 2, 5, 10) corresponding to a single, low-priority customer, while a higher level may affect a larger number of hosts (e.g., 100) corresponding to the same or larger number of customers associated with a broader priority level, etc. Broader impact thus refers to the number, scope, or breadth of applicability of a given action to hosts, instances, workloads, or customers.
[0144] Impact identification manager 706 may be configured to estimate a likely impact and / or estimated power reduction for one or more of the levels based, at least in part, on configuration data specifying the levels and corresponding actions and any suitable combination of power data 310, account data 312, or host / instance data, etc. For one or more response levels or one or more reduction actions corresponding to a given level, an estimated impact (e.g., what hosts, instances, workloads, customers are likely to be affected, what attributes are associated with the affected hosts, instances, workloads, and / or customers, how many hosts / instances / workloads / customers are affected, etc.) may be identified. In some embodiments, the estimated impact of a given action may be aggregated with the estimated impact for all actions at a given level to determine the estimated impact for a given level.
[0145] In some embodiments, impact identification manager 706 can be utilized to estimate impacts and / or estimate power reductions for all levels. In some embodiments, determining estimated power reductions can be an incremental process that includes identifying applicable resources (e.g., hosts, instances, and / or workloads), determining the current power consumption of the applicable resources, estimating the reduction for each response if the action is performed / implemented, and aggregating the individual estimated reductions for the individual responses to identify an estimated aggregate power reduction corresponding to the given action. This process can be repeated for each action at a given level, and the resulting estimated power consumption reductions for each action can be aggregated to determine the estimated aggregate power consumption reduction for the given level. The functionality of impact identification manager 706 can be invoked by implementation manager 708.
[0146] In some embodiments, the functionality of the implementation manager 708 can be invoked by the demand manager 710 based, at least in part, on monitoring or determining that a change to the aggregate power threshold is needed. The monitoring can include monitoring any suitable combination of the power data 310, the environmental data 314, the host / instance data 316, and / or the operational data 318 of FIG. 3 . The change to the aggregate power threshold can be implemented, and the amount by which the aggregate power threshold can be changed can be determined, at least in part, based on a number of triggering events. These triggering events can include, but are not limited to, 1) receiving a request from a governmental agency (e.g., a local power authority) requesting a power reduction indefinitely or for a period of time, 2) determining current and / or predicted environmental conditions (e.g., current ambient or external temperature, predicted ambient or external temperature), 3) determining current or predicted aggregate power consumption corresponding to devices in the physical environment, and 4) determining the current and / or predicted operational status of one or more temperature control units in the physical environment. The amount by which the aggregate power threshold can or should be changed can be determined, at least in part, based on a number of factors, including, but not limited to, the difference between ambient temperature and external temperature, the amount of power consumption attributable to or associated with a particular component (e.g., a particular temperature control unit), an amount or percentage specified in a request received from a government agency, or the difference between a current or predicted aggregate power consumption value and a current or predicted aggregate power threshold.
[0147] In some embodiments, upon determining the amount of change in the aggregate power threshold, demand manager 710 can invoke functionality of enforcement manager 708 to determine what response level is appropriate for the given change. As a non-limiting example, enforcement manager 708 can incrementally invoke functionality of impact identification manager 706 (e.g., via function call, etc.) to determine an estimated impact and / or estimated power reduction corresponding to a given level (e.g., VEPO response level 1 in FIG. 8 ). The resulting estimated impact and / or estimated power reduction calculated by the impact identification manager can be returned to enforcement manager 708. In some embodiments, enforcement manager 708 can obtain and / or implement a predefined scheme for determining the suitability of selecting one response level over another. As an example, environmental manager 708 may be configured to compare the estimated power reduction of a given level with a current aggregate power threshold identified by demand manager 710 to determine whether the estimated power consumption of the given level (e.g., VEPO response level 1) is sufficient (e.g., exceeds the amount of change required to reduce the current aggregate power consumption (identified and provided by power management service 402) to a value below the aggregate power threshold identified by demand manager 710). If sufficient, in some embodiments, enforcement manager 708 may evaluate the estimated impact according to a predefined scheme to determine whether the estimated impact is sufficient and / or acceptable. As a non-limiting example, even if the estimated power reduction is sufficient to reduce the current aggregate power consumption below the current aggregate power threshold, the response level may be rejected by enforcement manager 708 if the estimated impact exceeds the conditions of the predefined scheme (e.g., the number of affected hosts exceeds a threshold specified in the predefined scheme).If either the estimated power consumption and / or estimated impact of a response level does not pass the requirements / conditions of the predefined scheme, the implementation manager 708 may be configured to invoke the impact identification manager 706 to determine an estimated impact and / or estimated power reduction corresponding to a higher VEPO response level (e.g., VEPO response level 2).
[0148] The impact and / or power reduction estimates may be invoked by the implementation manager 708 and determined by the impact identification manager 706, one level at a time, until the implementation manager 708 finds a first level that meets the requirements / conditions of the predefined scheme with respect to the estimated impact and / or estimated power reduction. The level determined to be the lowest level (e.g., a level occurring higher in table 800) at which all impact conditions of the predefined scheme are met may be referred to as the “least impact resource level.” The level determined to be the lowest level (e.g., a level occurring higher in table 800) at which all power reduction conditions of the predefined scheme meet (e.g., below) the current aggregate power threshold may be referred to as the “least reduced resource level.” The level determined to be the lowest response level at which all conditions with respect to the estimated impact and estimated power reduction are met may be referred to as the “first sufficient response level.” The implementation manager 708 may be configured to select the response level that is first identified as the least impactful level, the least reduced level, or the first sufficient level according to a predefined scheme.
[0149] In some embodiments, enforcement manager 708 may cause the estimated impact and / or estimated power reduction for all levels (or some subset, such as the first five, next five, etc.) to be determined by impact identification manager 706. Enforcement manager 708 may be configured to select a particular response level based, at least in part, on comparing the estimated impact and / or estimated power reduction of each level to determine the one with the least impact (e.g., relatively the smallest impact), the least reduction (e.g., resulting in the smallest amount of power reduction sufficient to reduce the current aggregate power consumption below the current aggregate power threshold), or the one with the least impact and the least reduction (e.g., based at least in part on a weighting algorithm).
[0150] Any suitable data, such as estimated impacts and / or estimated reductions calculated by impact identification manager 706 (possibly by invoking functionality of computer services 306 and / or power management services 304 of FIG. 3), and / or current and / or forecast values based on those estimates, may be presented via interface 11, which will be discussed in more detail below.
[0151] As discussed herein, any suitable determination and / or identification of the demand manager 710, impact identification manager 706, and / or implementation manager 708 (or any suitable service or subroutine called by them) may rely on current data, historical data, and / or forecast data. As a non-limiting example, the demand manager 710 may determine that a change is needed and / or the amount by which the current aggregate power threshold should be changed based on the current operational status of the temperature control units 112 of FIG. 1 or the predicted operational status of those units during a future time period. This predicted status may be identified, at least in part, based on historical data (e.g., indicating one or more partial or complete failures or shutdowns of the temperature control units) and / or based on output provided by a machine learning model (e.g., a model trained in the manner described in connection with FIG. 12). In some embodiments, the amount of change determined by the demand manager 710 for the aggregate power threshold may be based on predefined formulas and / or tables. For example, formulas and / or tables may be available from which the capacity associated with the temperature control units can be determined and / or calculated. If used, the table may indicate the amount of power consumption attributable to or associated with a temperature control unit. Thus, a failure (e.g., partial and / or complete) of that temperature control unit may be associated with the entire, or some proportional, portion of the power consumption attributable to or associated with that temperature control unit (e.g., the amount of power consumption in the form of heat that the temperature control unit is configured to handle). Thus, the demand manager 710 may determine that a total failure of that temperature control unit requires a reduction in the aggregate power threshold by an amount equal to the overall power consumption attributable to or associated with the temperature control unit, as specified by the table.
[0152] In some embodiments, the implementation manager 708 may be configured to present, recommend, or automatically select a given response level from a set of possible response levels based, at least in part, on any suitable combination of estimated impact and / or estimated power reduction likely to occur if a reduction action corresponding to the given level is implemented (e.g., realized). Thus, in some embodiments, a particular level may be recommended and / or selected by the VEPO orchestration service 702 based on any suitable combination of: 1) determining a level that, if implemented, is likely to result in a sufficient, but not excessive, reduction in power consumption in resources of the physical environment (e.g., the smallest amount of power consumption reduction sufficient to reduce the current power consumption value below the current aggregate power threshold); 2) determining a level that includes the least severe set of actions; or 3) identifying the least impactful level or action (e.g., the set of affected resources and / or customers with the smallest number of potentially affected resources / customers, a priority or category indicating the lowest overall degree of importance, etc.). Determining the least severe or least impactful level or action may be determined independently of the estimated power consumption reduction likely to result from implementing the level and / or action, or determining the least severe and / or least impactful level / action may be determined from one or more levels that, if implemented given the current situation, are estimated to cause a reduction in power consumption sufficient to bring the power consumption value down to an aggregate value that falls within the current aggregate power threshold. In some embodiments, implementation manager 708 may cause any suitable data corresponding to current, predicted, or estimated power consumption and / or impact to be presented in the interface of FIG. 11. In some embodiments, selection of the particular level to implement may be made by user input provided in the user interface.Thus, in some embodiments, the implementation manager 708 may refrain from performing operations to implement a level reduction action until the user identifies and / or confirms the level to be implemented via user input provided in the user interface.
[0153] In some embodiments, performing these actions may include instructing the power management service 402 and / or the computer service 306 to perform the actions. By way of example, the enforcement manager 708 may provide identifiers for affected resources (e.g., hosts, instances, workloads, etc.) and reduction action data specifying estimated power reductions and / or actions to take for the power management service 402 (e.g., setting a power cap for a particular resource, setting a power cap for all low-priority hosts / instances / workloads, setting a power cap by a particular amount (e.g., 10%, by at least x amount, etc.) for a particular or all resources that meet a particular set of attributes), which may be configured to execute / enforce the corresponding power caps as instructed by the enforcement manager 708, as described above. Similarly, the compute service 306 may be provided with identifiers for the affected resources (e.g., hosts, instances, workloads, etc.) as well as curtailment action data specifying the action to be taken by the compute service 306 (e.g., shut down all idle hosts, shut down these specific hosts, suspend all low priority workloads, migrate and / or compress all low priority workloads to maximize the number of idle / free hosts, etc.), which may be configured to enforce / enforce the corresponding power cap as directed by the enforcement manager 708, as described above.
[0154] The demand manager 710 may be further configured to identify that a previous trigger that resulted in the response level being implemented has ceased to occur and may modify the aggregate power threshold based on identifying that the trigger condition has ceased to occur. For example, the demand manager 710 may identify (e.g., by monitoring the operational data 318) that a previously failed temperature control unit is no longer operational. As a result, the demand manager 710 may modify (e.g., increase) the aggregate power threshold by the amount that was attributed to the temperature control unit when it was fully operational. Following this modification, the implementation manager 706 may be configured to perform any suitable operation to reverse any previously taken curtailment actions, to the extent possible, depending on which response level was selected in response to the original trigger condition. If a power cap was utilized, the implementation manager 706 may instruct the power management service 402 to remove the previously applied power cap. If a workload and / or instance migration was performed, the implementation manager 706 may instruct the compute service 306 to attempt to migrate those resources back to the original host and / or instance that hosted the instance and / or workload, respectively. If resources are suspended and / or shut off, the enforcement manager 706 can instruct the compute service 306 to perform operations to resume instances and / or workloads and / or send signals to power on powered-down hosts. In this manner, the enforcement manager 706 can perform any suitable operations to recover from curtailment actions taken based on a previous selection of a given response level.
[0155] Figure 10 is a block diagram illustrating an example use case in which multiple response levels are applied based on a set of host current states, according to at least one embodiment. The example provided in Figure 10 assumes that the VEPO response levels and corresponding actions discussed in connection with Figure 8 are the current configuration data utilized by the VEPO Orchestration Service 702 of Figure 7. As shown, the following states are assumed:
[0156] Server 1004A and server 1004C are free. Server 1004D hosts resources (eg, instances) associated with customers in category level 1.
[0157] Server 1004I hosts BM instances associated with customers associated with priority level 2.
[0158] Server 1004J hosts the BM for a customer associated with priority level 3.
[0159] Server 1004K hosts VMs for a customer associated with priority level 1.
[0160] Server 1004B hosts one or more instances that are considered critical to recovery. In some embodiments, these instances may be associated with services 912 offered by a cloud provider.
[0161] As described above, the VEPO orchestration service 702 can incrementally or in parallel identify estimated impacts and / or estimated power reductions corresponding to each response level. The VEPO orchestration service 702 can identify (e.g., through functionality provided by the power management service 304 and / or the compute service 306) applicable responses to a given level (e.g., reduction actions corresponding to a given response level). By way of example, in the current example, the VEPO orchestration service 702 (via the power management service 304) can identify that VEPO Response Level 1 can be estimated to affect two servers, server 1004A and server 1004C, based on identifying servers 1004A and 1004C as a set of resources to which the redundancy actions of VEPO Response Level 1 apply (e.g., from non-cloud provider-based resources, such as resources associated with a user enclave, such as user enclave 902). In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount by which an estimated power reduction for a given level, if implemented / realized, would affect the aggregate power consumption of resources in the physical environment (e.g., a 10% reduction with no impact to customers, free BM capacity reduced to 0%, free CM capacity severely curtailed, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0162] The VEPO orchestration service 702 (via the power management service 304) can determine that VEPO Response Level 2 can be estimated to impact server 1004A and server 1004C based on identifying that server 1004D hosts resources to which VEPO Response Level 3 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with a user enclave, such as user enclave 902), as well as including VEPO Response Level 1 reduction actions. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount to which the estimated power reductions for VEPO Response Level 2, if implemented / implemented, would affect the aggregate power consumption of the resources of the physical environment (e.g., a 20% reduction with minimal impact to customers, up to 400 low-priority VMs affected, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO Response Levels.
[0163] The VEPO orchestration service 702 (via the power management service 304) can determine that VEPO Response Level 3 can be presumed to affect servers 1004A, 1004C, and 1004D based on identifying that the BM hosted by server 1004I and the VM hosted by server 1004J are resources to which VEPO Response Level 3 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with user enclaves, such as user enclave 902), as well as including VEPO Response Level 3 reduction actions. Server 1004K, and / or resources hosted by server 1004K can be excluded based on meeting the exclusions associated with the response level being evaluated. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount by which the estimated power reductions for VEPO response level 3, if implemented / realized, would affect the aggregate power consumption of resources in the physical environment (e.g., medium impact - 40% reduction, up to 200 category 1 or 2 BMs / VMs affected, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0164] In addition to identifying that the BM hosted by server 1004I and the VM hosted by server 1004J are resources to which VEPO Response Level 4 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with user enclaves, such as user enclave 902), the VEPO orchestration service 702 (via power management service 304) can determine that VEPO Response 4 can be estimated to affect all servers (e.g., the set of servers including servers 1004A and 1004C-1004L) except for server 1004B based on including the reduction actions of VEPO Response Levels 1-3. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount to which the estimated power reduction for VEPO Response Level 4, if executed / realized, would affect the aggregate power consumption of the resources of the physical environment (e.g., impact is significant with an 80% reduction, region down, etc.). Any suitable portion of this information may be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0165] FIG. 11 is a schematic diagram of an example user interface 1100, according to at least one embodiment. User interface 1100 may include any suitable number and types of graphical interface elements (e.g., drop-down boxes 1102-1108) for selecting and / or displaying VEPO data corresponding to a given region, availability domain (AD), building, and / or room, respectively. Other interface elements, such as radio buttons and / or edit boxes, may also be employed. The VEPO data presented in user interface 1100 may correspond to the region, AD, building, and room selected via graphical interface elements 1102-1108 (also referred to as "options" 1102-1108). As shown, VEPO data 1109 is displayed according to the values selected via options 1102-1108.
[0166] As described above, any suitable information utilized by the VEPO orchestration service 702, the power management service 304, and / or the compute service 306 may be presented via the user interface 1100. As a non-limiting example, any suitable combination of power data 310, account data 312, environment data 314, host / instance data 316, and / or operational data 318 may be presented in the user interface 1100.
[0167] While the specific content of the VEPO data 1109 displayed in the user interface 1100 may vary, as shown, the VEPO data 1109 includes VEPO level data (e.g., identifiers corresponding to the VEPO response levels of FIG. 8 presented in column 1110), current power consumption values (e.g., current power consumption values presented in column 1112), estimated reductions associated with each VEPO level (e.g., estimated VEPO reductions presented in column 1114), and estimated impacts associated with each VEPO level (e.g., estimated VEPO impacts presented in column 1116).
[0168] At 1118, a current aggregate power threshold (e.g., the current aggregate power threshold determined by demand engine 710 of FIG. 7) may be displayed. At 1120, a current power reduction required to meet the current aggregate power threshold (e.g., indicating a 5 kW reduction is required) may be calculated (e.g., by demand engine 710) based, at least in part, on determining the difference between the current aggregate power threshold (e.g., 11 kW) and the current power consumption (e.g., 16 kW) presented at 1112.
[0169] User interface 1100 may be configured to present estimated impact data corresponding to each of the VEPO response levels ("VEPO levels" for brevity) in column 1116. In some embodiments, areas corresponding to rows of column 1116 may be selectable to present an expanded view of estimated impact data corresponding to a particular VEPO level. Similarly, areas corresponding to rows of column 1114 may be selectable to present an expanded view of estimated VEPO reduction corresponding to a particular VEPO level.
[0170] User interface 1100 may include an indicator 1122 that indicates a recommended VEPO response level (as shown, VEPO Level 2 is recommended by the system). As discussed above, the recommended VEPO level may be identified by power orchestration service 702 based on the factors discussed above. In some embodiments, the recommended VEPO level (e.g., VEPO Level 2) may be selected based on determining that VEPO Level 2 is associated with the smallest amount of estimated VEPO reduction that meets the required reduction, as shown at 1120. In some embodiments, VEPO Level 2 may also or alternatively be recommended based, at least in part, on determining that the estimated impact of applying VEPO Level 2 provides the smallest impact relative to the estimated impact of VEPO Level 104.
[0171] In some embodiments, the area corresponding to each row of VEPO data 1109 may be selectable to select a given VEPO response level. Once selected, the user may be provided with an additional menu or option to instruct the system to proceed with the selected VEPO response according to the reduction action associated with that level.
[0172] FIG. 12 shows a flow illustrating an example method 1200 for training one or more machine learning models according to at least one embodiment. In some embodiments, the models 1202 can be trained (e.g., by the power management service 304 of FIG. 3 , the power orchestration service 302 of FIG. 3 , the compute service 306 of FIG. 3 , or a different device or system) using any suitable machine learning algorithm (e.g., supervised, unsupervised, etc.) and any suitable number of training datasets (e.g., training data 1208). A supervised machine learning algorithm refers to a machine learning task that involves learning an inferred function that maps inputs to outputs based on a labeled training dataset in which example input / output pairs are known. An unsupervised machine learning algorithm refers to a set of algorithms used to analyze and cluster an unlabeled dataset (e.g., unlabeled data 1210). These algorithms are configured to identify patterns or groupings of data without requiring human intervention. In some embodiments, any suitable number of models 1202 can be trained during the training phase 1204.
[0173] At least one of the models can be trained to identify a likelihood or confidence that the aggregate power consumption of downstream devices (e.g., hosts 324 of FIG. 3 ) exceeds a budget threshold corresponding to upstream devices (e.g., string-level devices such as string PDUs 204 of FIG. 2 ) from which power is distributed to any suitable combination of hosts 324 and / or servers 104, according to at least one embodiment. The training data 1208 for training one or more of the models 1202 may include any suitable combination of individual and / or aggregate power consumption values (current or historical) of power data 310 corresponding to one or more hosts (e.g., hosts 324 of FIG. 3 , etc., versus servers 104 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the model may be trained to identify a predicted amount corresponding to the aggregate power consumption of any suitable combination of hosts 324 and / or servers 104. In some embodiments, the training data 1208 may include current or historical power caps applied to each host, and the output provided by such model 1208 may identify a predicted power cap value for any suitable combination of hosts 324 and / or servers 104.
[0174] At least one of the models can be trained to identify predicted workloads for any suitable combination of hosts 324 and / or servers 104 according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify predicted workloads may include any suitable combination of current and / or historical host / instance data 316 corresponding to one or more hosts (e.g., host 324 of FIG. 3 for server 104 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models can be trained to identify a likelihood or confidence that a set of predicted workloads will be assigned to any suitable combination of hosts 324 and / or servers 104.
[0175] At least one of the models can be trained to identify predicted partial or complete failures of one or more temperature control units according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify such failures may include any suitable combination of current and / or historical environmental data 314 and / or operational data 318 corresponding to the physical location where the host 324 and / or server 104 are located, current and / or historical external temperatures occurring currently and / or in the past for areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or historical ambient locations associated with the physical location, weather forecasts corresponding to future time periods, and current and / or historical failures that have occurred to one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the model can be trained to identify the likelihood or confidence that a particular failure (partial or complete) of the temperature control system 112 will occur in a future period. In some embodiments, the model can be trained to identify the degree of failure for the temperature control system 112 (e.g., 80% failure, 100% failure, etc.).
[0176] At least one of the models can be trained to provide an estimated aggregate power reduction required due to partial and / or complete failure of one or more temperature control units (e.g., temperature control system 112) according to at least one embodiment. Training data 1208 for training one or more of the models 1202 to identify an estimated aggregate power reduction required may include any suitable combination of current and / or historical data of environmental data 314 and / or operational data 318 corresponding to the physical locations where the hosts 324 and / or servers 104 are located, current and / or historical external temperatures occurring currently and / or in the past for areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or historical ambient locations associated with the physical locations, weather forecasts corresponding to future time periods, current and / or historical failures that have occurred to one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ), current and / or historical aggregate power thresholds, and / or current and / or historical power consumption levels of any suitable combination of hosts 324 and / or servers 104. In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models may be trained to identify the likelihood or confidence that estimated aggregate power reductions will be needed during a future time period.
[0177] At least one of the models can be trained to predict future (e.g., ambient and / or external) temperatures according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify future temperatures may include any suitable combination of environmental data 314 and / or operational data 318 corresponding to the physical location where the host 324 and / or server 104 are located, current and / or past external temperatures occurring currently and / or in the past with respect to areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or past ambient locations associated with the physical location, weather forecasts corresponding to future time periods, current and / or past faults experienced by one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ), current and / or past aggregate power thresholds, and / or current and / or past power consumption levels of any suitable combination of the host 324 and / or server 104. In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models may be trained to identify the likelihood or confidence that a predicted temperature will occur during a future time period.
[0178] In general, models 1202 may include any suitable number of models. Models 1202 may be individually trained from training data 1208 described above to determine the likelihood or confidence or accuracy of the output provided by model 1202. In some embodiments, models 1202 may be configured to determine values corresponding to the examples provided above, where the likelihood and / or confidence values may indicate the degree of likelihood or confidence that the output provided by the model is accurate.
[0179] Generally, model 1202 can be trained during training phase 1204 using a supervised learning algorithm and labeled data 1206 to identify outputs described in the examples above. The likelihood value can be a binary indicator indicating the degree of likelihood, a percentage, a confidence value, or the like. The likelihood value can be a binary indicator indicating whether a particular budget amount is likely or unlikely to be breached, or the likelihood value can indicate the likelihood and / or confidence that a predicted output will become reality in the future. Labeled data 1206 can be any suitable portion of potential training data (e.g., training data 1208) that can be used to train the model to generate the outputs described above. Labeled data 1206 can include any suitable number of examples of current and / or historical data corresponding to any suitable number of devices in a physical environment (e.g., data center 102, server 104, temperature control system 112, PDUs of FIG. 2, etc.). In some embodiments, labeled data 1206 can include labels identifying known likelihoods and / or actual values. Using the labeled data 1206, a model (e.g., an inferred function) can be learned that maps an input (e.g., one or more instances of training data) to an output (e.g., a predicted value, or a likelihood / confidence value that the predicted value or output will occur in a future time period). In some embodiments, any suitable combination of the VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3, or separate services or systems, can train any suitable combination of the models 1202. In some embodiments, any suitable combination of the VEPO orchestration service 302, power management service 304, and / or compute service 306 can obtain a pre-trained version of the models 1202.
[0180] The models 1202, and the various types of those models described above, may include any suitable number of models trained using unsupervised learning techniques to identify likelihoods / confidences and / or predictive quantities corresponding to the examples provided above. Unsupervised machine learning algorithms are configured to learn patterns from untagged data. In some embodiments, the training phase 1204 may utilize an unsupervised machine learning algorithm to generate one or more of the models 1202. For example, the training data 1208 may include unlabeled data 1210 (e.g., any suitable combination of current and / or historical values corresponding to power data 310, account data 312, environmental data 314, host / instance data 316, operational data 318, etc.). The unlabeled data 1210 may be utilized with an unsupervised learning algorithm that segments entries in the unlabeled data 1210 into groups. The unsupervised learning algorithm may be configured to cluster similar entries into common groups. Examples of unsupervised learning algorithms may include clustering methods such as k-means clustering, DBScan, etc. In some embodiments, the unlabeled data 1210 can be clustered with the labeled data 1206 so that unlabeled instances in a given group can be assigned the same label as other labeled instances in that group.
[0181] In some embodiments, any suitable portion of the training data 1208 may be utilized to train the model 1202 during the training phase 1204. For example, 70% of the labeled data 1206 and / or unlabeled data 1210 may be utilized to train the model 1202. Once trained, or at any suitable point in time, the model 1202 may be evaluated to assess its quality (e.g., the accuracy of the output 1212 with respect to the label corresponding to the labeled data 1206). By way of example, a portion of the example labeled data 1206 and / or unlabeled data 1210 may be utilized as input to the model 1202 to generate the output 1212. By way of example, an example of the labeled data 1206 may be provided as input, and the corresponding output (e.g., output 1212) may be compared with a label already known to be associated with that example. If a portion of the output (e.g., a label) matches the example label, that portion of the output may be deemed accurate. Any suitable number of labeled examples can be utilized, and the number of correct labels can be compared to the total number of examples provided (and / or the total number of previously identified labels) to determine an accuracy value for a given model that quantifies the model's accuracy. For example, if 90 out of 100 input examples produce output labels that match previously known label examples, the model being evaluated can be determined to be 90% accurate.
[0182] In some embodiments, as the model 1202 is utilized with subsequent inputs, subsequent outputs generated by the model 1202 can be added to the corresponding inputs and used to retrain and / or update the model 1202 at 1216. In some embodiments, an example may not be used to retrain or update the model until a feedback procedure 1214 is performed. In the feedback procedure 1214, an example (e.g., an example including one or more historical consumption data instances corresponding to one or more devices and / or racks) and the corresponding output generated for that example by one of the models 1202 are presented to a user, who identifies whether the generated output (e.g., quantity and / or likelihood confidence value) is correct for a given example.
[0183] The training process shown in FIG. 12 (e.g., method 1200) can be performed any suitable number of times, at any suitable intervals and / or according to any suitable schedule, so that the accuracy of model 1202 improves over time.
[0184] In some embodiments, any suitable number and / or combination of models 1202 may be used to determine the output. In some embodiments, the power management service 402 may utilize any suitable combination of outputs provided by the models 1202 to determine whether a budget threshold for a given component (e.g., a column-level PDU) is likely to be breached (and / or likely to be breached by a certain amount). Thus, in some embodiments, models trained with any suitable combination of supervised and unsupervised learning algorithms may be utilized by the power management service 402.
[0185] FIG. 13 is a block diagram illustrating an example method for causing application of a corresponding set of curtailment actions associated with a response level, according to at least one embodiment. Method 1300 may be performed by one or more components of power orchestration system 300 of FIG. 3 or subcomponents thereof discussed in connection with FIGS. 4 and 7. By way of example, method 1300 may be performed, at least in part, by any suitable combination of VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3. The operations of method 1300 may be performed in any suitable order. Method 1300 may include more or fewer operations than those illustrated in FIG. 13.
[0186] Method 1300 may begin at 1302, where configuration data (e.g., configuration data 712 of FIG. 7) may be obtained (e.g., by power orchestration service 702 of FIG. 2). In some embodiments, the configuration data may include multiple response levels (e.g., VEPO response levels of FIG. 8), which specify the applicability of respective sets of curtailment actions to multiple hosts (e.g., via one or more curtailment actions corresponding to each response level). In some embodiments, a first response level of the multiple response levels specifies the applicability of a first set of curtailment actions to multiple hosts (e.g., based at least in part on one or more curtailment actions associated with that response level). By way of example, the VEPO response level of FIG. 8 may specify that all free / idle hosts and hypervisors that can be fully evacuated are applicable to the action (e.g., powering off). In this manner, the curtailment action indicates that a subset of resources (e.g., free / idle hosts and hypervisors that can be fully evacuated) are applicable to a given curtailment action (e.g., powering off) and, through association, a corresponding response level.
[0187] At 1304, a current value of aggregate power consumption for the multiple hosts may be obtained (e.g., by the power management service 402, which may provide such data to the power orchestration service 702). The current value of aggregate power consumption for the multiple hosts may be calculated, for example, by the power management service 402 based at least in part on the power data 310 of FIG. 3 and / or from power consumption values obtained by / from the rack PDUs 320 and / or from any suitable combination of hosts 324 (e.g., the hosts in the rack 319, all hosts, etc.).
[0188] At 1306, a current value of an aggregate power threshold for the multiple hosts may be obtained (e.g., by demand manager 710 of FIG. 7 ). In some embodiments, the aggregate power threshold may be an aggregation of all known power consumption-related capabilities of any suitable devices in the physical environment in which the multiple hosts are located. As a non-limiting example, the aggregate power threshold may be an aggregation of a budget power level and / or the amount of heat / power consumption that a device (e.g., a temperature control unit, such as temperature control unit 112A of FIG. 1 , is configured to handle / process when fully operational). In some embodiments, the current aggregate power threshold may be calculated by demand manager 710 in the manner described above in connection with the above figures and based on the factors and / or data described above. In some embodiments, the current aggregate power threshold may be dynamically changed (e.g., by demand manager 710) based on changing conditions of devices in the physical environment in which the multiple hosts are located (including, e.g., changing conditions of server 104, changing conditions of temperature control system 112, changing power consumption, changing temperatures, changing operational status of devices in the physical environment, etc.). In some embodiments, the current aggregate power threshold may be modified / changed based, at least in part, on receiving a request from a governmental entity (e.g., a local power authority) requesting / requiring a curtailment, where the nature of the requested curtailment may be specified in the request.
[0189] At 1308, a first response level may be selected (e.g., by implementation manager 708 of FIG. 7) from a plurality of response levels (e.g., VEPO response levels of FIG. 8) based at least on a difference between the current value of the aggregate power consumption and the current value of the aggregate power threshold. In some embodiments, the selection may be initiated, at least in part, based on receiving user input (e.g., at user interface 1100) indicating a selection of the first response level.
[0190] At 1310, the computer system may cause application of a first set of curtailment actions to at least one host of the plurality of hosts according to the selected first response level. In some embodiments, causing application of the first set of curtailment actions may include performing and / or directing the performance of any suitable operations described above in connection with the power orchestration service 302, the power management service 304, and / or the compute service 306. In some embodiments, the instructions, if used, may originate from any of the power orchestration service 302, the power management service 304, and / or the compute service 306, or from one of the other services instructing that service to perform one or more operations such as those described above in connection with FIGS. 3, 4, and 7.
[0191] FIG. 14 is a block diagram illustrating an example method for preemptively migrating workloads affected by a selected response level, according to at least one embodiment. Method 1400 may be performed by one or more components of power orchestration system 300 of FIG. 3 or subcomponents thereof discussed in connection with FIGS. 4 and 7. By way of example, method 1400 may be performed, at least in part, by any suitable combination of VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3. The operations of method 1400 may be performed in any suitable order. Method 1400 may include more or fewer operations than those illustrated in FIG. 14.
[0192] Method 1402 may begin at 1402, where configuration data (e.g., configuration data 712 of FIG. 7) may be obtained (e.g., by power orchestration service 702 of FIG. 2). In some embodiments, the configuration data may include multiple response levels (e.g., VEPO response levels of FIG. 8), which specify the applicability of respective sets of curtailment actions to multiple hosts (e.g., via one or more curtailment actions corresponding to each response level). In some embodiments, a first response level of the multiple response levels specifies the applicability of a first set of curtailment actions to multiple hosts (e.g., based at least in part on one or more curtailment actions associated with that response level). By way of example, the VEPO response level of FIG. 8 may specify that all free / idle hosts and hypervisors that can be fully evacuated are applicable to the action (e.g., powering off). In this manner, the curtailment action indicates that a subset of resources (e.g., free / idle hosts and hypervisors that can be fully evacuated) are applicable to a given curtailment action (e.g., powering off) and, through association, to the corresponding response level.
[0193] At 1404, a forecast of aggregate power consumption of the multiple hosts for a future time period may be obtained (e.g., by the power management service 402, which may provide such data to the power orchestration service 702). The forecast of aggregate power consumption of the multiple hosts may be calculated, for example, by the power management service 402 based at least in part on the power data 310 of FIG. 3 and / or from power consumption values obtained by / from the rack PDUs 320 and / or from any suitable combination of hosts 324 (e.g., hosts in rack 319, all hosts, etc.). In some embodiments, the forecast of aggregate power consumption of the multiple hosts may utilize historical power data. In further environments, the forecast of aggregate power consumption may utilize a pre-trained machine learning model (e.g., model 1202 of FIG. 12). The machine learning model may be pre-trained using a supervised or unsupervised machine learning algorithm and a training dataset (e.g., training data 1208), as described in connection with FIG. 12, to provide an output (e.g., a predicted value of aggregate power consumption) based, at least in part, on providing input (e.g., any suitable combination of power data 310, account data 312, environmental data 314, host / interface data 316, and / or operational data 318 of FIG. 3).
[0194] At 1406, predicted aggregate power thresholds for the multiple hosts may be obtained (e.g., by demand manager 710 of FIG. 7 ). In some embodiments, the predicted aggregate power threshold may be an aggregation of all known and / or predicted power consumption-related capabilities of any suitable devices in the physical environment in which the multiple hosts are located. As a non-limiting example, the predicted aggregate power threshold may be an aggregation of budget and / or predicted power levels and / or amounts or predicted amounts of heat / power consumption that a device (e.g., a temperature control unit, such as temperature control unit 112A of FIG. 1 , is configured to handle / handle if fully operational). In some embodiments, the predicted aggregate power threshold may be calculated by demand manager 710 in the manner described above in connection with the preceding figures and based on the factors and / or data described above. In some embodiments, the predicted aggregate power threshold may be determined, at least in part, based on any suitable combination of current and historical data for power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In further environments, the predicted value of aggregate power consumption may utilize a pre-trained machine learning model (e.g., model 1202 of FIG. 12 ). The machine learning model may be pre-trained using a supervised or unsupervised machine learning algorithm and a training data set (e.g., training data 1208) as described in connection with FIG. 12 to provide an output (e.g., a predicted value of aggregate power threshold) based, at least in part, on receiving input data (e.g., any suitable combination of power data 310, account data 312, environmental data 314, host / interface data 316, and / or operational data 318 of FIG. 3 ).
[0195] At 1408, a first response level may be selected (e.g., by implementation manager 708 of FIG. 7) from a plurality of response levels (e.g., VEPO response levels of FIG. 8) based at least on a difference between the current value of the aggregate power consumption and the predicted value of the aggregate power threshold. In some embodiments, the selection may be initiated at least in part based on receiving user input (e.g., at user interface 1100) indicating a selection of the first response level.
[0196] At 1410, the computer system may identify one or more workloads that (a) are currently executing on the plurality of hosts and (b) would be affected by application of a first set of reduction actions to the plurality of hosts according to a selected first response level. Identification of the one or more workloads may be performed in the manner described above in connection with impact identification manager 706 in FIG. 7 .
[0197] At 1412, the computer system may preemptively migrate affected workloads from multiple hosts to one or more other hosts in advance of a future time period. In some embodiments, these actions may be performed by the compute service 306 (e.g., based at least in part on instructions received by that service from the power orchestration service 302).
[0198] 15-18 illustrate a number of example environments that may be hosted by host 324 of FIG. 3. The environments illustrated in FIGS. 15-18 illustrate cloud computing, multi-tenant environments. As discussed above, cloud computing and / or multi-tenant environments, as well as other environments, may benefit from utilizing the power management techniques disclosed herein. These techniques enable data centers, including components implementing the environments described in connection with FIGS. 1-4 and 7, among others, to utilize the data center's power resources more efficiently than conventional techniques, which left large amounts of power unused.
[0199] IaaS Infrastructure Example Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, an IaaS provider can also offer various services associated with those infrastructure components (example services include billing software, monitoring software, logging software, load balancing software, clustering software, etc.). Accordingly, these services can be policy-driven, allowing IaaS users to implement policies that drive load balancing to maintain application availability and performance.
[0200] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on the VMs. The customer can then use the provider's services to perform a variety of functions, including distributing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0201] In most cases, the cloud computing model requires the participation of a cloud provider, which can be, but is not required to be, a third-party service that specializes in providing (e.g., offering, renting, selling) IaaS. An entity can also choose to deploy a private cloud and become its own provider of infrastructure services.
[0202] In some examples, IaaS deployment is the process of placing a new application, or a new version of an application, onto a prepared application server, etc. IaaS deployment may also include the process of preparing the server (e.g., installing libraries, daemons, etc.), which is often managed by the cloud provider below the hypervisor layer (e.g., server, storage, network hardware, and virtualization). Thus, the customer may be responsible for handling the deployment of the (OS), middleware, and / or application (e.g., in self-service virtual machines (e.g., that can be spun up on demand) etc.
[0203] In some instances, IaaS provisioning may refer to obtaining the computers or virtual hosts to be used and also installing the necessary libraries or services on those computers or virtual hosts. In most cases, deployment does not include provisioning, which must be performed first.
[0204] In some cases, IaaS provisioning presents two distinct challenges. First, there is the initial challenge of provisioning the initial set of infrastructure before anything is running. Second, there is the challenge of evolving the existing infrastructure (e.g., adding new services, modifying services, removing services, etc.) once everything is provisioned. In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., what resources depend on what and how each works together) can be described declaratively. In some cases, once the topology is defined, workflows can be generated to create and / or manage the different components described in the configuration files.
[0205] In some examples, the infrastructure can have many interconnected elements. For example, there can be one or more virtual private clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there can also be one or more inbound / outbound traffic group rules and one or more virtual machines (VMs) provisioned to define how inbound and / or outbound traffic on the network is configured. Other infrastructure elements, such as load balancers, databases, etc., can also be provisioned. The infrastructure can evolve over time as more infrastructure elements are desired and / or added.
[0206] In some cases, continuous deployment techniques can be employed to enable deployment of infrastructure code across various virtual computing environments. Additionally, the described techniques can enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, but often many, different production environments (e.g., across various different geographic locations, sometimes across the globe). However, in some examples, the infrastructure onto which the code will be deployed must first be set up. In some cases, provisioning can be done manually, utilizing provisioning tools to provision resources, and / or utilizing deployment tools to deploy the code once the infrastructure has been provisioned.
[0207] 15 is a block diagram 1500 illustrating an example IaaS architecture pattern, according to at least one embodiment. A service operator 1502 may be communicatively coupled to a secure host tenancy 1504, which may include a virtual cloud network (VCN) 1506 and a secure host subnet 1508. In some examples, the service operator 1502 may employ one or more client computing devices, which may be portable handheld devices (e.g., iPhones, mobile phones, iPads, computing tablets, personal digital assistants (PDAs)), or wearable devices (e.g., Google Glass head-mounted displays) running software such as Microsoft Windows Mobile and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and supporting the Internet, email, short message service (SMS), Blackberry, or other communication protocols. Alternatively, the client computing devices may be general-purpose personal computers, including, by way of example, personal and / or laptop computers running various versions of the Microsoft Windows operating system, the Apple Macintosh operating system, and / or the Linux operating system. The client computing devices may be workstation computers running any of a variety of commercially available UNIX or UNIX-like operating systems, including, without limitation, various GNU / Linux operating systems, such as Google Chrome OS.Alternatively or in addition, the client computing device may be any other electronic device, such as a thin client computer, an Internet-enabled gaming system (e.g., a Microsoft Xbox gaming console with or without a Kinect® gesture input device), and / or a personal messaging device, that can communicate over a network that has access to VCN 1506 and / or the Internet.
[0208] VCN 1506 may include a local peering gateway (LPG) 1510, which may be communicatively coupled to a secure shell (SSH) VCN 1512 via the LPG 1510 included in the SSH VCN 1512. The SSH VCN 1512 may include an SSH subnet 1514, which may be communicatively coupled to a control plane VCN 1516 via the LPG 1510 included in the control plane VCN 1516. The SSH VCN 1512 may also be communicatively coupled to a data plane VCN 1518 via the LPG 1510. The control plane VCN 1516 and the data plane VCN 1518 may be included in a service tenancy 1519, which may be owned and / or operated by the IaaS provider.
[0209] The control plane VCN 1516 may include a control plane demilitarized zone (DMZ) tier 1520 that acts as a perimeter network (e.g., a portion of an enterprise network between the enterprise intranet and an external network). DMZ-based servers have limited responsibility and can help mitigate contained breaches. Additionally, the DMZ tier 1520 may include one or more load balancer (LB) subnets 1522, a control plane app tier 1524 that may include an app subnet 1526, and a control plane data tier 1528 that may include a database (DB) subnet 1530 (e.g., a front-end DB subnet and / or a back-end DB subnet). LB subnet 1522 included in control plane DMZ tier 1520 can be communicatively coupled to app subnet 1526 included in control plane app tier 1524 and to an Internet gateway 1534 that can be included in control plane VCN 1516, and app subnet 1526 can be communicatively coupled to DB subnet 1530 included in control plane data tier 1528, to a service gateway 1536, and to a network address translation (NAT) gateway 1538. Control plane VCN 1516 can include service gateway 1536 and NAT gateway 1538.
[0210] The control plane VCN 1516 may include a data plane mirrored app tier 1540, which may include an app subnet 1526. The app subnet 1526 included in the data plane mirrored app tier 1540 may include a virtual network interface controller (VNIC) 1542 on which a compute instance 1544 can run. The compute instance 1544 may communicatively couple the app subnet 1526 of the data plane mirrored app tier 1540 to the app subnet 1526, which may be included in the data plane app tier 1546.
[0211] The data plane VCN 1518 may include a data plane app layer 1546, a data plane DMZ layer 1548, and a data plane data layer 1550. The data plane DMZ layer 1548 may include a LB subnet 1522 that can be communicatively coupled to an app subnet 1526 of the data plane app layer 1546 and an internet gateway 1534 of the data plane VCN 1518. The app subnet 1526 can be communicatively coupled to a service gateway 1536 of the data plane VCN 1518 and a NAT gateway 1538 of the data plane VCN 1518. The data plane data layer 1550 may also include a DB subnet 1530 that can be communicatively coupled to the app subnet 1526 of the data plane app layer 1546.
[0212] The internet gateway 1534 of the control plane VCN 1516 and the internet gateway 1534 of the data plane VCN 1518 can be communicatively coupled to a metadata management service 1552, which can be communicatively coupled to the public internet 1554. The public internet 1554 can be communicatively coupled to a NAT gateway 1538 of the control plane VCN 1516 and the NAT gateway 1538 of the data plane VCN 1518. The service gateway 1536 of the control plane VCN 1516 and the service gateway 1536 of the data plane VCN 1518 can be communicatively coupled to a cloud service 1556.
[0213] In some examples, the service gateway 1536 of the control plane VCN 1516 or the service gateway 1536 of the data plane VCN 1518 can make application programming interface (API) calls to the cloud service 1556 without traversing the public internet 1554. The API calls from the service gateway 1536 to the cloud service 1556 can be one-way. That is, the service gateway 1536 can make an API call to the cloud service 1556, and the cloud service 1556 can send the requested data to the service gateway 1536. However, the cloud service 1556 cannot initiate the API call to the service gateway 1536.
[0214] In some examples, secure host tenancy 1504 can be directly connected to service tenancy 1519, which may otherwise be isolated. Secure host subnet 1508 can communicate with SSH subnet 1514 through LPG 1510, which can enable bidirectional communication in an otherwise isolated system. By connecting secure host subnet 1508 to SSH subnet 1514, secure host subnet 1508 can access other entities in service tenancy 1519.
[0215] The control plane VCN 1516 can enable users of the service tenancy 1519 to set up or otherwise provision desired resources. The desired resources provisioned in the control plane VCN 1516 can be deployed or otherwise used in the data plane VCN 1518. In some examples, the control plane VCN 1516 can be isolated from the data plane VCN 1518, and the data plane mirror app layer 1540 of the control plane VCN 1516 can communicate with the data plane app layer 1546 of the data plane VCN 1518 via a VNIC 1542, which can be included in the data plane mirror app layer 1540 and the data plane app layer 1546.
[0216] In some examples, a user or customer of the system can make a request, for example, a request for a create, read, update, or delete (CRUD) operation, via the public internet 1554, which can communicate such a request to the metadata management service 1552. The metadata management service 1552 can communicate the request to the control plane VCN 1516 via the internet gateway 1534. The request can be received by the LB subnet 1522 included in the control plane DMZ tier 1520. The LB subnet 1522 may determine that the request is valid, and in response to this determination, the LB subnet 1522 can send the request to the app subnet 1526 included in the control plane app tier 1524. If the request is validated and requires a call to the public internet 1554, the call to the public internet 1554 can be sent to the NAT gateway 1538, which can make the call to the public internet 1554. Metadata that may be desired to be stored by the request can be stored in the DB subnet 1530.
[0217] In some examples, the data plane mirror app layer 1540 can facilitate direct communication between the control plane VCN 1516 and the data plane VCN 1518. For example, it may be desired that configuration changes, updates, or other suitable modifications be applied to resources included in the data plane VCN 1518. The control plane VCN 1516 can communicate directly with the resources included in the data plane VCN 1518 via the VNIC 1542, thereby performing configuration changes, updates, or other suitable modifications on the resources.
[0218] In some embodiments, the control plane VCN 1516 and the data plane VCN 1518 can be included in the service tenancy 1519. In this case, a user or customer of the system need not own or operate either the control plane VCN 1516 or the data plane VCN 1518. Instead, an IaaS provider can own or operate the control plane VCN 1516 and the data plane VCN 1518, both of which can be included in the service tenancy 1519. This embodiment can enable network isolation, thereby preventing users or customers from interacting with other users' or customers' resources. This embodiment can also enable users or customers of the system to store databases privately without having to rely on the public internet 1554, which may not have the desired level of threat protection for storage.
[0219] In another embodiment, LB subnet 1522 included in control plane VCN 1516 may be configured to receive signals from service gateway 1536. In this embodiment, control plane VCN 1516 and data plane VCN 1518 may be configured to be called by customers of the IaaS provider without calling the public internet 1554. Customers of the IaaS provider may desire this embodiment because databases used by the customers can be stored in service tenancy 1519, which can be controlled by the IaaS provider and isolated from the public internet 1554.
[0220] Figure 16 is a block diagram 1600 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1602 (e.g., service operator 1502 of Figure 15) can be communicatively coupled to a secure host tenancy 1604 (e.g., secure host tenancy 1504 of Figure 15), which can include a virtual cloud network (VCN) 1606 (e.g., VCN 1506 of Figure 15) and a secure host subnet 1608 (e.g., secure host subnet 1508 of Figure 15). VCN 1606 can include a local peering gateway (LPG) 1610 (e.g., LPG 1510 of Figure 15), which can be communicatively coupled to a secure shell (SSH) VCN 1612 (e.g., SSH VCN 1512 of Figure 15) via the LPG 1510 included in the SSH VCN 1612. SSH VCN 1612 can include SSH subnet 1614 (e.g., SSH subnet 1514 in FIG. 15 ), and SSH VCN 1612 can be communicatively coupled to control plane VCN 1616 (e.g., control plane VCN 1516 in FIG. 15 ) via LPG 1610 included in control plane VCN 1616. Control plane VCN 1616 can be included in service tenancy 1619 (e.g., service tenancy 1519 in FIG. 15 ), and data plane VCN 1618 (e.g., data plane VCN 1518 in FIG. 15 ) can be included in customer tenancy 1621, which may be owned or operated by a user or customer of the system.
[0221] The control plane VCN 1616 may include a control plane DMZ tier 1620 (e.g., control plane DMZ tier 1520 in FIG. 15 ) that may include a LB subnet 1622 (e.g., LB subnet 1522 in FIG. 15 ), a control plane app tier 1624 (e.g., control plane app tier 1524 in FIG. 15 ) that may include an app subnet 1626 (e.g., app subnet 1526 in FIG. 15 ), and a control plane data tier 1628 (e.g., control plane data tier 1528 in FIG. 15 ) that may include a database (DB) subnet 1630 (e.g., similar to DB subnet 1530 in FIG. 15 ). The LB subnet 1622 included in the control plane DMZ tier 1620 can be communicatively coupled to an app subnet 1626 included in the control plane app tier 1624 and to an Internet gateway 1634 (e.g., Internet gateway 1534 in FIG. 15 ) that can be included in the control plane VCN 1616, and the app subnet 1626 can be communicatively coupled to a DB subnet 1630 included in the control plane data tier 1628 and to a service gateway 1636 (e.g., service gateway 1536 in FIG. 15 ) and a network address translation (NAT) gateway 1638 (e.g., NAT gateway 1538 in FIG. 15 ). The control plane VCN 1616 can include the service gateway 1636 and the NAT gateway 1638.
[0222] The control plane VCN 1616 may include a data plane mirror app layer 1640 (e.g., data plane mirror app layer 1540 of FIG. 15 ), which may include an app subnet 1626. The app subnet 1626 included in the data plane mirror app layer 1640 may include a virtual network interface controller (VNIC) 1642 (e.g., VNIC 1542) on which a compute instance 1644 (e.g., similar to compute instance 1544 of FIG. 15 ) can run. The compute instance 1644 can facilitate communication between the app subnet 1626 of the data plane mirror app layer 1640 and the app subnet 1626, which may be included in the data plane app layer 1646 (e.g., data plane app layer 1546 of FIG. 15 ), via the VNIC 1642 included in the data plane mirror app layer 1640 and the VNIC 1642 included in the data plane app layer 1646.
[0223] An internet gateway 1634 included in the control plane VCN 1616 can be communicatively coupled to a metadata management service 1652 (e.g., metadata management service 1552 of FIG. 15 ), which can be communicatively coupled to a public internet 1654 (e.g., public internet 1554 of FIG. 15 ). The public internet 1654 can be communicatively coupled to a NAT gateway 1638 included in the control plane VCN 1616. A service gateway 1636 included in the control plane VCN 1616 can be communicatively coupled to a cloud service 1656 (e.g., cloud service 1556 of FIG. 15 ).
[0224] In some examples, data plane VCN 1618 can be included in customer tenancy 1621. In this case, the IaaS provider can provide a control plane VCN 1616 for each customer, and the IaaS provider can set up a unique compute instance 1644 for each customer, which is included in service tenancy 1619. Each compute instance 1644 can enable communication between the control plane VCN 1616, which is included in service tenancy 1619, and the data plane VCN 1618, which is included in customer tenancy 1621. The compute instance 1644 can enable resources provisioned in the control plane VCN 1616, which is included in service tenancy 1619, to be deployed to or otherwise used in the data plane VCN 1618, which is included in customer tenancy 1621.
[0225] In another example, an IaaS provider customer may have a database that resides in customer tenancy 1621. In this example, control plane VCN 1616 may include data plane mirror app tier 1640, which may include app subnet 1626. Data plane mirror app tier 1640 may reside in data plane VCN 1618, and data plane mirror app tier 1640 may not reside in data plane VCN 1618. That is, data plane mirror app tier 1640 has access to customer tenancy 1621, but data plane mirror app tier 1640 may not reside in data plane VCN 1618 or be owned or operated by the IaaS provider customer. Data plane mirror app tier 1640 may be configured to make calls to data plane VCN 1618, but may not be configured to make calls to any entities included in control plane VCN 1616. A customer may wish to deploy or otherwise use resources in the data plane VCN 1618 that have been provisioned in the control plane VCN 1616, and the data plane mirror app layer 1640 can facilitate the desired deployment or other use of the customer's resources.
[0226] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 1618. In this embodiment, the customer can determine what the data plane VCN 1618 can access, and the customer can restrict access from the data plane VCN 1618 to the public internet 1654. The IaaS provider may not be able to apply filters or otherwise control access from the data plane VCN 1618 to any external networks or databases. Applying filters and controls to the data plane VCN 1618 that the customer includes in the customer tenancy 1621 can help isolate the data plane VCN 1618 from other customers and from the public internet 1654.
[0227] In some embodiments, service gateway 1636 can call cloud services 1656 to access services that may not reside on the public internet 1654, on the control plane VCN 1616, or on the data plane VCN 1618. The connection between cloud services 1656 and the control plane VCN 1616 or data plane VCN 1618 may not be live or continuous. Cloud services 1656 may reside on different networks owned or operated by the IaaS provider. Cloud services 1656 may be configured to receive calls from service gateway 1636 and may not be configured to receive calls from the public internet 1654. Some cloud services 1656 may be isolated from other cloud services 1656, and control plane VCN 1616 may be isolated from cloud services 1656 that may not be in the same region as control plane VCN 1616. For example, control plane VCN 1616 may be located in “Region 1,” and cloud service “Deployment 15” may be located in Region 1 and Region 2. If a call to deployment 15 is made by service gateway 1636 included in control plane VCN 1616 located in region 1, the call may be sent to deployment 15 in region 1. In this example, control plane VCN 1616 or deployment 15 in region 1 may not be communicatively coupled or otherwise in communication with deployment 15 in region 2.
[0228] Figure 17 is a block diagram 1700 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1702 (e.g., service operator 1502 of Figure 15) can be communicatively coupled to a secure host tenancy 1704 (e.g., secure host tenancy 1504 of Figure 15), which can include a virtual cloud network (VCN) 1706 (e.g., VCN 1506 of Figure 15) and a secure host subnet 1708 (e.g., secure host subnet 1508 of Figure 15). VCN 1706 can include an LPG 1710 (e.g., LPG 1510 of Figure 15), which can be communicatively coupled to an SSH VCN 1712 (e.g., SSH VCN 1512 of Figure 15) via an LPG 1710 included in SSH VCN 1712. SSH VCN 1712 can include SSH subnet 1714 (e.g., SSH subnet 1514 in FIG. 15 ), and SSH VCN 1712 can be communicatively coupled to control plane VCN 1716 (e.g., control plane VCN 1516 in FIG. 15 ) via LPG 1710 included in control plane VCN 1716 and to data plane VCN 1718 (e.g., data plane 1518 in FIG. 15 ) via LPG 1710 included in data plane VCN 1718. Control plane VCN 1716 and data plane VCN 1718 can be included in service tenancy 1719 (e.g., service tenancy 1519 in FIG. 15 ).
[0229] The control plane VCN 1716 may include a control plane DMZ tier 1720 (e.g., control plane DMZ tier 1520 of FIG. 15 ) that may include a load balancer (LB) subnet 1722 (e.g., LB subnet 1522 of FIG. 15 ), a control plane app tier 1724 (e.g., control plane app tier 1524 of FIG. 15 ) that may include an app subnet 1726 (e.g., similar to app subnet 1526 of FIG. 15 ), and a control plane data tier 1728 (e.g., control plane data tier 1528 of FIG. 15 ) that may include a DB subnet 1730. LB subnet 1722 included in control plane DMZ tier 1720 can be communicatively coupled to app subnet 1726 included in control plane app tier 1724 and to an Internet gateway 1734 (e.g., Internet gateway 1534 in FIG. 15 ), which may be included in control plane VCN 1716, and app subnet 1726 can be communicatively coupled to DB subnet 1730 included in control plane data tier 1728 and to a service gateway 1736 (e.g., service gateway in FIG. 15 ) and a network address translation (NAT) gateway 1738 (e.g., NAT gateway 1538 in FIG. 15 ). Control plane VCN 1716 may include service gateway 1736 and NAT gateway 1738.
[0230] Data plane VCN 1718 may include a data plane app layer 1746 (e.g., data plane app layer 1546 in FIG. 15 ), a data plane DMZ layer 1748 (e.g., data plane DMZ layer 1548 in FIG. 15 ), and a data plane data layer 1750 (e.g., data plane data layer 1550 in FIG. 15 ). Data plane DMZ layer 1748 may include LB subnet 1722, which may be communicatively coupled to trusted app subnet 1760 and untrusted app subnet 1762 of data plane app layer 1746, which are included in data plane VCN 1718, as well as to Internet gateway 1734. Trusted app subnet 1760 may be communicatively coupled to service gateway 1736, which is included in data plane VCN 1718, NAT gateway 1738, which is included in data plane VCN 1718, and DB subnet 1730, which is included in data plane data layer 1750. The untrusted app subnet 1762 may be communicatively coupled to a service gateway 1736 included in the data plane VCN 1718 and to a DB subnet 1730 included in the data plane data layer 1750. The data plane data layer 1750 may include a DB subnet 1730 that may be communicatively coupled to a service gateway 1736 included in the data plane VCN 1718.
[0231] The untrusted app subnet 1762 may include one or more primary VNICs 1764(1)-(N), which may be communicatively coupled to tenant virtual machines (VMs) 1766(1)-(N). Each tenant VM 1766(1)-(N) may be communicatively coupled to a respective app subnet 1767(1)-(N), which may be included in a respective container egress VCN 1768(1)-(N), which may be included in a respective customer tenancy 1770(1)-(N). Each secondary VNIC 1772(1)-(N) may facilitate communication between the untrusted app subnet 1762 included in the data plane VCN 1718 and the app subnet included in the container egress VCN 1768(1)-(N). Each container egress VCN 1768(1)-(N) may include a NAT gateway 1738, which may be communicatively coupled to the public internet 1754 (e.g., public internet 1554 in FIG. 15 ).
[0232] The internet gateway 1734 included in the control plane VCN 1716 and the internet gateway 1734 included in the data plane VCN 1718 can be communicatively coupled to a metadata management service 1752 (e.g., metadata management system 1552 of FIG. 15 ), which can be communicatively coupled to the public internet 1754. The public internet 1754 can be communicatively coupled to a NAT gateway 1738 included in the control plane VCN 1716 and the NAT gateway 1738 included in the data plane VCN 1718. The service gateway 1736 included in the control plane VCN 1716 and the service gateway 1736 included in the data plane VCN 1718 can be communicatively coupled to cloud services 1756.
[0233] In some embodiments, data plane VCN 1718 can be integrated with customer tenancy 1770. This integration can be useful or desirable for an IaaS provider's customer in some cases, such as when they may want support when running code. A customer may provide code to run that may be disruptive, may communicate with other customer resources, or may otherwise cause undesirable effects. In response, the IaaS provider can determine whether to run the code that the customer has provided to the IaaS provider.
[0234] In some examples, a customer of an IaaS provider can grant temporary network access to the IaaS provider and request a function to be attached to data plane app layer 1746. The code that performs the function can run in VMs 1766(1)-(N), and the code may not be configured to run anywhere else on data plane VCN 1718. Each VM 1766(1)-(N) can be connected to one customer tenancy 1770. Each container 1771(1)-(N) contained in VM 1766(1)-(N) can be configured to run code. In this case, there can be double isolation (e.g., container 1771(1)-(N) can run code, and container 1771(1)-(N) can be contained in VMs 1766(1)-(N) that are at least in untrusted app subnet 1762). This can help prevent erroneous or otherwise unwanted code from damaging the IaaS provider's network or damaging a different customer's network. Containers 1771(1)-(N) can be communicatively coupled to customer tenancy 1770 and can be configured to send or receive data from customer tenancy 1770. Containers 1771(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 1718. When the code execution is complete, the IaaS provider can kill or otherwise discard containers 1771(I)-(N).
[0235] In some embodiments, trusted app subnet 1760 can execute code that may be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 1760 can be communicatively coupled to DB subnet 1730 and configured to perform CRUD operations on DB subnet 1730. Untrusted app subnet 1762 can be communicatively coupled to DB subnet 1730, but in this embodiment, the untrusted app subnet can be configured to perform read operations within DB subnet 1730. Containers 1771(1)-(N) that can be included in each customer's VMs 1766(1)-(N) and that can execute code from the customer may not be communicatively coupled to DB subnet 1730.
[0236] In other embodiments, the control plane VCN 1716 and the data plane VCN 1718 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 1716 and the data plane VCN 1718. However, communication may occur indirectly in at least one manner. An IaaS provider may establish an LPG 1710 that can facilitate communication between the control plane VCN 1716 and the data plane VCN 1718. In another example, the control plane VCN 1716 or the data plane VCN 1718 can make a call to a cloud service 1756 through the service gateway 1736. For example, a call from the control plane VCN 1716 to the cloud service 1756 may include a request for a service that can communicate with the data plane VCN 1718.
[0237] 18 is a block diagram 1800 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1802 (e.g., service operator 1502 of FIG. 15 ) can be communicatively coupled to a secure host tenancy 1804 (e.g., secure host tenancy 1504 of FIG. 15 ), which can include a virtual cloud network (VCN) 1806 (e.g., VCN 1506 of FIG. 15 ) and a secure host subnet 1808 (e.g., secure host subnet 1508 of FIG. 15 ). VCN 1806 can include an LPG 1810 (e.g., LPG 1510 of FIG. 15 ), which can be communicatively coupled to an SSH VCN 1812 (e.g., SSH VCN 1512 of FIG. 15 ) via the LPG 1810 included in the SSH VCN 1812. SSH VCN 1812 can include SSH subnet 1814 (e.g., SSH subnet 1514 in FIG. 15 ), and SSH VCN 1812 can be communicatively coupled to control plane VCN 1816 (e.g., control plane VCN 1516 in FIG. 15 ) via LPG 1810 included in control plane VCN 1816 and to data plane VCN 1818 (e.g., data plane 1518 in FIG. 15 ) via LPG 1810 included in data plane VCN 1818. Control plane VCN 1816 and data plane VCN 1818 can be included in service tenancy 1819 (e.g., service tenancy 1519 in FIG. 15 ).
[0238] The control plane VCN 1816 may include a control plane DMZ layer 1820 (e.g., the control plane DMZ layer 1520 of FIG. 15) that may include a LB subnet 1822 (e.g., the LB subnet 1522 of FIG. 15), a control plane app layer 1824 (e.g., the control plane app layer 1524 of FIG. 15) that may include an app subnet 1826 (e.g., the app subnet 1526 of FIG. 15), and a control plane data layer 1828 (e.g., the control plane data layer 1528 of FIG. 15) that may include a DB subnet 1830 (e.g., the DB subnet 1730 of FIG. 17). LB subnet 1822 included in control plane DMZ tier 1820 can be communicatively coupled to app subnet 1826 included in control plane app tier 1824 and to an Internet gateway 1834 (e.g., Internet gateway 1534 in FIG. 15 ), which may be included in control plane VCN 1816, and app subnet 1826 can be communicatively coupled to DB subnet 1830 included in control plane data tier 1828 and to service gateway 1836 (e.g., service gateway in FIG. 15 ) and network address translation (NAT) gateway 1838 (e.g., NAT gateway 1538 in FIG. 15 ). Control plane VCN 1816 may include service gateway 1836 and NAT gateway 1838.
[0239] Data plane VCN 1818 may include a data plane app layer 1846 (e.g., data plane app layer 1546 in FIG. 15 ), a data plane DMZ layer 1848 (e.g., data plane DMZ layer 1548 in FIG. 15 ), and a data plane data layer 1850 (e.g., data plane data layer 1550 in FIG. 15 ). Data plane DMZ layer 1848 may include a LB subnetwork 1822, which may be communicatively coupled to a trusted app subnetwork 1860 (e.g., trusted app subnetwork 1760 in FIG. 17 ) and an untrusted app subnetwork 1862 (e.g., untrusted app subnetwork 1762 in FIG. 17 ) of data plane app layer 1846 included in data plane VCN 1818, as well as to an Internet gateway 1834. Trusted app subnet 1860 may be communicatively coupled to service gateway 1836 included in data plane VCN 1818, NAT gateway 1838 included in data plane VCN 1818, and DB subnet 1830 included in data plane data layer 1850. Untrusted app subnet 1862 may be communicatively coupled to service gateway 1836 included in data plane VCN 1818 and DB subnet 1830 included in data plane data layer 1850. Data plane data layer 1850 may include DB subnet 1830 that may be communicatively coupled to service gateway 1836 included in data plane VCN 1818.
[0240] The untrusted app subnet 1862 may include primary VNICs 1864(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 1866(1)-(N) that reside within the untrusted app subnet 1862. Each tenant VM 1866(1)-(N) can execute code in a respective container 1867(1)-(N) and can be communicatively coupled to an app subnet 1826 that can be included in a data plane app tier 1846 that can be included in a container egress VCN 1868. Each secondary VNIC 1872(1)-(N) can facilitate communication between the untrusted app subnet 1862 included in the data plane VCN 1818 and the app subnet included in the container egress VCN 1868. The container egress VCN may include a NAT gateway 1838 that can be communicatively coupled to the public internet 1854 (e.g., public internet 1554 in FIG. 15 ).
[0241] The internet gateway 1834 included in the control plane VCN 1816 and the internet gateway 1834 included in the data plane VCN 1818 can be communicatively coupled to a metadata management service 1852 (e.g., metadata management system 1552 of FIG. 15 ), which can be communicatively coupled to the public internet 1854. The public internet 1854 can be communicatively coupled to a NAT gateway 1838 included in the control plane VCN 1816 and the NAT gateway 1838 included in the data plane VCN 1818. The service gateway 1836 included in the control plane VCN 1816 and the service gateway 1836 included in the data plane VCN 1818 can be communicatively coupled to cloud services 1856.
[0242] In some examples, the pattern illustrated by the architecture of block diagram 1800 in FIG. 18 may be considered an exception to the pattern illustrated by the architecture of block diagram 1700 in FIG. 17 and may be desirable for an IaaS provider's customers when the IaaS provider cannot communicate directly with the customers (e.g., in a disconnected region). The customers can access each of the containers 1867(1)-(N) contained in each customer's VMs 1866(1)-(N) in real time. The containers 1867(1)-(N) can be configured to call each of the secondary VNICs 1872(1)-(N) contained in the app subnet 1826 of the data plane app tier 1846, which can be contained in the container egress VCN 1868. The secondary VNICs 1872(1)-(N) can send the call to a NAT gateway 1838, which can send the call to the public Internet 1854. In this example, containers 1867(1)-(N) that a customer can access in real time can be isolated from control plane VCN 1816 and can be isolated from other entities included in data plane VCN 1818. Containers 1867(1)-(N) can also be isolated from resources of other customers.
[0243] In another example, a customer can use containers 1867(1)-(N) to invoke cloud service 1856. In this example, the customer can execute code in containers 1867(1)-(N) that requests a service from cloud service 1856. Containers 1867(1)-(N) can send the request to secondary VNICs 1872(1)-(N), which can send the request to a NAT gateway that can send the request to public internet 1854. Public internet 1854 can send the request to LB subnet 1822, which is included in control plane VCN 1816, via internet gateway 1834. In response to determining that the request is valid, LB subnet 1826 can send the request to app subnet 1826, which can send the request to cloud service 1856 via service gateway 1836.
[0244] It should be understood that the illustrated IaaS architectures 1500, 1600, 1700, 1800 may include components other than those shown. Additionally, the illustrated embodiments are only some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, an IaaS system may have more or fewer components than those shown, may combine two or more components, or may have a different configuration or arrangement of components.
[0245] In some embodiments, the IaaS systems described herein may include a suite of application, middleware, and database service offerings that are self-service, subscription-based, elastically scalable, reliable, highly available, and securely delivered to customers. One example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the present assignee.
[0246] 19 illustrates an example computer system 1900 upon which various embodiments may be implemented. System 1900 may be used to implement any of the computer systems described above. As shown, computer system 1900 includes a processing unit 1904 that communicates with multiple peripheral subsystems via a bus subsystem 1902. These peripheral subsystems may include a processing acceleration unit 1906, an I / O subsystem 1908, a storage subsystem 1918, and a communication subsystem 1924. Storage subsystem 1918 includes a tangible computer-readable storage medium 1922 and a system memory 1910.
[0247] The bus subsystem 1902 provides a mechanism for allowing the various components and subsystems of the computer system 1900 to communicate with each other as intended. While the bus subsystem 1902 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. The bus subsystem 1902 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which may be implemented as a Mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0248] Processing unit 1904, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 1900. Processing unit 1904 may include one or more processors. These processors may include single-core or multi-core processors. In some embodiments, processing unit 1904 may be implemented as one or more independent processing units 1932 and / or 1934, with a single-core or multi-core processor included in each processing unit. In other embodiments, processing unit 1904 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0249] In various embodiments, the processing unit 1904 may execute various programs according to program code and may maintain multiple simultaneously executing programs or processes. At any given time, some or all of the program code to be executed may reside in the processor 1904 and / or in the storage subsystem 1918. Through suitable programming, the processor 1904 may provide the various functions described above. The computer system 1900 may further include a processing acceleration unit 1906, which may include a digital signal processor (DSP), a special purpose processor, etc.
[0250] The I / O subsystem 1908 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may include, for example, a motion sensing and / or gesture recognition device such as a Microsoft Kinect® motion sensor that enables a user to control and interact with an input device such as a Microsoft Xbox® 360 game controller through a natural user interface using gestures and spoken commands. User interface input devices may also include an eye gesture recognition device such as a Google Glass® blink detector that detects eye activity from a user (e.g., "blinking" during picture taking and / or menu selection) and translates the eye gesture as input to an input device (e.g., Google Glass®). Additionally, the user interface input devices may include a voice recognition sensing device that allows a user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0251] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser distance measuring devices, and eye-tracking devices. Furthermore, user interface input devices may include medical imaging input devices, such as computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, positron emission tomography (PET) scanners, and medical ultrasound scanners. User interface input devices may also include audio input devices, such as MIDI keyboards, digital musical instruments, and the like.
[0252] User interface output devices may include display subsystems, indicator lights, or non-visual displays such as audio output devices. Display subsystems may be flat-panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touch screens, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 1900 to a user or to another computer. For example, user interface output devices may include various display devices that visually convey text, graphics, and audio / video information, such as, but not limited to, monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, and modems.
[0253] Computer system 1900 may include a storage subsystem 1918 that provides a tangible, non-transitory, computer-readable storage medium for storing software and data structures that provide the functionality of embodiments described in this disclosure. The software may include programs, code modules, instructions, scripts, etc., that, when executed by one or more cores or processors of processing unit 1904, provide the functionality described above. Storage subsystem 1918 may also provide a repository for storing data used in accordance with the present disclosure.
[0254] 19 , storage subsystem 1918 may include various components including system memory 1910, computer-readable storage medium 1922, and computer-readable storage medium reader 1920. System memory 1910 may store program instructions that are loadable and executable by processing unit 1904. System memory 1910 may also store data used during the execution of the instructions and / or data generated during the execution of the program instructions. A variety of different types of programs may be loaded into system memory 1910, including, but not limited to, client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), virtual machines, containers, etc.
[0255] System memory 1910 may also store operating system 1916. Examples of operating system 1916 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS operating systems. In some implementations in which computer system 1900 runs one or more virtual machines, the virtual machines, along with their guest operating systems (GOS), may be loaded into system memory 1910 and executed by one or more processors or cores of processing unit 1904.
[0256] The system memory 1910 may be configured differently depending on the type of computer system 1900. For example, the system memory 1910 may be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). Different types of RAM configurations may be provided, including static random access memory (SRAM), dynamic random access memory (DRAM), etc. In some implementations, the system memory 1910 may include a basic input / output system (BIOS), which contains the basic routines that help to transfer information between elements within the computer system 1900, such as during start-up.
[0257] Computer-readable storage media 1922 can represent remote, local, fixed, and / or removable storage devices, plus storage media, for temporarily and / or more permanently containing and storing computer-readable information used by computer system 1900, including instructions executable by processing unit 1904 of computer system 1900.
[0258] Computer-readable storage medium 1922 may include any suitable medium known or used in the art, including, but not limited to, storage media and communication media such as volatile and nonvolatile, removable and non-removable media, implemented in any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, Digital Versatile Disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or other tangible computer-readable medium.
[0259] By way of example, the computer-readable storage medium 1922 may include a hard disk drive that reads from or writes to non-removable, non-volatile magnetic media, a magnetic disk drive that reads from or writes to removable, non-volatile magnetic disks, and an optical disk drive that reads from or writes to removable, non-volatile optical disks such as CD-ROMs, DVDs, and Blu-Ray® disks or other optical media. The computer-readable storage medium 1922 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD disks, digital video tapes, etc. The computer-readable storage media 1922 can also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, volatile memory-based SSDs such as magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 1900.
[0260] Machine-readable instructions executable by one or more processors or cores of processing unit 1904 may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include physically tangible memory or storage devices, including volatile and / or non-volatile memory storage devices. Examples of non-transitory computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard drives, floppy drives, removable memory drives (e.g., USB drives), or other types of storage devices.
[0261] The communications subsystem 1924 provides an interface to other computer systems and networks. The communications subsystem 1924 serves as an interface for receiving data from the computer system 1900 and transmitting data from the computer system 1900 to other systems. For example, the communications subsystem 1924 may enable the computer system 1900 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 1924 may include a radio frequency (RF) transceiver component for accessing a wireless voice and / or data network (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution), WiFi (IEEE 802.11 family of standards, or other mobile communications technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some embodiments, the communications subsystem 1924 may provide wired network connectivity (e.g., Ethernet) in addition to or instead of a wireless interface.
[0262] In some embodiments, the communications subsystem 1924 may also receive incoming communications in the form of structured and / or unstructured data feeds 1926, event streams 1928, event updates 1930, etc., on behalf of one or more users who may use the computer system 1900.
[0263] By way of example, the communications subsystem 1924 may be configured to receive data feeds 1926 in real time from users of social networks and / or other communications services, such as web feeds such as Twitter® feeds, Facebook® updates, Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party sources.
[0264] Additionally, the communications subsystem 1924 may be configured to receive data in the form of a continuous data stream, which may include an event stream 1928 of real-time events and / or event updates 1930, which may be continuous or infinite in nature with no apparent end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0265] The communications subsystem 1924 may also be configured to output structured and / or unstructured data feeds 1926, event streams 1928, event updates 1930, etc. to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 1900.
[0266] The computer system 1900 can be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google® Glass head-mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or any other data processing system.
[0267] Because the nature of computers and networks is constantly changing, the description of computer system 1900 shown in the figure is intended merely as a specific example. Many other configurations are possible, having more or fewer components than the system shown in the figure. For example, customized hardware may be used, and / or particular elements may be implemented in hardware, firmware, software (including applets), or a combination. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other manners and / or methods for implementing various embodiments.
[0268] The embodiments may be implemented by using a computer program product comprising a computer program / instructions which, when executed by a processor, cause the processor to perform any of the methods described in this disclosure.
[0269] While specific embodiments have been described, various modifications, variations, alternative constructions, and equivalents are encompassed within the scope of the present disclosure. The embodiments are not limited to operating in one specific data processing environment, but can freely operate in multiple data processing environments. Furthermore, while the embodiments have been described using a particular sequence of transactions and steps, it should be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described sequence of transactions and steps. Various features and aspects of the above-described embodiments may be used individually or jointly.
[0270] Furthermore, while embodiments have been described using particular combinations of hardware and software, it should be recognized that other combinations of hardware and software are within the scope of the present disclosure. Embodiments may be implemented exclusively in hardware, exclusively in software, or using a combination thereof. Various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, when a component or service is described as being configured to perform certain operations, such configuration may be achieved, for example, by designing electronic circuitry to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations, or any combination thereof. Processes may communicate using various techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.
[0271] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and alterations may be made without departing from the broader spirit and scope as set forth in the appended claims. Accordingly, while specific disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.
[0272] In the context of describing the disclosed embodiments (particularly in the context of the claims which follow), use of the terms "a," "an," and "the" and similar referents should be construed to encompass both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" should be construed as open-ended terms (i.e., meaning "including, but not limited to"), unless otherwise noted. The term "connected" should be construed as contained within, attached to, or joined to one another, either partially or as a whole, even if there is intervening material. The recitation of ranges of values herein, unless otherwise indicated herein, is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. Unless otherwise indicated herein or clearly contradicted by context, all methods described herein can be performed in any suitable order. Any and all examples provided herein, or the use of exemplary language (e.g., "etc.") are intended merely to clarify the embodiments and do not limit the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0273] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as generally used to indicate that an item, term, etc. can be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is not intended to, and should not, generally imply that some embodiments require that at least one of X, at least one of Y, or at least one of Z, respectively, be present.
[0274] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Such variations will be readily apparent to those skilled in the art, and the present disclosure may be practiced in ways other than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, unless otherwise indicated herein, any combination of the above-described elements in all possible variations thereof is encompassed by the present disclosure.
[0275] Example embodiments of the present disclosure can be described in light of the following provisions. Clause 1. A method is disclosed. The method may include a computer system obtaining configuration data including a plurality of response levels that specify applicability of a respective set of curtailment actions to a plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of the first set of curtailment actions to the plurality of hosts. The method may include the computer system obtaining current values of aggregate power consumptions of the plurality of hosts. The method may include the computer system obtaining current values of aggregate power thresholds of the plurality of hosts. The method may include selecting a first response level from the plurality of response levels based at least on a difference between the current values of the aggregate power consumptions and the current values of the aggregate power thresholds. The method may include causing application of the first set of curtailment actions to at least one host of the plurality of hosts in accordance with the selected first response level.
[0276] Clause 2. The method of clause 1, further including the computer system executing a first set of reduction actions on the plurality of hosts according to the selected first response level.
[0277] Clause 3. The method of clause 1 or clause 2, further including: 1) determining whether the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold; and 2) selecting from a plurality of response levels in response to determining that the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold. The selected response level is applicable to a subset of the plurality of hosts.
[0278] Clause 4. The method of any of clauses 1-3, wherein the first response level specifies applicability of a first set of reduction actions to the plurality of hosts based at least on host attributes.
[0279] Clause 5. The method of clause 4, wherein the host attributes include a priority level. Clause 6. The method of any of clauses 1 through 5, wherein each successive response level of the plurality of response levels specifies an increasingly severe set of mitigation actions to be applied to the plurality of hosts.
[0280] Clause 7. The method of clause 6, wherein a first power cap value applied to the first host by a first curtailment action according to a first response level is lower than a second power cap value applied to the first host by a different curtailment action according to a second response level.
[0281] Clause 8. The method of clause 6, wherein: 1) the first response level specifies application of a first reduction action to a first subset of the plurality of hosts associated with a priority level lower than the first level; 2) the second response level specifies application of a second reduction action to a second subset of the plurality of hosts associated with a priority level lower than the second level; 3) the second level is higher than the first level; and 4) the second subset of the plurality of hosts does not include the first subset of the plurality of hosts.
[0282] Clause 9. The method of clause 6, wherein the selected first response level is the response level of the plurality of response levels associated with the least severe set of reduction actions that achieves a new value of the aggregate power consumption that is less than or equal to the current value of the aggregate power threshold.
[0283] Clause 10. The method of clause 6, wherein selecting a first response level from a plurality of response levels based at least on a difference between a current value of the aggregate power consumption and a current value of the aggregate power threshold includes: 1) determining that a first reduction in aggregate power consumption resulting from the first response level is greater than the difference; 2) determining that a second reduction in aggregate power consumption resulting from the second response level is less than the difference; and 3) selecting the first response level.
[0284] Clause 11. The method of any of clauses 1 to 10, wherein the current value of the aggregate power consumption breaches the current value of the aggregate power threshold due to a decrease in the aggregate power threshold, and the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold for a period of time.
[0285] Clause 12. The method of clause 11, wherein the decrease in aggregate power threshold is caused by at least one of 1) a degradation or failure of a temperature control system for a physical environment containing the plurality of hosts, or 2) a governmental reduction in power supply, or 3) an increase in external temperature.
[0286] Clause 13. The method of any of clauses 1-12, further comprising: 1) determining that during a first time period, a first value of the aggregate power threshold exceeds a current value of the aggregate power consumption; and 2) determining that during a current time period following the first time period, the current value of the aggregate power consumption exceeds the current value of the aggregate power threshold. In some embodiments, the current value of the aggregate power threshold during the current time period is less than the first value of the aggregate power threshold during the first time period.
[0287] Clause 14. The method of any of clauses 1 through 13, wherein the first curtailment action includes at least one of 1) enforcing a power cap on the host, 2) migrating workloads from the host, or 3) shutting down or suspending the host.
[0288] Clause 15. The method of any of clauses 1 to 14, further including the computer system obtaining a new value of the aggregate power threshold and performing recovery from the first response level based on at least the new value of the aggregate power threshold.
[0289] Clause 16. The method of clause 15, wherein performing recovery from the first response level includes at least one of: 1) quiescing power capping on the host; 2) migrating workloads back to the host; or 3) restarting a previously shut down or paused host.
[0290] Clause 17. The method of any of clauses 1 to 16, wherein a first response level from a plurality of response levels is performed in response to receiving user input at a user interface, the selected first response level being associated with a reduction that exceeds a difference between a current value of the aggregate power consumption and a current value of the aggregate power threshold.
[0291] Clause 18. A method is disclosed. The method may include a computer system obtaining configuration data including a plurality of response levels specifying applicability of respective sets of curtailment actions to a plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of the first set of curtailment actions to the plurality of hosts. The method may include the computer system obtaining a forecast value of aggregate power consumption of the plurality of hosts for a future time period. The method may include the computer system obtaining a forecast value of aggregate power thresholds of the plurality of hosts for a future time period. The method may include selecting a first response level from the plurality of response levels based on at least a forecast difference between the forecast value of aggregate power consumption and the forecast value of aggregate power thresholds. The method may include identifying one or more workloads (a) currently executing on the plurality of hosts and (b) that would be affected by applying the first set of curtailment actions to the plurality of hosts in accordance with the selected first response level. The method may include preemptively migrating affected workloads from the plurality of hosts to one or more other hosts before the future time period.
[0292] Clause 19. The method of clause 18 may further include determining a predicted aggregate power threshold value based on at least one of 1) the health of a temperature control system for a physical environment including the plurality of hosts, or 2) an announced or anticipated governmental power supply reduction, or 3) an expected increase in external temperature.
[0293] Clause 20. The method of clause 18 or clause 19 may further include determining a forecast of aggregate power consumption based on a historical pattern of aggregate power consumption. Further includes:
[0294] Clause 21. The method of clauses 18 through 20, wherein identifying one or more workloads that a) are currently executing on a plurality of hosts and b) would be affected by applying a first set of reduction actions to the plurality of hosts according to a selected first response level includes 1) determining that a particular reduction action will be applied to the first host according to a selected second first level, 2) determining that the first workload is currently executing on the first host, and 3) determining that the first workload will be affected by applying the first set of reduction actions to the plurality of hosts according to the selected first response level.
[0295] Clause 22. A system is disclosed. The system may include a memory configured to store instructions and one or more processors configured to execute the instructions to perform a method according to any of clauses 1 to 21.
[0296] Clause 23. A non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium can store instructions that, when executed by a processor, cause the processor to perform the method of any of clauses 11 to 21.
[0297] All references cited herein, including publications, patent applications, and patents, are incorporated by reference to the same extent as if each individual reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0298] While the foregoing specification has described aspects of the disclosure with reference to specific embodiments thereof, those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure can be used individually or jointly. Moreover, the embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. It is a method, The computer system includes obtaining configuration data that includes multiple response levels specifying the applicability of each set of reduction actions to multiple hosts, The first response level among the plurality of response levels specifies the applicability of a first set of reduction actions to the plurality of hosts, and the method further, The computer system obtains the current value of the aggregated power consumption of the multiple hosts, The computer system obtains the current value of the aggregated power threshold of the multiple hosts, At a minimum, the first response level is selected from the plurality of response levels based on the difference between the current value of the aggregated power consumption and the current value of the aggregated power threshold, A method comprising: causing the application of a first set of reduction actions to at least one host among the plurality of hosts, according to the selected first response level.
2. The method according to claim 1, further comprising the computer system performing a first set of reduction actions on the plurality of hosts according to a selected first response level.
3. To determine whether the current value of the aggregated power consumption exceeds the current value of the aggregated power threshold, The method further includes making a selection from a plurality of response levels in response to determining that the current value of the aggregated power consumption exceeds the current value of the aggregated power threshold, The method according to claim 1, wherein the selected response level is applicable to a subset of the plurality of hosts.
4. The method according to claim 1, wherein the first response level specifies the applicability of the first set of reduction actions to the plurality of hosts, based at least on host attributes.
5. The method according to claim 4, wherein the host attribute includes a priority level.
6. The method according to claim 1, wherein each consecutive response level of the plurality of response levels specifies a progressively more severe set of reduction actions to be applied to the plurality of hosts.
7. The method according to claim 6, wherein the first power limit applied to the first host by the first reduction action according to the first response level is lower than the second power limit applied to the first host by a different reduction action according to the second response level.
8. The first response level specifies the application of the first reduction action to a first subset of the multiple hosts associated with a priority level lower than the first level, The second response level specifies the application of the second reduction to a second subset of the multiple hosts associated with a priority level lower than the second level, The second level is higher than the first level. The method according to claim 6, wherein the second subset of the plurality of hosts does not include the first subset of the plurality of hosts.
9. The method according to claim 6, wherein the selected first response level is a response level associated with the least severe set of reduction actions among the plurality of response levels that achieve a new value of aggregated power consumption that is less than or equal to the current value of the aggregated power threshold.
10. At a minimum, selecting the first response level from the plurality of response levels based on the difference between the current value of the aggregated power consumption and the current value of the aggregated power threshold is: It is determined that the first reduction in aggregated power consumption resulting from the first response level is greater than the difference, It is determined that the second reduction in aggregated power consumption resulting from the second response level is smaller than the difference, The method according to claim 6, comprising selecting the first response level.
11. The method according to claim 1, wherein the current value of the aggregated power consumption exceeds the current value of the aggregated power threshold due to a decrease in the aggregated power threshold, and the current value of the aggregated power consumption exceeds the current value of the aggregated power threshold for a certain period of time.
12. The aforementioned decrease in the aggregated power threshold is Degradation or failure of the temperature control system for the physical environment including the aforementioned multiple hosts, or Government cuts to electricity supply, or The method according to claim 11, which is caused by at least one of the following: an increase in external temperature.
13. During the first period, it is determined that the first value of the aggregated power threshold exceeds the current value of the aggregated power consumption, The present period following the first period further includes determining that the current value of the aggregated power consumption exceeds the current value of the aggregated power threshold, The method according to claim 1, wherein the current value of the aggregated power threshold during the current period is less than the first value of the aggregated power threshold during the first period.
14. The method according to claim 1, wherein the first reduction action includes at least one of 1) enforcing a power cap on a host, 2) migrating the workload from the host, or 3) shutting down or suspending the host.
15. The computer system acquires a new value for the aggregated power threshold, The method according to claim 1, further comprising performing recovery from the first response level based on at least the new value of the aggregated power threshold.
16. The method according to claim 15, wherein performing the recovery from the first response level includes at least one of: 1) suspending power limit settings for the host; 2) transferring the workload back to the host; or 3) restarting the host which was previously shut down or suspended.
17. The method according to claim 1, wherein the application of the first set of reduction actions is performed in response to receiving user input in the user interface, and the selected first response level is associated with a reduction amount exceeding the difference between the current value of the aggregated power consumption and the current value of the aggregated power threshold.
18. It is a method, The computer system includes obtaining configuration data that includes multiple response levels specifying the applicability of each set of reduction actions to multiple hosts, The first response level among the plurality of response levels specifies the applicability of a first set of reduction actions to the plurality of hosts, and the method further, The computer system obtains a predicted value of the aggregated power consumption of the multiple hosts for a future period, The computer system obtains predicted values for the aggregated power thresholds of the multiple hosts for the future period, At a minimum, the first response level is selected from the plurality of response levels based on the predicted difference between the predicted value of the aggregated power consumption and the predicted value of the aggregated power threshold. (a) Identifying one or more workloads that are currently running on the plurality of hosts and (b) are likely to be affected by applying the first set of reduction actions to the plurality of hosts according to the selected first response level; A method comprising proactively migrating the affected workload from the plurality of hosts to one or more other hosts before the aforementioned future period.
19. The integrity of the temperature control system for the physical environment including the aforementioned multiple hosts, or Government-announced or anticipated reductions in electricity supply, or The method according to claim 18, further comprising determining the predicted value of the aggregated power threshold based on at least one of the expected rise in external temperature.
20. The method according to claim 18, further comprising determining the predicted value of the aggregated power consumption based on past patterns of the aggregated power consumption.
21. a) Identifying one or more workloads that are currently running on the plurality of hosts and b) are likely to be affected by applying the first set of reduction actions to the plurality of hosts according to the selected first response level, Determining that the specific reduction action is applied to the first host according to the selected second first level, Determining that the first workload is currently running on the first host, The method according to claim 18, comprising determining that the first workload is affected by applying a first set of reduction actions to the plurality of hosts according to the selected first response level.
22. Memory configured to store instructions, A system comprising one or more processors configured to execute the instructions described in any one of claims 1 to 21.
23. A program that, when executed by a processor, includes instructions causing the processor to perform the method according to any one of claims 1 to 21.