Techniques for Orchestrated Load Shedding
The described method addresses suboptimal load shedding in data centers by using workload identification and response levels to minimize power consumption impact, ensuring efficient and disruptive-free power management.
Patent Information
- Application Number
- JP2025526337
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-06-21
- Filing Date
- 2023-08-11
- Publication Date
- 2025-12-16
AI Technical Summary
Traditional methods of load shedding in data centers lack understanding of workloads and customers, leading to suboptimal power management and potential disruptions, and there is a need for a more intelligent and efficient approach to reduce power consumption while minimizing impact on customers and devices.
A dynamically orchestrated method that identifies workloads and response levels, applying curtailment actions based on workload impact and power reduction estimates, using machine learning for predictive power management and workload migration.
This approach allows for efficient power reduction with minimal disruption by selecting the least impactful response level, balancing power supply and demand, and avoiding unnecessary power failures.
Smart Images

Figure 2025540609000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 423,762, entitled "Orchestrated DC-Scale Load Shedding," filed November 8, 2022; U.S. Provisional Patent Application No. 63 / 439,576, entitled "Orchestrated DC-Scale Load Shedding," filed January 18, 2023; and U.S. Patent Application No. 18 / 338,962, entitled "Techniques for Orchestrated Load Shedding," filed June 21, 2023, the contents of which are incorporated herein by reference in their entireties for all purposes.
[0002] FIELD OF THE INVENTION The present disclosure generally relates to techniques for orchestrating the reduction of power consumption in a data center. The disclosed systems, methods, devices, and services enable a dynamically orchestrated approach to constraining power consumption in a data center while ensuring that the least impactful response is adopted with respect to customers, hosts, and / or workloads. [Background technology]
[0003] background Data centers consist of a power infrastructure that provides numerous safety features according to a hierarchical power distribution. Power is supplied by a local utility and allocated according to this power distribution hierarchy to various components of the data center, including power distribution units (PDUs) (e.g., transformers, distribution panels, busways, rack PDUs, etc.) and power-consuming devices (e.g., servers, network devices, etc.). This ensures that the power consumed by all downstream devices adheres to the power limits of each upstream device. During peak demand and / or component failures, the data center may be unable to handle the demand effectively, potentially resulting in widespread power failures and, at a minimum, significant disruptions to downstream devices, resulting in reduced processing capacity within the data center and / or power failures at the data center, thereby degrading the user experience. In a worst-case scenario, a power failure in one data center can trigger cascading power failures of other devices within the same or other data centers as workloads are redistributed in an attempt to recover from the initial outage. Additionally, in some embodiments, external factors (eg, mandates by government regulations) may require that power consumption at a particular data center be reduced for a particular period of time. Summary of the Invention [Problem to be solved by the invention]
[0004] Traditional methods of managing load shedding have involved operators manually shutting down host racks or individual devices one by one to reduce power consumption. Alternatively, traditional techniques include pressing emergency shutoff switches or buttons configured to turn off power to all or most of the data center. Operators determining which components to power down typically lack understanding of the workloads or customers affected by their actions, or the extent of the impact of taking those actions. Traditional approaches such as those described herein result in suboptimal approaches with respect to the affected customers and workloads. More advanced approaches may have less impact on customers and / or devices / workloads while still allowing for the power reduction desired in a given situation. Therefore, it is desirable to improve power management techniques, particularly with respect to load shedding, to reduce power consumption to an impact and amount that is sufficient and avoids the potential hazards of traditional approaches.
[0005] Quick Overview In the following description, for purposes of explanation, specific details are set forth in order to provide a thorough understanding of some embodiments. However, it will be apparent that various embodiments may be practiced without these specific details. The illustrations and description are not intended to be limiting. [Means for solving the problem]
[0006] Some embodiments may include a method. The method may include a computer system identifying respective sets of workloads executing on each of a plurality of hosts. The method may include the computer system identifying a plurality of response levels that specify applicability of the respective sets of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels may specify applicability of the first set of curtailment actions to a plurality of resources. The method may include determining a first estimate for power reduction resulting from the first response level based at least on (a) the respective sets of workloads executing on each of the plurality of hosts and (b) applicability of the first set of curtailment actions to the plurality of hosts according to the first response level. The method may include selecting a first response level from the plurality of response levels based at least on the first estimate for power reduction resulting from the first response level. The method may include applying the first set of curtailment actions to the plurality of hosts according to the selected first response level.
[0007] In some embodiments, the method may further include determining a second estimate of a power reduction resulting from a second response level of the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of a second set of curtailment actions at the plurality of hosts according to the second response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on the first estimate of a power reduction resulting from the first response level.
[0008] In some embodiments, determining a first estimate for power reduction resulting from the first response level includes any suitable combination of: 1) determining that a first set of reduction actions are applicable to a first subset of the plurality of hosts; 2) identifying a respective set of workloads executing on each of the first subset of hosts; and / or 3) determining the power consumption of the respective set of workloads executing on each of the first subset of hosts.
[0009] In some embodiments, determining the first estimate for the power reduction resulting from the first response level further includes determining a sum of the respective power consumptions of the respective sets of workloads executing on each of the first subset of hosts as the first estimate for the power reduction resulting from the first response level.
[0010] In some embodiments, determining the first estimate for the power reduction resulting from the first response level further includes any suitable combination of: 1) determining an estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts after application of the first set of reduction actions to the first subset of hosts; and / or 2) determining the first estimate for the power reduction resulting from the first response level based on (a) the estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts and (b) the estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts after application of the first set of reduction actions to the first subset of hosts.
[0011] In some embodiments, the method further includes 1) determining a difference between current values of aggregate power consumption of the plurality of hosts and current values of aggregate power thresholds of the plurality of hosts, and 2) determining that a first value for a power reduction resulting from a first response level is greater than the difference. In some embodiments, selecting the first response level is based at least on determining that the first value for a power reduction resulting from the first response level is greater than the difference.
[0012] In some embodiments, the respective sets of workloads executing on each of the plurality of hosts and the plurality of response levels specifying applicability of the respective sets of curtailment actions to the plurality of hosts are identified based, at least in part, on at least one of: 1) identifying a degradation or failure of a temperature control system for a physical environment including the plurality of hosts; or 2) identifying a governmental power supply curtailment; or 3) identifying an increase in external temperature.
[0013] A second method is disclosed herein. The second method may include a computer system identifying a respective set of workloads executing on each of a plurality of hosts. The second method may include the computer system identifying a plurality of response levels that specify applicability of a respective set of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of a first set of curtailment actions to a plurality of resources such that a first curtailment action of the first set of curtailment actions is applicable to a first resource of the plurality of resources. The second method may include determining a first estimate of a predicted impact resulting from the first response level based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions on the plurality of hosts according to the first response level. In some embodiments, the impact on the workloads is determined based on at least one of a priority of the hosts, or a number of affected hosts, a number of affected workloads, a priority level of the affected workloads, a number of affected customers, or a priority level of the affected customers. The second method may include selecting a first response level from a plurality of response levels based at least on a first estimate of an impact on a workload resulting from the first response level, and applying a first set of reduction actions to a plurality of hosts according to the selected first response level.
[0014] In some embodiments, the second method may include determining a second estimate of an impact on the workload resulting from a second response level among the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) applicability of a second set of reduction actions at the plurality of hosts according to the second response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on the first estimate of an impact on the workload resulting from the first response level and the second estimate of an impact on the workload resulting from the second response level.
[0015] In some embodiments, the first set of reduction actions are more severe than the second set of reduction actions, and the first estimate of the impact to the workload resulting from the first response level is less than the second estimate of the impact to the workload resulting from the second response level.
[0016] In some embodiments, determining the first estimate of the impact on workloads resulting from the first response level includes at least one of: 1) determining that a first set of reduction actions are applicable to a first subset of the plurality of hosts; 2) identifying a respective set of workloads executing on each of the first subset of hosts as workloads affected by the first response level; or 3) determining at least one of (a) a number of workloads affected by the first response level, or (b) a priority type of the workloads affected by the first response level.
[0017] In some embodiments, determining the first estimate of the impact on workloads resulting from the first response level includes at least one of: 1) determining that a first set of reduction actions are applicable to a first subset of the plurality of hosts; 2) identifying respective sets of workloads running on each of the first subset of hosts; 3) identifying respective customers for the respective sets of workloads running on each of the first subset of hosts as customers affected by the first response level; or 4) determining at least one of (a) a number of customers affected by the first response level, or (b) a priority type of customers affected by the first response level.
[0018] In some embodiments, the second method further includes determining a second estimate of a power reduction resulting from the first response level based at least on (a) the respective sets of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions at the plurality of hosts according to the first response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on the first estimate of an impact on the workload resulting from the first response level and the second estimate of a power reduction resulting from the first response level.
[0019] In some embodiments, the selected first response level is a response level among the plurality of response levels that is associated with the least workload impact that achieves a power reduction equal to or greater than the difference between the current value of the aggregate power consumption of the plurality of hosts and the current value of the aggregate power threshold of the plurality of hosts.
[0020] A third method is disclosed herein. The third method may include a computer system determining a predicted set of workloads executing on each of a plurality of hosts during a future time period. The third method may include the computer system identifying a plurality of response levels specifying applicability of a respective set of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of a first set of curtailment actions to a plurality of resources. The third method may include determining a first estimate for power reduction resulting from the first response level based at least on (a) the predicted set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions on the plurality of hosts according to the first response level. The third method may include selecting a first response level from the plurality of response levels based at least on the first estimate for power reduction resulting from the first response level. A third method may include (a) identifying one or more workloads that are currently executing on a plurality of hosts and (b) that would be affected by applying a first set of reduction actions to the plurality of hosts according to a selected first response level. The third method may include preemptively migrating the affected workloads from the plurality of hosts to one or more other hosts in advance of a future time period.
[0021] In some embodiments, determining the respective predicted set of workloads executing on each of the plurality of hosts during the future time period is based on historical patterns of workloads executing on the plurality of hosts.
[0022] In some embodiments, determining the respective predicted set of workloads executing on each of the plurality of hosts during the future time period is performed using a machine learning model that has been pre-trained using a supervised learning algorithm to predict the set of workloads based, at least in part, on historical workload data provided as training data.
[0023] In some embodiments, the third method includes at least one of: 1) a computer system obtaining a forecast of aggregate power consumption of a plurality of hosts over a future time period; 2) a computer system obtaining a forecast of aggregate power thresholds of a plurality of hosts over a future time period; or 3) selecting from a plurality of response levels in response to determining that the forecast of aggregate power consumption exceeds the forecast of aggregate power thresholds.
[0024] In some embodiments, the third method includes at least one of: 1) a computer system identifying a predicted failure of a temperature control system associated with a plurality of hosts over a future time period; 2) the computer system obtaining a predicted value of an aggregate power threshold of the plurality of hosts over a future time period based, at least in part, on identifying the predicted failure; or 3) selecting from a plurality of response levels in response to identifying the predicted failure.
[0025] In some embodiments, identifying predicted faults in the temperature control system utilizes a machine learning model that is pre-trained using training data that includes historical data associated with a plurality of hosts. In some embodiments, the machine learning model is trained using a supervised learning algorithm to identify predicted faults from input data.
[0026] In some embodiments, the historical data in the training data includes temperature control system data, historical power consumption data corresponding to the set of hosts, and historical aggregate power thresholds.
[0027] Systems, devices, and computer media are disclosed, each of which may include one or more memories capable of storing instructions corresponding to the methods disclosed herein. The instructions may be executed by one or more processors of the disclosed systems and devices to perform the methods disclosed herein. One or more computer programs may be configured to perform specific operations or acts corresponding to the described methods by including instructions that, when executed by a data processing device, cause the device to perform those acts. [Brief explanation of the drawings]
[0028] [Figure 1] FIG. 1 illustrates an example physical environment (e.g., a data center or portion thereof) including various components, according to at least one embodiment. [Figure 2] FIG. 1 is a simplified diagram of an exemplary power distribution infrastructure including various components of a data center, according to at least one embodiment. [Figure 3] FIG. 1 illustrates an example architecture of an exemplary power orchestration system configured to orchestrate power consumption reductions applicable to various resources in a physical environment (e.g., a data center, a data center room, etc.), according to at least one embodiment. [Figure 4] FIG. 1 illustrates an example architecture of a power management service for detecting and constraining excessive power consumption, according to at least one embodiment. [Figure 5] 3 illustrates an example power distribution hierarchy corresponding to the arrangement of components in FIG. 2, according to at least one embodiment. [Figure 6] FIG. 1 is a flow diagram illustrating an example method for managing excessive power consumption, according to at least one embodiment. [Figure 7]FIG. 1 illustrates an example architecture of a VEPO orchestration service for orchestrating power consumption constraints and / or power shutdown tasks, according to at least one embodiment. [Figure 8] 1 is a table illustrating an example set of response levels and a respective set of reduction actions for each response level, according to at least one embodiment. [Figure 9] FIG. 1 illustrates an example range of components affected by one or more response levels utilized by a VEPO orchestration service to orchestrate power consumption constraints and / or power shutdown tasks, according to at least one embodiment. [Figure 10] FIG. 1 is a block diagram illustrating an example use case in which multiple response levels are applied based on the current state of a set of hosts, according to at least one embodiment. [Figure 11] 1 is a schematic diagram of an example user interface according to at least one embodiment. [Figure 12] FIG. 1 is a flow diagram illustrating an example method for training one or more machine learning models, according to at least one embodiment. [Figure 13] FIG. 1 is a block diagram illustrating an example method for implementing a response level from a plurality of response levels based at least in part on an estimated power reduction, according to at least one embodiment. [Figure 14] FIG. 1 is a block diagram illustrating an example method for implementing a response level from a plurality of response levels based at least in part on a predicted impact, according to at least one embodiment. [Figure 15] FIG. 1 is a block diagram illustrating an example method for preemptively migrating workloads affected by selected response levels, according to at least one embodiment. [Figure 16]FIG. 1 is a block diagram illustrating one pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 17] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 18] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 19] FIG. 1 is a block diagram illustrating another pattern for implementing a cloud infrastructure system as a service, according to at least one embodiment. [Figure 20] FIG. 1 is a block diagram illustrating an example computer system according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0029] Detailed Description In the following description, various embodiments are described. For purposes of explanation, specific configurations and details are set forth in order to provide a thorough understanding of the embodiments. However, it will also be apparent to one skilled in the art that the embodiments may be practiced without the specific details. Additionally, well-known features may be omitted or simplified so as not to obscure the described embodiments.
[0030] This disclosure relates to managing power consumption and orchestrating power consumption reductions within an environment (e.g., in hosts or power distribution units of a data center, or portions thereof). More particularly, techniques are described for enabling orchestrated load shedding to implement power consumption reductions due to current conditions and according to aggregate power thresholds.
[0031] It is desirable to balance the supply and consumption of power in a data center. If the power consumed exceeds the available supply, the balance can be restored by increasing the power supply or slowing down the rate of power consumption. If this balance is not maintained and a component attempts to consume more than the available supply, a circuit breaker may trip and disconnect the component from the supply.
[0032] The maximum aggregate power threshold for a particular environment can depend on many factors, including, but not limited to, the current power consumption values of hosts in the environment, the operational status of associated components such as temperature control systems (e.g., HVAC, chillers, etc.), and environmental conditions (e.g., ambient temperature, external temperature outside the environment, etc.). Conventional systems are not configured to manage power consumption based on these factors. When power consumption peaks and a power failure is imminent, conventional techniques primarily utilize manual intervention to shut down or suspend hosts. This leads to suboptimal load shedding strategies because operators implementing load shedding typically lack understanding of the effects of their actions. For example, when selecting devices to shut down or suspend, operators have little, if any, knowledge of customers, workloads, instances, or hosts, or their respective priorities. A power failure can cause an interruption to operations performed by components in a data center. For example, a website hosted on a server in a data center crashes if a tripped circuit breaker disconnects the server from the data center's power source. A non-optimal load shedding strategy may be inefficient and may result in more than sufficient power reduction to reduce aggregate power consumption by a desired amount. Additionally or alternatively, a non-optimal load shedding strategy may result in a broader / larger impact on the underlying hosts, instances, workloads, and corresponding customers than may be desirable.
[0033] The balance between power supply and power consumption in a data center can be managed by maintaining an equilibrium between available supply and demand and / or consumption. Because power is often statically allocated in long-term contracts with electric utilities, increasing or decreasing the power supply in a data center may sometimes be infeasible. While supply may be statically allocated, power consumption in a data center can change, sometimes drastically. As an overly simplistic example, as the number of threads executed by a server's processor increases, the server's power consumption may increase. The ambient temperature within a data center can affect the power consumed by the data center's cooling system. The cooling system operates at a higher load, consuming more power as it operates to reduce the ambient temperature experienced within the data center. If the cooling system fails, the data center may be unable to withstand the heat generated at the current power consumption level. Demand caused by some components in a data center, such as uninterruptible power supplies, power distribution units, cooling systems, or busways, can be difficult to regulate. However, some components, such as servers, virtual machines, and / or bare metal instances, can be more easily constrained. Additionally, workloads and / or instances can be migrated to other hosts and / or instances to concentrate resources on a smaller subset of hosts, thereby resulting in more idle and / or free hosts. Power reductions can then be applied to the idle / free hosts to ensure minimal impact (e.g., number of hosts, customers, instances, workloads).
[0034] Many conventional power management techniques utilize power capping to constrain operation in power-consuming devices (e.g., servers, network devices, etc.) in a data center. When power capping is utilized, a power capping limit can be used to constrain the power consumed by the device. The power capping limit is used to constrain (e.g., throttle) operation in a server to ensure that the server's power consumption does not exceed the power cap limit. Using power capping ensures that the allocated power limit of upstream devices is not breached, and each upstream device is provisioned according to a worst-case scenario in which each downstream device is estimated to consume its respective allocated maximum power. However, these downstream devices may often consume less power than their allocated maximum power, leaving at least a portion of the power allocated to the upstream device unused. These techniques waste valuable power and limit the density of power-consuming devices that can be utilized in a data center.
[0035] An efficient power infrastructure within a data center is necessary to increase provider profit margins, manage scarce power resources, and make the services the data center offers more environmentally friendly. Data centers that include components hosting multi-tenant environments (e.g., public clouds) can experience higher-than-average consumption because not all tenancies of the cloud are in use simultaneously. To improve data center efficiency and resource utilization, data center providers may increase servers and / or tenancies so that power consumption across all power-consuming devices approaches the data center's allocated power capacity. However, in some cases, reducing the gap between allocated power capacity and power consumption increases the risk of tripping circuit breakers and losing the ability to utilize computing resources. The techniques described herein minimize the frequency with which downstream devices are constrained, allowing these devices to utilize previously unused power while maintaining a high degree of safety regarding avoiding power failures. Technical effects The disclosed systems and methods provide an automated, dynamically orchestrated approach to constraining power consumption (e.g., via power capping, workload migration, etc.), suspending hosts, migrating instances and / or hosts, and / or shutting down hosts to achieve desired power reductions. In some embodiments, at least some of the functions performed toward these actions are user-selectable and / or based on user input. The described technology provides various response levels. Each response level can be associated with a set of reduction actions (referred to herein as “actions” for brevity). Each response level (referred to herein as “levels” for brevity) may provide a set of increasingly severe reduction actions to apply. If an aggregate power threshold is breached, a response level may be selected (e.g., by the system, based on user input, etc.) based on an estimated power reduction corresponding to each of the response levels. When implemented, the least severe and / or least impactful level sufficient to bring the aggregate power consumption of the data center below the aggregate power threshold may be selected (e.g., by the system, based on user input, etc.).
[0036] The estimated impact of applying the reduction action associated with the level may be determined at runtime based on the current attributes of the workload implemented by the virtual machine and / or bare metal instance, the customer with which the workload and / or affected host is associated, and the priorities corresponding to those workloads, hosts, and / or customers. Demand (e.g., corresponding to an aggregate power threshold) may change dynamically based on operational status data and environmental data corresponding to various components of the data center. These factors may be utilized to identify changes in aggregate power thresholds (e.g., aggregate power consumption / heat / demand thresholds that the components of the data center can collectively manage) as they occur in power management capabilities within the data center. The selection of a response level may be triggered by real-time changes in demand, via request (e.g., by user request, by a request submitted by a government agency mandating a reduction in power consumption, etc.), or by any suitable trigger, and may be based on the impact of implementing the corresponding reduction action.
[0037] By utilizing these techniques described herein, power constraints, migration tasks, and / or suspend or shutdown tasks employed can be tailored to the current state and resources (e.g., hosts, instances, workloads) within the datacenter, such that excessive curtailment is mitigated or avoided entirely. These techniques provide real-time capabilities to reduce the impact of power management response actions on customers, hosts, instances, and / or workloads, while ensuring that the risk of power failure is effectively avoided. The systems and methods described herein provide a more efficient and effective power orchestration approach than conventional systems, which can result in a more satisfying user experience. The particular action to implement may be selected (e.g., automatically by the system, through user selection, etc.) based, at least in part, on any suitable combination of: 1) maximizing overall power consumption reduction, 2) identifying / implementing the least severe response level, 3) identifying / implementing the least impactful response level, and / or 4) identifying a response level whose estimated power reduction is sufficient to reduce the current aggregate power consumption value below the current or projected aggregate power threshold while providing the least amount of excessive power reduction (e.g., reduction beyond that required to lower the current aggregate power consumption below the current aggregate power threshold). In this manner, the present technology provides a more flexible and intelligent approach to shedding loads in various situations and / or due to various triggering events.
[0038] 1 illustrates an example environment (e.g., environment 100) including various components, according to at least one embodiment. Environment 100 may be a physical environment, such as a data center (e.g., data center 102) or a portion thereof (e.g., a room in a data center, such as room 110A). Environment 100 may include a dedicated space for hosting any suitable number of servers, such as servers 104A-104P (collectively referred to as "servers 104"), and infrastructure for hosting those servers, such as networking hardware, a cooling system (also referred to as a "temperature control system"), and storage devices. Servers 104A-104P may also be referred to as "hosts." Networking hardware (not shown) in data center 100 allows remote users to interact with the servers over a network (e.g., the Internet). Any suitable number (e.g., 10, 14, 21, 42, etc.) of servers 104 may be held in various racks, such as racks 106A-106H (collectively referred to as "racks 106"). Racks 106 may include frames or enclosures in which corresponding sets of servers are positioned and / or mounted.
[0039] Various subsets of racks 106 can be organized into groups referred to as “rows” (e.g., rows 108A-108D, collectively referred to as “rows 108”). In some implementations, rows 108 can include any suitable number of racks (e.g., 5, 8, 10, up to 10, etc.) arranged (e.g., within a threshold distance of each other). In other implementations, rows can be organizational units, and racks comprising a given row can be located in different locations (not necessarily within a threshold distance of each other). As an example, rows 108 can be located in rooms (e.g., room 110A, room 110N, etc.). A room (e.g., room 110A) can be a section of a building or a physical enclosure or physical environment in which any suitable number of racks 106 are located. In other embodiments, a room can be an organizational unit, and rooms can be located in different physical locations, or multiple rooms can be located in a single section of a building.
[0040] Various temperature control systems (e.g., temperature control systems 112A-112N, temperature control systems 114A-114N, etc.) can be configured to manage the ambient temperature of a data center or portion thereof. As a non-limiting example, temperature control systems 112A-112N may be associated with room 110A, and temperature control systems 114A-114N may be associated with room 110N. Any suitable number of temperature control systems may be associated with a data center and / or portion thereof. In some embodiments, these temperature control systems may be any suitable heating, ventilation, and air conditioning (HVAC) devices (e.g., air conditioning units), chillers (e.g., water circulators that control temperature by circulating a liquid such as water), or the like. In some embodiments, each temperature control system may be associated with a corresponding amount of heat it is capable of and / or configured to manage (e.g., the amount of heat generated by a corresponding amount of power consumption of servers 104).
[0041] FIG. 2 illustrates a simplified diagram of an exemplary power distribution infrastructure 200 including various components (e.g., components of the data center 102 of FIG. 1 ) according to at least one embodiment. The power distribution infrastructure 200 can be connected to a utility power source (not shown), and power can be initially received by one or more uninterruptible power supplies (uninterruptible power supply (UPS) 202). In some embodiments, power may be received at the UPS from a utility company via an on-site power substation (not shown) configured to establish suitable voltage levels for distributing power throughout the data center. The UPSs 202 may each include dedicated batteries or generators to provide emergency power in the event of a failure of the input power source. The UPSs 202 can monitor the input power and provide backup power if a drop in input power is detected.
[0042] The power distribution infrastructure 200 may include any suitable number of intermediate power distribution units (PDUs) (e.g., intermediate PDUs 204) that connect to and receive power / electricity from the UPS 202. Any suitable number of intermediate PDUs 204 may be disposed between a UPS (of the UPS 202) and any suitable number of string PDUs (e.g., string PDUs 206). A power distribution unit (e.g., intermediate PDUs 204, string PDUs 206, rack PDUs 208, etc.) may be any suitable device configured to control and distribute power / electricity. Example power distribution units may include, but are not limited to, a main distribution panel, a distribution panel, a remote power panel, a bus bar, a power strip, a transformer, etc. Power may be provided from the UPS 202 to the intermediate PDUs 204. The intermediate PDUs 204 may distribute power to downstream components of the power distribution infrastructure 200 (e.g., string PDUs 206).
[0043] Power distribution infrastructure 200 may include any suitable number of string power distribution units (including string PDUs 206). A string PDU may include any suitable PDU (e.g., remote power panels, bus bars / ways, etc.) disposed between an intermediate PDU (e.g., a PDU among intermediate PDUs 204) and one or more rack PDUs (e.g., rack PDU 208A, rack PDU 208N, collectively referred to as “rack PDUs 208”). A “string PDU” refers to a PDU configured to distribute power to one or more strings of devices (e.g., string 210 including servers 212A-212D, collectively referred to as “servers 212”). As discussed above, a string (e.g., string 210) may include any suitable number of racks (e.g., racks 214A-214N, collectively referred to as “racks 214”) in which servers 212 are located.
[0044] Power distribution infrastructure 200 may include any suitable number of rack power distribution units (including rack PDUs 208). A rack PDU may include any suitable PDU positioned between a column PDU (e.g., column PDU 206) and one or more servers (e.g., servers 212A, 212B, etc.) corresponding to a rack (e.g., rack 214A, which is an example of rack 106 in FIG. 1). A "rack PDU" refers to any suitable PDU configured to distribute power to one or more servers in a rack. A rack (e.g., rack 214A) may include any suitable number of servers 212. In some embodiments, rack PDU 208 may include an intelligent PDU further configured to monitor, manage, and control consumption across multiple devices (e.g., rack PDU 208A, servers 212A, 212B, etc.).
[0045] The servers 212 (each an example of a server 104 in FIG. 1 ) may each include a power controller (power controllers 216A-216D, collectively referred to as “power controller 216”). A power controller refers to any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in that device. The power controllers 216 may individually monitor the power consumption of each server on which they operate. The power controllers 216 may each be configured to implement power capping to constrain the power consumption in their respective servers. Implementing power capping includes any suitable combination of monitoring the power consumption in the server, determining whether to constrain (e.g., constrain within a range, constrain, etc.) the power consumption in the server (e.g., based at least in part on a comparison of the server's current power consumption with a stored power capping limit), and constraining / constraining the power consumption in the server (e.g., using dynamic voltage and frequency scaling of processors and memory to throttle the server's power consumption). The implementation of power capping is sometimes referred to as "power capping."
[0046] The data center 102 of FIG. 1 may include various components shown in the power distribution infrastructure 200. By way of example, room 110A may include one or more busways (each an example of a row PDU 206). A busbar (also referred to as a "busway") refers to a duct of conductive material through which electrical power can be distributed (e.g., within room 110A). The busway may receive power from the power distribution units of the intermediate PDUs 204 and provide power to one or more racks (e.g., rack 106A, FIG. 1, rack 106B, FIG. 1, etc.) associated with a row (e.g., row 108A, FIG. 1). Each power infrastructure component that distributes / provides power to other components also consumes a portion of the power passing through it. This loss may be caused by heat loss due to the power flowing through that component or by power consumed directly by the component (e.g., power consumed by processors in a rack PDU).
[0047] FIG. 3 illustrates an example architecture of an exemplary power orchestration system 300 configured to orchestrate power consumption reductions applicable to various resources of a physical environment (e.g., a data center, a data center room, etc.) according to at least one embodiment. The term “resource” may be considered to include hosts, workloads (e.g., virtual machines and / or bare metal instances running those workloads), and / or customers associated with those hosts and / or workloads / instances. Power orchestration system 300 may be configured to monitor power consumption levels (e.g., current individual and / or aggregate power consumption values corresponding to various hosts, such as host 324, each of which is an example of server 104 in FIG. 1). Based at least in part on that monitoring, power orchestration system 300 may manage power consumption within the physical environment such that an aggregate power threshold is implemented. As used herein, an aggregate power threshold represents the maximum amount of power consumption that can be managed by components of system 300 given current conditions. The aggregate power thresholds may be dynamically adjusted as conditions change (e.g., based on government agency mandates to reduce power consumption, based on actual or predicted environmental conditions, or based on actual or predicted current power consumption values, based on the operational status of various components such as temperature control systems in the physical environment, etc.). The specific impact (e.g., scope, applicability) of changes made to implement the current aggregate power thresholds may change based on real-time conditions.Some example actions that may be taken to enforce the current aggregate power threshold may be setting a power cap (e.g., enforcing a maximum cap on power consumption in a particular device, such as one or more of the hosts 324), or otherwise allocating a budgeted amount of power to any suitable device (e.g., the host 324, the PDUs 202, 204, 206, 208 of FIG. 2, etc.), pausing or shutting down the host and / or instance (e.g., a VM or BM) and / or workloads running on the instance, migrating a workload from one instance to another, or migrating a workload from one host to another.
[0048] The power orchestration system 300 may include various components such as those shown in FIG. 3. For example, the power orchestration system 300 may include a VEPO orchestration service 302. The VEPO orchestration service 302 may be configured to obtain various data from which impact and mitigation actions may be determined. By way of example, the VEPO orchestration service 302 may be configured to obtain any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318.
[0049] The power data 310 may include any suitable budget / allocation values for the host 324 or PDUs, including the rack PDU 320 (an example of the rack PDU 208 in FIG. 2 ), power cap values, and / or current power consumption values indicating the amount of power currently being consumed by the device. The power data 310 may be stored in a location accessible to the VEPO orchestration service 302 and / or the power data 310 may be obtained from sources and / or services configured to manage and / or obtain such data. In some embodiments, the power management service 304 may be configured to obtain the power data 310 and store such data in a location from which the VEPO orchestration service 302 can retrieve the data. In some embodiments, the power management service 304, or another component configured to obtain such data, may provide the power data 310 directly to the VEPO orchestration service 302 according to a predefined schedule, frequency, or periodicity, or via request.
[0050] The account data 312 may include any suitable attributes of a customer or any suitable attributes of hosts and / or instances associated with the customer. Some example attributes may include a category associated with the customer's account (e.g., a free tier category indicating that the customer is using the service free of charge and may have limited features or resources; a free trial category indicating that the customer is using the service free of charge for a limited period of time). The category may correspond to a priority associated with the customer, or a priority may be assigned to the customer, host, or workload in other manners. The priority may indicate a relative degree of importance, such as low priority, medium priority, and high priority, although other priority schemes are contemplated. Another example attribute may include an identifier associated with the customer, host, instance, and / or workload. The account data 312 may be stored in a location accessible to the VEPO orchestration service 302 from which the account data 312 can be retrieved, or the account data 312 may be retrieved from a service and / or source configured to manage and / or retrieve such data (e.g., an account service in a cloud computing environment, not shown). In some embodiments, account data 310 to VEPO orchestration service 302 is directly from a source, or component configured to obtain such data, according to a predefined schedule, frequency, or periodicity, or via request.
[0051] Environmental data 314 may include any suitable data associated with the physical environment or environmental data indicative of external conditions. By way of example, environmental data 314 may include an ambient temperature reading for the physical environment, an external temperature outside the physical environment, a design point indicative of an external temperature the physical environment is designed to withstand, or the like. At least a portion of environmental data 314 may be collected from one or more sensors (not shown) configured to measure particular conditions, such as the ambient temperature within the physical environment or an external temperature occurring outside the physical environment. In some embodiments, at least a portion of environmental data 314 may be provided by a weather source, such as a weather forecast, stored within storage of power orchestration system 300 or may be retrievable from an external source, such as a weather forecast website. In some embodiments, environmental data 314 may include tables and / or protocols from which curtailed capacity / capacity may be determined / identified. As an example, the environmental data 314 may include a table or mapping indicating that a particular difference between the ambient temperature and temperatures occurring outside the physical environment (referred to as "external" temperatures) is associated with a particular reduced capacity of power consumption that a component of the system, or individual components, may monitor. The environmental data 314 may be stored in a location accessible to the VEPO orchestration service 302 from which the environmental data 314 can be retrieved, or the environmental data 314 may be obtained from a service and / or source configured to manage and / or obtain such data (e.g., a metrics service of a cloud computing environment, a power management service 304, etc.). In some embodiments, the environmental data 314 to the VEPO orchestration service 302 is directly from a source or component configured to obtain such data according to a predefined schedule, frequency, or periodicity, or via request.
[0052] Host / instance data 316 may include any suitable data associated with a host and / or instance. Host / instance data 316 may include workload metadata identifying workloads executing on a host and / or through a particular instance. In some embodiments, host / instance data 316 may identify a corresponding customer or corresponding priority of a host and / or instance and / or workload, etc. At least a portion of host / instance data 316 may initially be obtained and / or maintained by a separate service (e.g., compute service 306, an example of a compute service control plane in a cloud computing environment). Host / instance data 316 may be stored in a location accessible to VEPO orchestration service 302 and from which host / instance 316 can retrieve it, or host / instance 316 may obtain it from a service and / or source configured to manage and / or obtain such data (e.g., compute service 306, etc.). In some embodiments, host / instance data 316 may be provided directly to VEPO orchestration service 302 from a source, or a component configured to obtain such data, according to a predefined schedule, frequency, or periodicity, or via request.
[0053] The operational data 318 may include any suitable data associated with the operational status or state corresponding to one or more devices or components of the physical environment. As a non-limiting example, the operational data 318 may include the operational status of one or more temperature control systems (e.g., temperature control system 112 of FIG. 1 ). In some embodiments, the operational data 318 may indicate which temperature control systems are operational and / or the operational capabilities of those systems. In some embodiments, the environmental data 314 may include tables and / or protocols by which reduced capacity / capacity can be determined / identified. As an example, the operational data 318 may include a table or mapping indicating that a failure (e.g., complete or partial) of a particular component (e.g., a particular temperature control system) is associated with a particular reduced capacity of power consumption that the component of the system as a whole, or individual components of the system, can manage. At least a portion of the operational data 318 may initially be obtained and / or maintained by a separate service (e.g., power management service 304, etc.). The operational data 318 may be stored in a location accessible to the VEPO orchestration service 302 from which the operational data 318 can be retrieved, or the operational data 318 may be obtained from a service and / or source configured to manage and / or obtain such data (e.g., compute service 306, etc.). In some embodiments, the operational data 318 may be provided directly to the VEPO orchestration service 302 from a source or component configured to obtain such data according to a predefined schedule, frequency, or periodicity, or via request.
[0054] It should be understood that the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 may include current data indicating current values and / or historical data indicating corresponding past values. In some embodiments, future attributes of the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 may be predicted based, at least in part, on these past values. In some embodiments, machine learning may be utilized to predict any suitable portion of these future attributes. Exemplary methods for predicting future values of the power data 310, the account data 312, the environmental data 314, the host / instance data 316, or the operational data 318 are discussed in more detail with respect to FIG. 12 . In some embodiments, the VEPO orchestration service 302 can be configured to aggregate any suitable combination of power data 310, account data 312, environment data 314, host / instance data 316, or operational data 318 data into a table, mapping, or database from which the data can be filtered or sorted to identify a subset of resources (e.g., hosts, instances, workloads, customers) for a given set of constraints (e.g., low priority workloads, free tier customers, etc.).
[0055] The VEPO orchestration service 302 may be configured to identify and / or modify aggregate power thresholds associated with the physical environment based, at least in part, on any suitable combination of current / aggregate power consumption values (actual or predicted) associated with any suitable combination of hosts 324, received requests related to governmental entities (e.g., local power authorities) mandating / requiring specific or overall power reductions, possibly over a specific period of time (e.g., the next 24 hours), current or predicted environmental conditions (e.g., current or predicted ambient temperature, current or future external temperature, etc.), or current or predicted operational status corresponding to temperature control systems (e.g., current or predicted total / partial failure), etc.
[0056] The VEPO orchestration service 302, through associated functionality or functionality provided by other systems and / or services, can be configured to manage actual power consumption values corresponding to hosts / instances / workloads in the physical environment. Managing actual power consumption values can include setting power caps, pausing workloads, instances, hosts, shutting down workloads / instances / hosts, migrating instances from one host to another, or migrating workloads from one instance and / or host to another, etc.
[0057] The VEPO orchestration service 302 can obtain configuration data (e.g., mappings, protocol sets, rules, etc.) corresponding to multiple response levels. Each response level can be associated with a corresponding set of curtailment actions (e.g., power capping, migration, suspend, shutdown, etc.) that can be performed on components of the physical environment (e.g., servers 104, PDUs 202, 204, 206, 208, etc.). In some embodiments, each set of curtailment actions corresponding to a particular level is associated with implementing a different possible reduction in the aggregate power consumption of the data center. In some embodiments, the multiple response levels indicate increasing severity when curtailment actions are performed to reduce the aggregate power consumption of the data center. "Severity" is intended to refer to the relative degree to which an action is disruptive. As an example, an action to set a power cap on a server can be considered less severe than an action to completely shut down the server because in the former, the server can still provide some processing power, albeit at a reduced capacity, whereas in the latter, the server does not provide processing power. An example set of response levels is discussed in more detail with respect to FIG.
[0058] The VEPO orchestration service 302 can be configured to utilize configuration data corresponding to the specification of response levels and their corresponding reduction actions to determine the impact of applying a given level of reduction action to hosts / instances / workloads (collectively, “resources”) in a physical environment. The impact of applying a given action may depend on the current state of the resources at the time of execution. In some embodiments, identifying the impact of a given reduction action may include identifying a set of resources and / or customers to which the action, if implemented, is applicable. Thus, estimated impact is intended to refer to the scope or applicability of a given action or level. Identifying the scope and / or applicability of a given action or level may include identifying the specific resources and / or customers affected by implementing the action or level and / or any suitable attributes associated with the applicable resources and / or customers. As an overly simplistic example, the VEPO orchestration service 302 may identify that if an action is taken to shut down all idle hosts, X number of idle hosts among the hosts 324 will be affected, that those hosts are associated with a particular customer or number of customers, and / or that the hosts and / or customers are associated with other attributes such as category (e.g., free tier), priority (e.g., high priority), etc.
[0059] The VEPO orchestration service 302 can be configured to estimate the likely impact and / or actual power reduction for one or more of the levels based, at least in part, on configuration data specifying the levels and corresponding actions and any suitable combination of power data 310, account data 312, or host / instance data, etc. For one or more response levels or one or more reduction actions corresponding to a given level, an estimated impact (e.g., what hosts, instances, workloads, customers are likely to be affected, what attributes are associated with the affected hosts, instances, workloads, and / or customers, how many hosts / instances / workloads / customers are affected, etc.) can be identified. In some embodiments, the estimated impact of a given action can be aggregated with the estimated impact for all actions at a given level to determine the estimated impact for a given level.
[0060] The VEPO orchestration service 302 can be configured to determine an estimated impact and / or actual power reduction likely to occur for one or more of the levels based, at least in part, on current, historical, or forecasted data (e.g., current, historical, or forecasted values corresponding to power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318).
[0061] In some embodiments, the VEPO orchestration service 302 can be configured to present, recommend, or automatically select a given response level from a set of possible response levels based, at least in part, on any suitable combination of estimated impact and / or estimated power reduction likely to occur if a reduction action corresponding to the given level is implemented (e.g., realized). Thus, in some embodiments, a particular level can be recommended and / or selected by the power orchestration system 300 based on any suitable combination of: 1) determining a level that, if implemented, is likely to result in a sufficient, but not excessive, reduction in power consumption in resources of the physical environment (e.g., the smallest amount of power consumption reduction sufficient to reduce current power consumption values below the current aggregate power threshold); 2) determining a level that includes the least severe set of actions; or 3) identifying the least impactful level or action (e.g., the fewest number of potentially affected resources / customers, the set of affected resources and / or customers having a priority or category indicating the lowest degree of overall importance, etc.). Determining the least severe level or action, or the level or action having the least impact, may be determined independently of the estimated power consumption reduction likely to result from implementing that level and / or action, or determining the least severe level / action and / or the level / action having the least impact may be determined from one or more levels that, if implemented given the current conditions, are estimated to result in a power consumption reduction sufficient to bring the power consumption value down to an aggregated value that falls within the current aggregate power threshold.
[0062] The power management service 304 can be configured to provide functionality for identifying power cap values for hosts and / or instances based, at least in part, on aggregate power thresholds, budget / allocated power thresholds associated with PDUs, and / or current individual and / or aggregate power consumption values. In some embodiments, the power management service 304 can be configured to identify power caps and / or resources to which specific values of those power caps should apply, based, at least in part, on any suitable combination of current, historical, or predicted values of power data 310, account data 312, environmental data 314, and host / instance data 316. In some embodiments, the VEPO orchestration service 302 can trigger and utilize functionality provided by the power management service 304 for any suitable combination of identifying / changing budget / allocated power corresponding to one or more resources (e.g., hosts, instances, PDUs, etc.), identifying which or how many resources are affected, and identifying estimated power reductions that are likely to occur if a given level of action is implemented. In some embodiments, the power management service 304 can perform some or all of these functions.
[0063] The VEPO orchestration service 302 and / or the power management service 304 can host a user interface 308. The user interface 308 can be configured to provide any suitable application metadata corresponding to a combination of current individual and / or power consumption values corresponding to any suitable number or type of resources (e.g., hosts, instances, workloads, etc.) or customers, any suitable attributes (e.g., priority, category, etc.) associated with those resources or customers, a predicted state (e.g., predicted aggregate power threshold change), a current state (e.g., current aggregate power threshold), an estimated impact (e.g., number, identifier, or other attribute associated with the affected host / instance / workload customer), and / or an estimated power reduction (e.g., an estimated reduction that is expected to occur if a level action is applied), and / or an identification of an action corresponding to a given level. The user interface 308 can present any suitable combination of this application metadata. The application metadata can correspond to any suitable combination of available response levels. In some embodiments, application metadata corresponding to current / projected power consumption values and / or current / projected aggregate power thresholds can be displayed. In some embodiments, the application metadata presented in the user interface 308 can correspond to an estimated power consumption reduction and / or estimated impact for any suitable number of levels. In some embodiments, the VEPO orchestration service 302 can provide a recommendation for a particular level (e.g., via the provided application metadata corresponding to that level or a reduction action associated with that level), and confirmation / rejection of the recommended level can be entered via a selectable option in the user interface 308. If confirmed, the VEPO orchestration service 302 can be configured to perform an operation to effectuate implementation of the confirmed level and its corresponding action.In some embodiments, user interface 308 may present application metadata for any suitable number of response levels and / or corresponding reduction actions, and one or more options may be provided in user interface 308 to enable user selection of a particular level and / or action.
[0064] Upon receiving user input indicating the selection of a particular level and / or action, or confirmation of the selection of a particular level and / or action, the VEPO orchestration service 302 can perform operations to effectuate implementation of the confirmed level and its corresponding action. Effecting or implementing the confirmed level and / or its corresponding action may include instructing any suitable component of the power orchestration system (e.g., power management service 304, compute service 306) to implement any suitable portion of the level and / or corresponding action. Instructing one or more components may include providing any suitable data to the instructed component, such as any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and operational data 318. In some embodiments, the directed component may be configured to implement an action or level according to data provided by VEPO orchestration service 302, or the directed component may obtain any suitable portion of data from where power data 310, account data 312, environment data 314, host / instance data 316, or operational data 318 are stored.
[0065] As a non-limiting example, if the action at the level at which reduction or impact is to be estimated includes power capping, the power management service 304 can be configured to estimate a power cap, estimate individual or aggregate power consumption reductions, and / or implement the estimated power cap to effect the estimated individual / aggregate power consumption reductions. In some embodiments, the power cap can be implemented directly by the power management service 304 or by instructing the compute service 306. Other actions, such as migrating instances or workloads, pausing instances, workloads, or hosts, shutting down hosts, or preventing (at least temporarily) future allocation or launch of instances and / or workloads, can be identified by the VEPO orchestration service 302 and / or the compute service 306 and implemented by the VEPO orchestration service 302 through functionality provided by the compute service 306. In some embodiments, the compute service 306 can be configured to identify potentially affected hosts / instances / workloads / customers and provide such information to the VEPO orchestration service 302 and / or the power management service 304.
[0066] Compute service 306 can be configured to communicate with a baseboard management controller (BMC) / integrated lights out manager (ILOM) 322 of rack PDU 320. ILOM is one example of a BMC used for illustrative purposes. In some embodiments, BMC / ILOM 322 may be a service processor dedicated to monitoring the physical status of a host machine, in this example, rack PDU 320. Similarly, host 324 (e.g., a host in rack 319, to which rack PDU 320 manages power according to the power distribution hierarchy discussed in connection with FIG. 5 ) may include BMC / ILOM 328, which may be a service processor dedicated to monitoring the physical status of the host, in this example, one of hosts 324. BMC / ILOM 322 and BMC / ILOM 328 can be configured to, among other things, manage and / or enforce budget power and / or power caps in the device in which BMC / ILOM 328 operates or for downstream devices. For example, BMC / ILOM 322 may receive power cap values to be utilized by BMC / ILOM 328 to throttle power consumption for a given host. BMC / ILOM 322 may distribute the power caps to the corresponding BMC / ILOM 328 with which the power caps are associated. In some embodiments, BMC / ILOM 322 may distribute power caps stored on the applicable hosts, but enforcement of these power caps is not realized until further instructions are provided by BMC / ILOM 322 to BMC / ILOM 328. BMC / ILOM 328 may enforce power caps at the host level and / or instance level and / or workload level for one or more workloads (not shown) running on the instances.
[0067] BMC / ILOM 322 to BMC / ILOM 328 can be used to instruct devices (e.g., rack PDU 320 and / or hosts 324) to suspend operation of hosts / instances / workloads or to shut them down. In some embodiments, BMC / ILOM 322 can instruct BMC / ILOM 328 to resume operation of hosts / instances / workloads. If shut down, a device (e.g., via its BMC / ILOM) can be instructed (e.g., by Compute Services 306 and / or by an upstream device such as its PDU) to start up from shutdown. This can include Compute Services 306 sending instructions from BMC / ILOM 322 to BMC / ILOM 328 to suspend and / or shut down hosts / instances / workloads.
[0068] In some embodiments, compute service 306 may perform any suitable operation, such as migrating a workload from one instance to another (e.g., to an instance on the same host, to another instance running on a different host, etc.), migrating an instance from one host to another (e.g., to a host in the same rack, to a host in a different rack, etc.), migrating an instance / workload back to the host that originally hosted the previously migrated instance / workload, etc. In some embodiments, compute service 306 may be configured to ensure that a host / instance / workload is not assigned to a host and / or instance that has a level and / or action applied that has been achieved or is in the process of being achieved.
[0069] Any suitable combination of devices (e.g., host 324, rack PDU 320) may include one or more power controllers (e.g., BMC / ILOM 328, BMC / ILOM 322, respectively). A power controller may be any suitable hardware or software component configured to operate in a device (e.g., a server) and monitor and / or manage power consumption in the device. A power controller may individually monitor the power consumption of each device on which it operates. The power controllers may each be configured to implement power capping limits to constrain power consumption in the respective device. Implementing power capping limits may include monitoring power consumption in the device, determining whether to constrain (e.g., constrain within a range, constrain, etc.) the power consumption in the device (e.g., based at least in part on a comparison of the device's current power consumption with a stored power cap limit), and constraining / constraining the power consumption in the device (e.g., using dynamic voltage and / or frequency scaling in the device's processor and / or memory to throttle power consumption in the device). Any suitable action associated with power capping may be implemented by the power controller.
[0070] In some embodiments, a power controller (e.g., BMC / ILOM 322) can communicate with another power controller (e.g., BMC / ILOM 328) corresponding to one of the hosts 324) via a direct connection (e.g., via a cable) and / or via a network. The power controller (e.g., BMC / ILOM 328) can provide power consumption data indicating the device's current power consumption (e.g., cumulative power consumption over a period of time, current power consumption rate, etc.). The power consumption data can be provided to the power controller (e.g., BMC / ILOM 322) at any suitable frequency, periodicity, or according to a predefined schedule or event (e.g., when a predefined consumption threshold is breached, when a change in consumption rate reaches a threshold, when one or more predefined conditions are determined to be met, when thermal attributes of the device are determined, etc.).
[0071] In some embodiments, a power controller (e.g., BMC / ILOM 328) may receive a power cap value (also referred to as a "power cap") from a power controller (e.g., BMC / ILOM 322). In some embodiments, additional data may be provided along with the power cap value. By way of example, an indication of whether the power cap value should be applied immediately may be included with the power cap value. In some embodiments, the received power cap may be applied / enforced immediately by default. In other embodiments, the received power cap may not be applied / enforced immediately by default.
[0072] When applying / enforcing a power cap, the power controller (e.g., BMC / ILOM 328) may monitor power consumption at the device. This may include utilizing a metering device or software configured to identify / calculate power consumption data for the device (e.g., cumulative power consumption over a period of time, a current power consumption rate, a current change in power consumption rate over a time window, etc.). As part of applying / enforcing a power cap (also referred to as “power capping”), the power controller (e.g., BMC / ILOM 328) may determine whether to constrain (e.g., limit within a range, constrain, etc.) the power consumption at the device (e.g., based at least in part on comparing the device's current consumption rate with a stored power cap value). When constraining power consumption at the device, the power controller (e.g., BMC / ILOM 328) may limit / constrain the power consumption at the device (e.g., using dynamic voltage and frequency scaling on the processor and / or memory to throttle power consumption at the device). In some embodiments, the power controller (e.g., BMC / ILOM 328) can execute instructions to limit / constrain the power consumption of a device when the device's current consumption data (e.g., cumulative consumption, consumption rate over a time window, etc.) approaches an enforced power cap (e.g., breaches a threshold below the power cap). When constraining / limiting (also referred to as "throttling") power consumption at a device, the power controller (e.g., BMC / ILOM 328) ensures that the device's power consumption stays below the power consumption indicated by the power cap. The power controller (e.g., BMC / ILOM 328) can be configured to constrain power at a host generally or according to the instance and / or workload to which the power cap may be applied.
[0073] In some embodiments, the power controller of the host 324 can be configured to allow the host 324 to run unconstrained (e.g., without constraints based on power consumption and power caps) until instructed to apply / enforce a power cap by a power controller (e.g., BMC / ILOM 328). In some embodiments, the power controller (e.g., BMC / ILOM 328) may receive a power cap value that it may or may not be instructed to enforce later, but once received, the power cap value may be stored in memory without being utilized in power management at the device. Thus, in some embodiments, the power controller (e.g., BMC / ILOM 328) may not initiate a power capping operation (e.g., the comparison and determination described above) until instructed to do so (e.g., via an indication provided by the power controller 448). The indication to begin enforcing the power cap may be received along with the power cap value or may be received as a separate communication from the power controller 448.
[0074] According to power distribution hierarchy 500, each power controller (e.g., BMC / ILOM 322) of one or more PDUs can be configured to distribute, manage, and monitor power for any suitable number of devices (e.g., hosts 324 in rack 319). In some embodiments, a power controller (e.g., BMC / ILOM 322) may be a computing agent or program installed on a given PDU (e.g., rack PDU 320 of FIG. 3). A power controller (e.g., BMC / ILOM 322) can be configured to obtain power consumption data from hosts 324 in rack 319 that correspond to the devices with which it is associated (e.g., devices for which a given PDU is configured to distribute, manage, or monitor power). The power controller (e.g., BMC / ILOM 322) may receive the consumption data according to a predefined frequency, periodicity, or schedule implemented by the power controller (e.g., BMC / ILOM 328), and / or the power controller (e.g., BMC / ILOM 322) may receive the consumption data in response to requesting the consumption data from the power controller (e.g., BMC / ILOM 328). The power controller (e.g., BMC / ILOM 322) may be configured to request the consumption data from the power controller (e.g., BMC / ILOM 328) according to a predefined frequency, periodicity, or schedule implemented by the power controller (e.g., BMC / ILOM 322).
[0075] The power controller (e.g., BMC / ILOM 322) may transmit received consumption data to the power management service 402 (directly or via the compute service 306) at any suitable time, according to any suitable frequency, periodicity, or schedule, or as a result of one or more predefined conditions being met (e.g., the rate of change in the individual / cumulative / aggregate consumption of the hosts 324 breaches a threshold). In some embodiments, the power controller (e.g., BMC / ILOM 322) may aggregate power consumption data received from the hosts 324 (e.g., via each device's power controller (e.g., BMC / ILOM 328)) and then transmit the aggregated consumption data to the power management service 402. In some embodiments, the consumption data may be aggregated by host, instance type / category / priority, workload type / category / priority, or customer.
[0076] The power controller (e.g., BMC / ILOM 322) may receive one or more power cap values corresponding to the host(s) 324 (e.g., from the implementation manager 410 of the power management service 402) at any suitable time. These power cap values may be calculated by the power management service 402 (e.g., via the constraint specification manager 408). In some embodiments, the power controller (e.g., BMC / ILOM 322) may receive, along with the power cap values, a timing value indicating a duration to be used for the timer. The power controller (e.g., BMC / ILOM 322) may be configured to generate and start a timer having an associated duration / period corresponding to the timing value. When the timer expires, indicating that a period corresponding to the period has elapsed, the power controller (e.g., BMC / ILOM 322) may transmit data to one or more hosts including an indicator or other suitable data that instructs the one or more hosts to proceed with power capping using the power cap value previously stored in each device. In some embodiments, the power cap value may be provided along with the indicator. In other embodiments, the power cap value may be provided by the power controller (e.g., BMC / ILOM 322) to the device's power controller (e.g., BMC / ILOM 328) immediately upon receiving the power cap value from the power management service 402. Thus, in some embodiments, the power cap value is sent by the power controller (e.g., BMC / ILOM 322) to the power controller (e.g., BMC / ILOM 328) and stored in memory, but is not enforced by the power controller (e.g., BMC / ILOM 322) until the power controller (e.g., BMC / ILOM 328) receives subsequent data from the power controller (e.g., BMC / ILOM 322) that instructs the power controller (e.g., BMC / ILOM 328) to begin power capping.
[0077] The VEPO orchestration service 302, the power management service 304, the compute service 306, the rack PDU 320, and the hosts 324 (collectively referred to as the "devices of FIG. 3") can communicate over one or more wired or wireless networks (e.g., network 808). Storage devices on which the power data 310, the account data 312, the environment data 314, the host / instance data 316, and the operational data 318 are stored can also communicate with the VEPO orchestration service 302, the power management service 304, the compute service 306, the rack PDU 320, and the hosts 324 over one or more wired or wireless connections (e.g., over one or more networks). In some embodiments, the network can include any one or combination of many different types of networks, such as a cable network, the Internet, a wireless network, a cellular network, and other private and / or public networks.
[0078] 3 may be any suitable type of computing device, such as, without limitation, a server device, a network device, or any suitable device in a data center. In some embodiments, some of the devices (e.g., rack PDU 320 and host 324) are arranged in a power distribution hierarchy, such as power distribution hierarchy 500 discussed in connection with FIG. 5. In some embodiments, host 324 may correspond to and be represented by a level 1 node of power distribution hierarchy 500. Rack PDU 320 (the “rack leader”) may correspond to and be represented by a higher-level node (e.g., a level 2 node) of the power distribution hierarchy.
[0079] Each of the devices of Figure 3 may include at least one memory. Each of the processors may be implemented in hardware, computer-executable instructions, firmware, or a combination thereof, as appropriate. The computer-executable instructions or firmware implementation of the processor of the devices of Figure 3 may include computer-executable or machine-executable instructions written in any suitable programming language to perform the various functions described.
[0080] The memory of the devices of Figure 3 can store program instructions that are loadable and executable on each processor of a given device, as well as data generated during the execution of these programs. Depending on the configuration and type of user computing device, the memory can be volatile (such as random access memory (RAM)) and / or non-volatile (such as read-only memory (ROM), flash memory, etc.). The devices of Figure 3 can also include additional removable and / or non-removable storage devices, including but not limited to magnetic storage devices, optical disks, and / or tape storage devices. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computing device. In some implementations, the memory can individually include multiple different types of memory, such as static random access memory (SRAM), dynamic random access memory (DRAM), or ROM.
[0081] Referring more particularly to the contents of the memory, the memory may include an operating system, one or more data stores, and one or more application programs, modules, instances, workloads, or services, such as a VEPO orchestration service 302, a power management service 304, and a compute service 306.
[0082] The devices of Figure 3 may include communication connections that allow the devices to communicate with each other over a network (not shown). The devices of Figure 3 may also include I / O devices such as a keyboard, mouse, pen, voice input device, touch input device, display, speakers, printer, etc.
[0083] 4 illustrates an example architecture 400 of a power management service 402 (an example of power management service 304 of FIG. 3 ) according to at least one embodiment. The power management service may include an input / output processing manager 404 that may be configured to receive and / or transmit any suitable data between the power management service 402 and any other suitable device of FIG. 3 . As shown, the power management service 402 may include a consumption monitoring manager 406, a constraint specification manager 408, and an enforcement manager 410, although more or fewer computing components or subroutines may likewise be utilized.
[0084] Consumption monitor manager 406 may be configured to determine current individual and / or aggregate power consumption values (e.g., corresponding to examples of host 324 in FIG. 3 and server 104 in FIG. 1). Constraint identification manager 408 may be configured to perform any suitable function corresponding to identifying one or more power cap values for a host, instance, or workload (e.g., host 324). Enforcement manager 410 may be configured to enforce and / or trigger enforcement of power caps on any suitable device, according to at least one embodiment. Allocating power may refer to the process of allocating a budgeted amount of power (e.g., expected amount of load / power consumption) to any suitable component. A power cap is an example of a budgeted amount of power.
[0085] The functionality described in connection with power management service 402 may, in some embodiments, be performed by one or more virtual machines implemented within a hosted computing environment (e.g., in hosts 324 corresponding to any suitable number of servers 104 of FIG. 1). A hosted computing environment may include one or more rapidly provisioned and released computing resources, which may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment. A number of example cloud computing environments are provided and described in more detail below with respect to FIGS. 16-19.
[0086] The power management service 450 (e.g., consumption monitor manager 406) can be configured to receive / obtain consumption data from a PDU (e.g., from rack PDU 320 from a power controller such as BMC / ILOM 322 of FIG. 3). As described above, the consumption data may be aggregated or cumulative for one or more devices. As a non-limiting example, an instance of consumption data received by the power management service 402 may include cumulative / aggregated consumption data for all devices associated with (e.g., managed by) a given PDU. By way of example, the PDU providing consumption data may be a rack PDU (e.g., rack PDU 320), and the consumption data provided by that PDU may include cumulative / aggregated and / or individual consumption data values for all devices in the rack (e.g., hosts in hosts 324 associated with a given rack 319). In some embodiments, the cumulative / aggregated consumption data value may further include the power consumption of the PDU. From this consumption data, the power management service 402 can access individual and / or aggregate or cumulative power consumption data for the individual device or all of the devices with which that instance of consumption data is associated. The power management service 402 can be configured to perform any suitable operation (e.g., aggregating the consumption data, calculating power cap values, calculating timing values, determining budget power values, determining whether power capping is necessary (e.g., when the aggregated consumption has breached or is likely to breach the budget power, etc.) based, at least in part, on data representing the power distribution hierarchy 500 of FIG. 5 and / or any suitable configuration data specifying the placement of components within the data center (e.g., indicating which devices distribute power to which devices). By way of example, the configuration data may include data representing the power distribution hierarchy 500.
[0087] FIG. 5 illustrates an example power distribution hierarchy 500 corresponding to the arrangement of components in FIG. 2 , according to at least one embodiment. Power distribution hierarchy 500 may represent the arrangement of any suitable number of components of an electric power system, such as the power distribution infrastructure components discussed in connection with power distribution infrastructure 200 of FIG. 2 . Power distribution hierarchy 500 may include any suitable number of nodes organized according to any suitable number of levels. Each level may include one or more nodes. A root level (e.g., level 5) may include a single root node of power distribution hierarchy 500. Each node of power distribution hierarchy 500 may represent a corresponding component of power distribution infrastructure 200. A set of one or more nodes at a given level may be derived from a particular node corresponding to a higher level of power distribution hierarchy 500. A set of lower-level components represented by a lower-level (e.g., Level 1) node can receive power distributed through a higher-level component represented by a Level 2 node (e.g., a component upstream from the lower-level component), which in turn receives power from a higher-level component represented by a Level 3 node, which in turn receives power from a higher-level component represented by a Level 4 node. In some embodiments, all components of the power system receive power initially distributed through a component corresponding to a Level 5 (e.g., the top level of the power distribution hierarchy 500) node (e.g., node 502, the root node). Node 502 can receive power from a utility power source (e.g., a local electric utility system).
[0088] As shown, power distribution hierarchy 500 includes node 502 at level 5. In some embodiments, node 502 may represent an uninterruptible power supply, such as UPS 202 of FIG. 2. A component corresponding to node 502 may distribute / supply power to a component corresponding to node 504 at level 4. A component (e.g., a lower-level component) that receives power from a higher-level component (a component represented by a node at a higher level than the level of a node representing a lower-level component in power distribution hierarchy 500) may be considered subordinate to the higher-level component. In some embodiments, node 304 may represent a component, such as one of intermediate PDUs 204 of FIG. 2 (e.g., a power distribution panel), that is subordinate to the UPS represented by node 502. The component represented by node 504 may be configured to distribute / supply power to components represented by nodes 506 and 508, respectively.
[0089] Level 3 nodes 506 and 508 may each represent a respective component (e.g., a respective column PDU) in FIG. 2. The components (e.g., column PDU 206, bus bars) corresponding to node 506 may distribute / supply power to the components (e.g., rack PDU 208A in FIG. 2) corresponding to node 510 and the components (e.g., rack PDU 208N in FIG. 2) corresponding to node 512. The components (e.g., rack PDU 208A) corresponding to node 510 may distribute / supply power to the components (e.g., rack PDU 208A) corresponding to node 510 in FIG. 2 (representing servers 212A and 212B in FIG. 2, respectively), which may be monitored / managed by components of servers 212A and 212B, such as power controllers 216A and 216B. The components corresponding to nodes 514 and 516 may be located in the same rack.
[0090] A component corresponding to node 512 (e.g., rack PDU 208A) may distribute / supply power to a component corresponding to level 1 node 518 (e.g., server 214C in FIG. 2 including power controller 216C). Level 1 nodes 514, 516, and 518 may be located in and / or associated with the same column (e.g., column 210 in FIG. 2).
[0091] Returning to node 508 in level 3, node 508 (e.g., a different column PDU, a remote power panel) can distribute / supply power to components corresponding to node 510 (e.g., rack PDU 208A in FIG. 2 ) and components corresponding to node 512 (e.g., rack PDU 208N in FIG. 2 ). Components corresponding to node 520 (e.g., a rack PDU) can distribute / supply power to components corresponding to nodes 522 and 526 in level 1 (e.g., corresponding to respective servers including respective power controllers). Components corresponding to nodes 524 and 526 can be located in the same rack. Components corresponding to node 522 (e.g., another rack PDU) can distribute / supply power to components corresponding to nodes 528 and 530 in level 1 (e.g., corresponding to respective servers including respective power controllers). Components corresponding to nodes 528 and 530 can be located in the same rack.
[0092] The particular number of components (e.g., corresponding to level 1 nodes) that receive distributed power from components at a higher level (e.g., corresponding to level 2 nodes) may differ from the number shown in FIG. 5. The particular number of levels in the power distribution hierarchy may vary depending on the particular arrangement of components used in a given data center. It is contemplated that each non-root level (e.g., levels 1-4 in the example of FIG. 5) may include a different number of nodes representing a different number of components than the number of nodes shown in each non-root level in FIG. 5. The nodes may be arranged in different but similar configurations to those shown in FIG. 5.
[0093] Returning to FIG. 4 , the power management service 402 (e.g., the constraint identification manager 408) can calculate aggregated / cumulative consumption data to identify aggregated / cumulative consumption data for one or more devices at higher levels of the power distribution hierarchy 500. By way of example, the power management service 402 (e.g., the constraint identification manager 408) can utilize consumption data provided by one or more rack PDUs to calculate aggregated / cumulative consumption data for column devices. For example, consumption data corresponding to devices represented by the level 1 nodes of FIG. 5 that share a common set of nodes up to level 3 of the hierarchy. In this case, the common level 3 node represents a column PDU, such as a bus bar. In some embodiments, the consumption data provided to the power management service 402 can include consumption data indicating the power consumption of one or more rack PDUs (e.g., example components corresponding to level 2 of the power distribution hierarchy 500). The consumption data provided by the rack PDUs can be provided according to any suitable frequency, periodicity, or schedule, or in response to a request sent by the power management service 402. The power management service 402 may obtain consumption data associated with any suitable number of devices in FIG. 3 , such as a host 324 and / or a rack of devices corresponding to one or more racks (e.g., rack 319 ), and may aggregate or calculate any suitable consumption data corresponding to any suitable number of components represented by any suitable node and / or level of the power distribution hierarchy 500 .
[0094] The power management service 402 (e.g., constraint identification manager 408) can be configured to calculate one or more power caps (e.g., power cap values for one or more servers to which power is distributed by a higher-level device (e.g., a column-level device in this example) based, at least in part, on the allocated power values (e.g., amounts of budgeted power) for that higher-level component. The higher-level component can correspond to any suitable level (e.g., levels 2-5) of the power distribution hierarchy 500 other than the lowest level (e.g., level 1). As an example, the power management service 402 can calculate power cap values for servers to which power is distributed by a given column device (e.g., bus bar). These calculations may be based on consumption data provided by rack PDUs to which power is distributed by the column devices. In some embodiments, the power management service 402 may store the consumption data for subsequent use. The power management service 402 may utilize historical consumption data when calculating these power cap values. In some embodiments, the power management service 402 may obtain, utilize, and / or train one or more machine learning models from the historical consumption data to identify specific power cap values for one or more hosts 324 (e.g., components corresponding to level 1 of the power distribution hierarchy). These techniques are discussed in more detail with respect to FIG. 12.
[0095] The power management service 402 (e.g., the enforcement manager 410) can calculate timing values for timers (e.g., timers that may be initiated and managed by a PDU, such as the rack PDU 320). The timing values can be calculated based, at least in part, on any suitable combination of the rate of change of the power consumption of the higher-level components, the direction (increase / decrease) of the change in the power consumption of the higher-level components, the power spike tolerance of the higher-level components and / or the system 300 as a whole, or the associated enforcement time for the lower-level devices (e.g., an estimated / known delay between when each of the devices (e.g., the host 324) is instructed to enforce a power cap and when each of the devices (e.g., the host 324) will actively enforce the cap (e.g., the first time that the power consumption at that device is constrained / throttled, or at least a determination is made whether to constrain / throttle).
[0096] The power management service 402 (e.g., implementation manager 410) may communicate the calculated timing values to the PDU 402 at any suitable time. In some embodiments, the power management service 450 may first determine a power cap for a given higher-level component (e.g., a column component) regardless of the consumption occurring for other components at the same level (e.g., other column components). In some embodiments, while the timer is being initialized or elapses, or at any suitable time, the power management service 402 may process consumption data corresponding to other same-level components to determine whether it is more desirable to set a power cap for one or more downstream ones of those same-level components. In some embodiments, the power management service 402 may utilize priorities associated with devices and / or workloads running on those devices to determine power cap values (e.g., priorities identified based at least in part on the host / instance data 316 of FIG. 3 ). The power management service 402 may be configured to prioritize power cap settings for lower priority devices / workloads while leaving higher priority devices / workloads unconstrained. In some embodiments, the power management service 402 can be configured to prioritize power capping of a set of highest consuming devices (e.g., across a row, across multiple rows, etc.) while allowing lower consuming devices to operate unconstrained. In some embodiments, a particular consumption of a device can remain unconstrained even if the device is included in the set of highest consuming devices if the priority associated with that device and / or workload is high (higher than the priority associated with other devices and / or workloads). Thus, the power management service 402 can be configured to prioritize the priority of a device / workload over the power consumption of that particular device.
[0097] In some embodiments, for a given string, power cap values may be initially determined by the power management service 402 (e.g., by the constraint identification manager). These power cap values may be provided to a power controller (e.g., BMC / ILOM 320 or another agent or component of the rack PDU 320) of the rack PDU, which may then distribute the power cap values to a power controller (e.g., BMC / ILOM 328) of the host 324 for storage. The power management service 402 (e.g., constraint identification manager 408) may process consumption data for other string devices in the same string to determine power cap values for devices corresponding to a different string. This may be advantageous because devices managed by another string may not be consuming their budgeted power, leaving some amount of power unused for devices at that string level. In some embodiments, the power management service 402 (e.g., constraint identification manager 408) may be configured to determine whether it may be more advantageous to set power caps for devices in one string while allowing at least some devices in another string to operate unconstrained. Determining the benefit of a set of power cap values may be based, at least in part, on minimizing the estimated impact (e.g., the number of devices that should be power capped), minimizing the priority values associated with devices that should be power capped, and maximizing the number of devices associated with a particular high priority value that will not be power capped. The power management service 402 may utilize a predefined protocol (e.g., a set of rules) to determine whether implementing a power cap value that it has already sent to one PDU is substantially more advantageous / beneficial than a different power cap setting that it has identified based on processing consumption data from multiple PDUs associated with one or more other strings.
[0098] The power management service 402 (e.g., the enforcement manager 410) can be configured to perform enforcement of a set of power cap values determined to be most advantageous (based on a predefined set of rules) to ensure that one or more circuit breakers in the data center do not trip and / or that the aggregate power threshold is enforced. If the current / aggregate power consumption level exceeds the current aggregate power threshold, the power management service 402 can be configured to perform or trigger the execution of an action to bring the current / aggregate power consumption level below the current aggregate power threshold.
[0099] In some embodiments, power cap-based monitoring, identification, and / or constraints may be determined by the power management service 402 directly based on power data 310, account data 312, environmental data 314, host / instance data 316, operational data 318, etc.
[0100] In some embodiments, the power management service 402 can be configured to identify and / or enforce power caps based, at least in part, on instructions and / or constraints or other applicable data provided by the VEPO orchestration service 302. As a non-limiting example, the VEPO orchestration service 302 can filter aggregated data of any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318 based on a reduction action associated with a level (e.g., "set a 10% power cap on all low-priority hosts") to determine an estimated impact of implementing the action, including the number and / or subset of resources (e.g., hosts, instances, workloads, customers) to which the action will apply if implemented. As an example, the VEPO orchestration service 302 can identify numbers and / or identifiers corresponding to all low-priority hosts from any suitable combination of aggregated data of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the identifier and / or instruction (e.g., set a 10% power cap) can be provided to the power management service 402, received through the input / output processing module 404, and implemented by the enforcement manager 410. As another non-limiting example, any suitable portion of metadata associated with a curtailment action (e.g., "set a 10% power cap on all low priority hosts") can be provided to the power management service 402, and the constraint identification manager 408 can be configured to identify the particular host / instance / workload from the metadata associated with the curtailment action. Stated another way, any suitable combination of the VEPO orchestration service 302 and / or the power management service 402 (e.g., the power management service 304) can be configured to identify the estimated impact of a given action.In some embodiments, the power management service 402 (e.g., the constraint identification manager 408) can be configured to identify estimated power reduction values corresponding to individual or aggregate power consumption of one or more devices (e.g., the host 324) taking into account the hosts / instances / workloads / customers expected to be affected if the action (or level) is implemented. If a set of power caps previously provided to a power controller (e.g., the BMC / ILOM 322 and / or the BMC / ILOM 322) is determined to be most advantageous, the power management service 402 can send data to the power controller to cause the power controller to send an indication to the device to initiate power capping based on the previously allocated power caps. Alternatively, if the power management service 402 determines a new set of power caps is more (or most) advantageous from a power management perspective, the power management service 402 can send data to the power controller (e.g., the BMC / ILOM 322) to cancel a timer. In some embodiments, canceling the timer may cause previously allocated power caps to be deleted by instruction from the power controller (e.g., sent in response to canceling the timer), or they may time out by default (e.g., according to a predefined period). The power management service 402 may transmit new power caps in appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was transmitted) with an indication that the power caps should be immediately enforced by the corresponding devices. These PDUs may transmit the power caps to receiving devices with instructions to immediately initiate power capping operations (including, for example, monitoring consumption relative to the power cap values, determining whether the cap is based on the power cap values and the device's current consumption, and capping the device's operation or leaving such operation unconstrained based on the determination).
[0101] In some embodiments, if the timer expires (e.g., the period corresponding to the timing value has passed) on the original PDU and a cancellation has not been received from the power management system 402, the power controller (e.g., BMC / ILOM 322) can automatically instruct the device (e.g., via BMC / ILOM 328) to initiate a power capping operation at the stored power cap value. This technique ensures a fail-safe in case the power management service 402 fails to instruct the PDU or cancel the timer for any reason.
[0102] The techniques described above allow more lower-level devices to operate unconstrained by minimizing the number and / or frequency at which device power consumption is capped. Furthermore, devices downstream of a given device in the hierarchy can be allowed to peak above the device's budgeted power while still ensuring that the device's maximum power capacity (a value higher than the budget value) is not exceeded. This allows the consumption of lower-level devices to allow the consumption levels of higher devices to operate within a buffer of power that conventional systems would leave unused. The present system and described techniques allow for more efficient utilization of the power distribution components of system 400, reducing waste while ensuring that power failures are avoided.
[0103] In some embodiments, a set of power cap values determined to be most advantageous (e.g., based at least in part on any suitable combination of maximum expected reduction and / or minimum estimated impact, etc.) may be selected by the power management service 402, or a particular set of power cap values (previously identified by the power management service 402 and provided to the VEPO orchestration service 302) may be provided from and / or implemented by the power management service 402. In some embodiments, the particular power cap values identified may be identified at least in part based on maximizing power consumption reductions, identifying sufficient power cap-related reductions, individually or collectively, along with other actions corresponding to reduction actions or levels associated with lowering the device's aggregate power consumption value below the current aggregate power threshold.
[0104] 6 is a flow diagram illustrating an example method 600 for managing excessive power consumption, according to at least one embodiment. Method 600 may include more or fewer operations than those shown or described with respect to FIG. 6. The operations may be performed in any suitable order. Any suitable portion of the operations described with respect to power management service 606 may additionally or alternatively be performed by VEPO orchestration service 302 and / or compute service 306 of FIG. 3.
[0105] Method 600 may begin at 610, where consumption data may be received and / or obtained (e.g., upon request) by PDU 604 from devices 602. Devices 602 may be examples of hosts 324 of FIG. 3 (e.g., multiple servers) and may be located within a common server rack (e.g., rack 319). PDU 604 is an example of rack PDU 208A of FIG. 2 and / or rack PDU 320 of FIG. 3. Each of devices 602 may be a device to which PDU 604 distributes power. PDU 608 may be an example of another rack PDU associated with the same column as PDU 604 and therefore receiving power from the same column PDU as PDU 604 (e.g., column PDU 206 of FIG. 2, corresponding to node 506 of FIG. 5). The consumption received from PDU 608 may correspond to the power consumption of devices 612. In some embodiments, PDU 608 and at least some portions of devices 612 correspond to a different column than the column corresponding to devices 602. In some embodiments, any suitable number of instances of consumption data (referred to as “device consumption data”) received by PDU 604 may relate to a single one of devices 602. The same may be true for device consumption data received by PDU 608. In some embodiments, PDU 604 and PDU 608 may aggregate and / or perform calculations from the received device consumption data instances to generate rack consumption data corresponding to each respective PDU. For example, PDU 605 may generate rack consumption data from device consumption data provided by device 602. The rack consumption data may include device consumption data instances received by PDU 604 at 610. Similarly, the rack consumption data provided by PDU 608 may include device consumption instances received from device 612. The rack consumption data generated by PDU 604 and / or PDU 608 may be generated, at least in part, based on the power consumption of PDU 604 and / or PDU 608, respectively.
[0106] At 614, power management service 606 (an example of power management service 402 in FIG. 4 , power management service 304 in FIG. 3 ) may receive and / or obtain (e.g., via a request) rack consumption data from any suitable combination of PDU 604 and PDU 608 (which are also examples of PDU 208 in FIG. 2 ), not necessarily simultaneously. In some embodiments, power management service 606 may continuously receive / obtain rack consumption data from PDU 604 and PDU 608. PDU 604 and PDU 608 may similarly receive and / or obtain (e.g., via a request) device consumption data from devices 602 and 612, respectively (e.g., from their corresponding power controllers, each power controller being an example of BMC / ILOM 328 in FIG. 3 ). PDU 608 may correspond to a PDU associated with the same column as PDU 604.
[0107] At 616, the power management service 606 may determine whether to calculate a power cap value for the device 602 based, at least in part, on maximum and / or budget power amounts associated with another PDU (e.g., a string-level device not shown in FIG. 5 , such as a bus bar) and consumption data received from PDU 604 and / or PDU 608. In some embodiments, the power management service 606 may access these maximum and / or budget power amounts associated with a string-level PDU (e.g., string PDU 206 of FIG. 2 , not shown here) from stored data, or the maximum and / or budget power amounts may be received by the power management service 606 at any suitable time from the PDU to which the maximum and / or budget power amounts are associated (the string-level PDU). The power management service 606 may aggregate rack consumption data from PDU 604 and / or PDU 608 to determine cumulative power consumption values for row-level devices (e.g., total power consumption from devices downstream of the row-level device within the last time window, a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, a period of time required to initiate power capping on one or more of devices 602 and / or 612, etc.). In some embodiments, the power management service 606 may generate additional cumulative power consumption data from at least a portion of the cumulative power consumption data values and / or using historical device consumption data (e.g., cumulative or individual consumption data). For example, the power management service 606 may calculate any suitable combination of a change in consumption rate from the perspective of the row-level device, a direction of change in consumption rate from the perspective of the row-level device, etc. based, at least in part, on historical consumption data (e.g., historical device consumption data and / or historical rack consumption data corresponding to devices 602 and / or 612).
[0108] Alternatively, in some embodiments, the power management service 606 may receive 616 from the VEPO orchestration service 402 any suitable data and / or instructions identifying any suitable portion of the curtailment action (e.g., setting power caps only on low-priority hosts), and / or any suitable constraints, and / or any suitable devices to which the constraints apply. The power management service 606 may identify a subset of hosts for which power caps should be calculated (e.g., in this example, from only low-priority host candidates). For example, an action to set power caps only on low-priority hosts may result in the power management service 606 identifying, from all low-priority hosts (e.g., subset of hosts 324), the number of power caps for those subset of low-priority hosts.
[0109] The power management service 606 may determine that a power cap should be calculated if the cumulative power consumption (e.g., the power consumption corresponding to devices 602 and 612) exceeds the budgeted power amount associated with the column-level PDU. If the cumulative power consumption does not exceed the budgeted power amount associated with the column-level PDU, the power management service 606 may determine that a power cap should not be calculated, and method 600 may end. Alternatively, the power management service may determine that a power cap should be calculated due to the cumulative power consumption of downstream devices (e.g., devices 602 and 612) exceeding the budgeted power amount associated with the column-level PDU, and may proceed to 618.
[0110] In some embodiments, the power management service 606 may additionally or alternatively determine that a power cap value should be calculated at 616 based, at least in part, on providing historical consumption data (e.g., device consumption data and / or rack consumption data corresponding to devices 602 and / or 612) as input to one or more machine learning models. The machine learning models may be trained using any suitable supervised or unsupervised machine learning algorithm to identify, from the historical consumption data provided as input, a likelihood that consumption corresponding to a device associated with the historical consumption data will exceed a budgeted amount (e.g., a budgeted amount of power allocated to a row-level device). The machine learning models may be trained using a corresponding training dataset including historical consumption data instances. Training of these machine learning models is discussed in more detail with respect to FIG. 12 . Using the techniques described above, the power management service 506 may identify that a power cap value should be calculated based on cumulative device / rack consumption levels that exceed the budgeted power associated with row-level devices and / or based on determining, from the output of the machine learning model, that the cumulative device / rack consumption levels likely exceed the budgeted power associated with row-level devices. Determining that a device / rack consumption level is likely to exceed the row-level device's power budget can be determined based on receiving an output from a machine learning model indicating the likelihood (e.g., likely / unlikely) or by comparing an output value (e.g., a percentage, confidence value, etc.) indicating the likelihood of exceeding the row-level device's power budget with a predefined threshold. An output value indicating the likelihood of exceeding the predefined threshold can enable the power management service 606 to determine that power capping is warranted and that a power cap value should be calculated.
[0111] At 618, the power management service 606 can calculate power cap values for any suitable number of devices 602 and / or devices 612. By way of example, the power management service 606 can utilize device consumption data and / or rack consumption data corresponding to each of the devices 602 and 612. In some embodiments, the power management service 606 can determine the difference between the cumulative power consumption (calculated from aggregating the device and / or rack consumption data corresponding to devices 602 and 612) of a column of devices (e.g., device 602 and any of the devices 612 in the same column) and the budgeted power associated with the column-level devices. This difference can be used to identify the amount by which power consumption should be constrained at the devices associated with the column. The power management service 606 can determine power cap values for any suitable combination of devices 602 and / or 612 for devices in the same column based on identifying power cap values that, when implemented by the devices, would reduce power consumption to a value less than the budgeted power associated with the column-level devices.
[0112] In some embodiments, the power management service 606 can determine specific power cap values for any suitable combination of devices 602 and 612 corresponding to the same column based, at least in part, on consumption data and / or priority values associated with those devices and / or workloads associated with the power consumption of those devices. In some embodiments, these priority values may be provided as part of device consumption data provided by the devices whose consumption data are associated, or the priority values may be ascertained, at least in part, based on the type of device, the type of workload, etc., obtained from the consumption data or any other suitable source (e.g., from separate data accessible to the power management service 606). Using the priorities associated with each device or workload, the power management service 606 can calculate power cap values such that it prioritizes setting power caps on devices associated with lower priority workloads over setting power caps on devices with higher priority workloads. In some embodiments, the power management service 606 may calculate power cap values based, at least in part, on prioritizing setting power caps on devices consuming at higher rates (e.g., the set of highest consuming devices) over setting power caps on other devices consuming at lower rates. In some embodiments, the power management service 606 may calculate power cap values based, at least in part, on a combination of factors including the consumption value for each device and the priority associated with the device or the workload being executed by the device. The power management service 606 may identify power cap values for devices that entirely avoid capping high-consuming and / or high-priority devices or that cap those devices to a lesser extent than devices that consume less power and / or are associated with a lower priority.
[0113] At 620, the power management service 620 may calculate timing values corresponding to durations for timers that the PDU (e.g., PDU 604) should initialize and manage. The timing values may be calculated based, at least in part, on the rate of change of power consumption from the perspective of the column-level devices, the direction of change of power consumption from the perspective of the column-level devices, or any suitable combination of the time during which power capping should be performed at each of the devices to which the power cap calculated at 618 is associated. As an example, the power management service 620 may determine a timing value that is greater than timing values for smaller increases in power consumption from the perspective of the column-level devices based on determining a relatively large increase in power consumption from the perspective of the column-level devices. Thus, a larger increase in the power consumption rate may result in a smaller timing value (corresponding to a shorter timer), while a smaller increase in the power consumption rate may result in a larger timing value (corresponding to a longer timer). Similarly, the calculated first timing value may be smaller than the calculated second timing value if the time during which power capping should be performed at each of the devices to which the power cap is associated is shorter relative to the calculation of the first timing value. Thus, the faster a device for which power capping is intended can implement the power cap, the shorter the timing value may be. In some embodiments, a timer is not calculated for power caps determined via data provided or directed by the VEPO orchestration service 302.
[0114] At 622, the calculated power cap value calculated at 618 and / or the timing value calculated at 620 may be transmitted to the PDU 604. Although not shown, corresponding power cap values and / or timing values for any of the devices 612 in the same row may be transmitted at 620. Although not shown, in some embodiments, the transmission at 622 may be based at least in part on instructions received from the VEPO orchestration service 302 to enforce / implement the power cap value calculated at 618.
[0115] At 624, if a timing value was provided at 622, the PDU 604 may use the timing value to initialize a timer having a duration corresponding to the timing value. As noted above, in instances where the power management service 606 is directed by the VEPO orchestration service 302, the timing value need not be provided at 622 or stored at 624.
[0116] At 626, the PDU 604 may transmit the power cap value (if received at 622) to the device 602. The device 602 may store the received power cap value at 628. In some embodiments, the power cap value provided at 626 may include an indicator that enforcement should not begin, or may not provide an indicator, and the device 602 may not initiate power capping by default.
[0117] At 630, the power management service 606 may perform a high-level analysis of device / rack consumption data received from any of the devices 612 corresponding to a different row than the row to which the device 602 corresponds. As part of this process, the power management service 606 may identify unused power associated with other row-level devices. If unused power exists, the power management service 606 may calculate a new set of power cap values based, at least in part, on consumption data corresponding to at least one other row of devices. In some embodiments, any suitable power cap values identified by the power management service 606 may be provided to the VEPO orchestration service 302 at any suitable time. As discussed above, the new set of power cap values may be advantageous because devices managed by another row-level device may not be consuming their budgeted power, leaving unused power for that row-level device. In some embodiments, the power management service 606 can be configured to determine whether it may be more advantageous to allow at least some of the devices 602 (and possibly some of the devices 612 corresponding to the same column as device 602) to operate unconstrained while setting caps on devices in another column. The benefits of each power capping approach can be calculated based, at least in part, on minimizing the number of devices for which power caps should be set, minimizing the priority values associated with the devices for which power caps should be set, maximizing the number of devices associated with a particular high priority value for which no power caps should be set, etc. The power management service 606 can utilize a predefined scheme or set of rules to determine whether implementing a power cap value it has already determined for one column (e.g., corresponding to the power cap value transmitted in 622) is effectively more advantageous than a different power cap setting identified based on processing consumption data from multiple columns.In some embodiments, each set of power cap values identified by the power management service 606 may be beneficial, and the most advantageous set of power caps may be selected by the power management service 606 (e.g., according to a set of rules, according to user input selecting a particular response level, etc.).
[0118] The power management service 606 can be configured to perform the implementation of the set of power cap values (or the set of power cap values instructed by the VEPO orchestration service 302) that is determined to be more (or most) favorable to ensure that one or more circuit breakers in the data center do not trip and / or that aggregate power thresholds are implemented. If the set of power cap values previously provided to the PDU 604 is determined to be less favorable than the set of power cap values calculated in 630, the power management service 606 can immediately send data to the PDU 604 to cause the PDU 604 to send an indication to the device 602 to immediately begin power capping based on the previously allocated power cap. Alternatively, the power management service 606 may take no further action, allowing a timer in the PDU 604 to expire. Any suitable action discussed as being performed by the power management service 606 based on a determination or identification made by the power management service 606 may alternatively be triggered by direction of the VEPO orchestration service 302.
[0119] If the power management service 606 determines that the new set of power caps is more (or most) advantageous from a power management perspective, the power management service 606 may send data to the power controller (e.g., the BMC / ILOM 322 in FIG. 3 ) to cancel the timer. The previously allocated power caps may be deleted by instruction from the power controller (e.g., the BMC / ILOM 322) (e.g., sent in response to canceling the timer), or the power caps may time out by default according to a predefined period. The power management service 606 may send the new power caps in appropriate PDUs (which may include or exclude the same PDUs in which the previous set of power caps was sent) with an indication that the power caps should be immediately enforced by the corresponding devices. These PDUs may send the power caps to the receiving devices with instructions to immediately initiate power capping operations (including, for example, monitoring consumption with respect to the power cap values, determining whether the caps are based on the power cap values and the device's current consumption, and capping or leaving such operations unconstrained based on the determination).
[0120] If it is determined that the power cap value calculated at 630 is more favorable from a power management perspective than the power cap value calculated at 618 and ultimately stored at the device 602 at 628, the method 600 proceeds to 632, where the power cap value calculated at 630 may be transmitted to a PDU 608 (e.g., any of the PDUs 608 that manage the device to which the power cap value is associated). In some embodiments, the power management service 606 may provide an indication that the power cap value transmitted at 632 should be implemented immediately. The operation described at 632 may be triggered by the power management service 606 based on a determination / identification made by the power management service 606 or in response to receiving an instruction from the VEPO orchestration service 302.
[0121] In response to receiving the power cap values and indications at 632, the PDU 608 that distributes power to the devices to which the power cap values are associated can transmit the power cap values and indications to the devices to which the power cap values are associated. At 636, the receiving device among the devices 612 can perform a power capping operation without delay based on receiving the indication, which includes 1) determining whether to limit / constrain operation based on the current consumption of a given device as compared to the power cap value provided for that device, and 2) limiting / constraining or not limiting / constraining the power consumption of that device within the range.
[0122] In some embodiments, the power management service 606 may transmit data at 622 canceling the timer and / or power cap value transmitted based, at least in part, on determining that the power cap value calculated at 630 is more favorable (or most favorable) than that calculated at 618. In some embodiments, the PDU 604 may be configured to cancel the timer at 640. In some embodiments, the PDU 604 may transmit data at 642 that causes the device 602 to delete the power cap value from memory at 644.
[0123] At 646, in the situation where a more favorable set of power cap values is not found as described above or the power management service 606 has not sent a cancellation as described above in connection with 638, the PDU 604 may determine that the timer started at 624 has expired (e.g., the duration corresponding to the timing value provided at 622 has elapsed). As a result, the PDU 604 may send an indication at 648 to the device 602 instructing the device 602 to initiate power capping based on the power cap values stored at 628 in the device 602. Upon receiving this indication at 648, the device 602 may initiate power capping at 650 according to the power cap values stored at 628. In some embodiments, the VEPO orchestration service 302 may be configured to control the operations of 646-650, and the operations of 646-650 may not be performed by the power management service 606 unless directed to do so by the VEPO orchestration service 302.
[0124] In some embodiments, the PDU 604 may be configured to cancel the timer at any suitable time after the timer is started at 624 if the PDU 604 determines that the device 602 has reduced its power consumption. In some embodiments, this may cause the PDU 604 to send data to the device 602 to cause the device 602 to discard the previously stored power cap value from local memory.
[0125] In some embodiments, the power cap value implemented in any suitable device may be allowed to time out or may be replaced at any suitable time with a power cap value calculated by the power management service 606. In some embodiments, the power management service 606 may send a power cap cancellation or replacement to any suitable device via the corresponding rack PDU. If a power cap cancellation is received, the device may delete the previously stored power cap value, allowing the device to resume unconstrained operation. If a replacement power cap value is received, the device may store the new power cap value and implement the new power cap value immediately or when instructed by its rack PDU. In some embodiments, being instructed by its PDU to implement a new power cap value may cause a device (e.g., power controller 446) of the devices to replace the power cap value previously utilized to implement the power cap setting with the new power cap value that the device is instructed to implement.
[0126] Any suitable operations of method 600 may be performed continuously or at any suitable time to manage excess consumption (e.g., a situation in which the consumption of a server corresponding to a column-level device exceeds the column-level device's budgeted power) to enable more efficient use of previously unused power while avoiding power outages due to circuit breaker tripping. It should be understood that similar operations may be performed by power management service 606 with respect to any suitable level of power distribution hierarchy 500. While examples have been provided with respect to monitoring power consumption corresponding to column-level devices (e.g., devices represented by level 3 nodes of power distribution hierarchy 500), similar operations may be performed with respect to any suitable level of higher-level devices (e.g., components corresponding to any of levels 2-5 of power distribution hierarchy 500). As discussed above, any suitable function / operation of power management service 606 may be triggered, at least in part, based on a determination / determination performed by power management service 606, or any suitable function / operation of power management service 606 may be triggered via direction by VEPO orchestration service 302.
[0127] 7 illustrates an example architecture of a VEPO orchestration service 702 (an example of VEPO orchestration service 302 of FIG. 3) configured to orchestrate management of host power consumption relative to aggregate power thresholds, according to at least one embodiment. VEPO orchestration service 702 may include an input / output processing manager 704 that may be configured to receive and / or transmit any suitable data between VEPO orchestration service 702 and any other suitable device of FIG. 3. As shown, power management service 402 may include a demand manager 710, an impact identification manager 706, and an implementation manager 708, although more or fewer computing components or subroutines may likewise be utilized.
[0128] Demand manager 710 can be configured to monitor and / or modify aggregate power thresholds associated with a physical environment (e.g., data center 102, room 110A, column 108A in FIG. 1, etc.). Monitoring may include monitoring any suitable combination of power data 310, environmental data 314, host / instance data 316, and / or operational data 318 in FIG. 3. The modification of the aggregate power thresholds can be performed, and the amount by which the aggregate power thresholds can be modified can be determined, at least in part, based on a number of triggering events. These triggering events may include, but are not limited to, 1) receiving a request from a governmental agency (e.g., a local power authority) requesting a power reduction indefinitely or for a period of time; 2) determining current and / or predicted environmental conditions (e.g., current ambient or external temperature, predicted ambient or external temperature); 3) determining current or predicted aggregate power consumption corresponding to devices in the physical environment; and 4) determining the current and / or predicted operational status of one or more temperature control units in the physical environment. The amount by which the aggregate power threshold can or should be changed can be determined, at least in part, based on a number of factors, including, but not limited to, the difference between ambient temperature and external temperature, the amount of power consumption attributable to or associated with a particular component (e.g., a particular temperature control unit), an amount or percentage specified in a request received from a government agency, or the difference between a current or predicted aggregate power consumption value and a current or predicted aggregate power threshold.
[0129] The impact identification manager 706 can be configured to generate (directly or by invoking functionality of another service or process, such as the compute service 306 and / or the power management service 304) an estimated impact corresponding to the response level and / or one or more instances of a reduction action corresponding to the response level. The VEPO orchestration service 702 can obtain configuration data 712 that specifies a set of multiple VEPO levels and their corresponding one or more reduction actions.
[0130] 8 is a table 800 illustrating an example set of response levels and a respective set of reduction actions for each response level, according to at least one embodiment. The data in table 800 can be provided in different ways, such as via a mapping or embedded within the program code of the VEPO orchestration service 702. Each response level in the set of response levels can be associated with a corresponding set of reduction actions that can be performed on multiple resources (e.g., hosts, instances, workloads, etc.) of the data center, with each corresponding set of reduction actions associated with implementing a different reduction in the aggregate power consumption of the data center. In some embodiments, the multiple response levels indicate increasing severity when the reduction actions are performed to reduce the aggregate power consumption of the data center.
[0131] By way of example, table 800 includes four VEPO response levels, although any suitable number of response levels may be utilized. VEPO response level 1 may be associated with one curtailment action. Specifically, the curtailment action specifies that all free and / or idle hosts and hypervisors that can be fully evacuated (e.g., via migration) should be shut down (e.g., powered off).
[0132] In some embodiments, VEPO Response Level 1 may include stopping or preventing any launch or assignment of instances and / or workloads to the affected hosts. In some embodiments, VEPO Response Level 1 may include mitigation actions including migrating instances and / or workloads from one host to another, or from one instance to another (in the case of workload migration).
[0133] VEPO Response Level 2 may be associated with multiple actions, such as any / all of the mitigation actions described above in connection with VEPO Response Level 1, as well as one or more additional mitigation actions, such as an additional mitigation action corresponding to powering off hypervisors hosting instances and / or workloads associated with Category 1 customers (e.g., free tier customers, free trial customers, etc.).
[0134] VEPO Response Level 3 may be associated with multiple actions, such as any / all of the reduction actions associated with VEPO Response Levels 1 and / or 2, as well as one or more additional reduction actions. By way of example, VEPO Response Level 3 may include an additional reduction action corresponding to powering off hypervisors (e.g., after evacuating higher priority VMs) hosting Category 2 bare metal instances (e.g., Pay-As-You-Go and / or Enterprise BMs) and Category 3 virtual machines (e.g., VMs associated with the Pay-As-You-Go and Enterprise categories). In some embodiments, the additional reduction action may specify excluding from consideration BMs and / or VMs associated with Priority Level 1 customers and / or cloud services (e.g., critical customer and cloud provider services).
[0135] VEPO Response Level 4 may be associated with multiple actions, such as any / all of the mitigation actions associated with VEPO Response Levels 1, 2, and / or 3, as well as one or more additional mitigation actions. By way of example, VEPO Response Level 4 may include an additional mitigation action corresponding to powering off all compute hosts in a customer enclave except those that remain in a critical state and / or are deemed critical for recovery.
[0136] The number of response levels and / or specific reduction actions illustrated in FIG. 8 are not intended to limit the scope of the present disclosure. Any suitable number of response levels may be employed, each corresponding to any suitable number of reduction actions. The reduction actions may include or exclude hosts, instances, workloads, and / or customers based on any suitable attributes associated with them, such as based on priority, category, status, and / or future recovery needs. In some embodiments, the hosts, instances, and workloads considered as candidates for applying these reduction actions may be based on limited ranges, as described in more detail with respect to FIG. 9.
[0137] 8 include some degree of overlapping reduction actions, but this is not a requirement. In some embodiments, the reduction actions at each level may include a lesser or greater degree of overlap, including no overlap, resulting in completely unique actions associated with each level.
[0138] 9 is a diagram 900 illustrating example ranges for components affected by one or more response levels utilized by a VEPO orchestration service (e.g., VEPO orchestration service 702 of FIG. 7 ) to orchestrate power consumption constraints, migration, suspension, and / or power shutdown tasks, according to at least one embodiment. The diagram shows a user enclave 902 and a cloud service enclave 904. The user enclave 902 may include compute instances running workloads corresponding to services 906 (or any suitable workloads) operating within a customer overlay (e.g., overlay 908). The instances and / or workloads corresponding to services 906 may be associated with users 910.
[0139] In contrast, cloud service enclave 904 may include compute instances running workloads corresponding to services 912 (e.g., control and / or data plane services such as those discussed in connection with the service tenancies of Figures 16-19), and / or resources such as object storage and virtual cloud networking.
[0140] In some embodiments, a host 324 may host any suitable combination of instances and / or workloads corresponding to services 906 associated with a user (e.g., a customer) and / or services 912 and / or resources associated with a cloud provider, but only hosts, instances, and / or workloads corresponding to customer enclaves (e.g., compute resources) may be considered as candidates to which curtailment actions may be applied. In some embodiments, application of curtailment actions to services and / or resources associated with a cloud provider may be completely avoided (at least temporarily) because those services and / or resources are needed for later recovery.
[0141] Returning now to FIG. 7 , the functionality described in connection with power management service 402 may, in some embodiments, be performed by one or more virtual machines implemented in a hosted computing environment (e.g., in hosts 324 corresponding to any suitable number of servers 104 of FIG. 1 ). A hosted computing environment may include one or more rapidly provisioned and released computing resources, which may include computing, networking, and / or storage devices. A hosted computing environment may also be referred to as a cloud computing environment. A number of example cloud computing environments are provided and described in more detail below with respect to FIGS. 16-19 .
[0142] The VEPO orchestration service 702, through associated functionality or functionality provided by other systems and / or services, can be configured to manage power consumption values corresponding to hosts / instances / workloads in the physical environment. Managing power consumption values can include setting power caps, pausing workloads, instances, hosts, shutting down workloads / instances / hosts, migrating instances from one host to another, or migrating workloads from one instance and / or host to another, etc.
[0143] Impact identification manager 706 can be configured to utilize configuration data 712 (e.g., table 800 in FIG. 8 ) corresponding to the specification of numerous response levels and their corresponding reduction actions to determine the estimated impact of applying a given level of reduction action to hosts / instances / workloads (collectively, “resources”) in a physical environment. The estimated impact of applying a given action and / or level may depend on the current state of the resources at the time of execution. In some embodiments, determining the estimated impact of a given reduction action may include identifying a set of resources and / or customers to which the action, if implemented, is applicable. Thus, estimated impact is intended to refer to the scope or applicability of a given action or level. Determining the scope and / or applicability of a given action or level may include identifying the specific resources and / or customers affected by implementing the action or level and / or identifying any suitable attributes associated with the applicable resources and / or customers. As an overly simplistic example, impact identification manager 706 may identify that if an action to shut down all idle hosts is taken, X number of idle hosts among hosts 342 will be affected, that the hosts are associated with a particular customer or number of customers, and / or that the hosts and / or customers are associated with other attributes such as category (e.g., free tier) and priority (e.g., high priority). The curtailment actions may provide increasingly more severe actions and / or increasingly widespread impacts. As an example of increasingly more severe actions, a lower level may specify an action to set a power cap for a particular host associated with a particular attribute, while a higher level may specify to shut down that host completely.An example of increasingly broader impact may include a lower level affecting a small number of hosts (e.g., 2, 5, 10) corresponding to a single, low-priority customer, while a higher level may affect a larger number of hosts (e.g., 100) corresponding to the same or larger number of customers associated with a broader priority level, etc. Broader impact thus refers to the number, scope, or breadth of applicability of a given action to hosts, instances, workloads, or customers.
[0144] Impact identification manager 706 may be configured to estimate a likely impact and / or estimated power reduction for one or more of the levels based, at least in part, on configuration data specifying the levels and corresponding actions and any suitable combination of power data 310, account data 312, or host / instance data, etc. For one or more response levels or one or more reduction actions corresponding to a given level, an estimated impact (e.g., what hosts, instances, workloads, customers are likely to be affected, what attributes are associated with the affected hosts, instances, workloads, and / or customers, how many hosts / instances / workloads / customers are affected, etc.) may be identified. In some embodiments, the estimated impact of a given action may be aggregated with the estimated impact for all actions at a given level to determine the estimated impact for a given level.
[0145] In some embodiments, impact identification manager 706 can be utilized to estimate impacts and / or estimate power reductions for all levels. In some embodiments, determining estimated power reductions can be an incremental process that includes identifying applicable resources (e.g., hosts, instances, and / or workloads), determining the current power consumption of the applicable resources, estimating the reduction for each response if the action is performed / implemented, and aggregating the individual estimated reductions for the individual responses to identify an estimated aggregate power reduction corresponding to the given action. This process can be repeated for each action at a given level, and the resulting estimated power consumption reductions for each action can be aggregated to determine the estimated aggregate power consumption reduction for the given level. The functionality of impact identification manager 706 can be invoked by implementation manager 708.
[0146] In some embodiments, the functionality of the implementation manager 708 can be invoked by the demand manager 710 based, at least in part, on monitoring or determining that a change to the aggregate power threshold is needed. The monitoring can include monitoring any suitable combination of the power data 310, the environmental data 314, the host / instance data 316, and / or the operational data 318 of FIG. 3 . The change to the aggregate power threshold can be implemented, and the amount by which the aggregate power threshold can be changed can be determined, at least in part, based on a number of triggering events. These triggering events can include, but are not limited to, 1) receiving a request from a governmental agency (e.g., a local power authority) requesting a power reduction indefinitely or for a period of time, 2) determining current and / or predicted environmental conditions (e.g., current ambient or external temperature, predicted ambient or external temperature), 3) determining current or predicted aggregate power consumption corresponding to devices in the physical environment, and 4) determining the current and / or predicted operational status of one or more temperature control units in the physical environment. The amount by which the aggregate power threshold can or should be changed can be determined, at least in part, based on a number of factors, including, but not limited to, the difference between ambient temperature and external temperature, the amount of power consumption attributable to or associated with a particular component (e.g., a particular temperature control unit), an amount or percentage specified in a request received from a government agency, or the difference between a current or predicted aggregate power consumption value and a current or predicted aggregate power threshold.
[0147] In some embodiments, upon determining the amount of change in the aggregate power threshold, demand manager 710 can invoke functionality of enforcement manager 708 to determine what response level is appropriate for the given change. As a non-limiting example, enforcement manager 708 can incrementally invoke functionality of impact identification manager 706 (e.g., via function call, etc.) to determine an estimated impact and / or estimated power reduction corresponding to a given level (e.g., VEPO response level 1 in FIG. 8 ). The resulting estimated impact and / or estimated power reduction calculated by the impact identification manager can be returned to enforcement manager 708. In some embodiments, enforcement manager 708 can obtain and / or implement a predefined scheme for determining the suitability of selecting one response level over another. As an example, environmental manager 708 may be configured to compare the estimated power reduction of a given level with a current aggregate power threshold identified by demand manager 710 to determine whether the estimated power consumption of the given level (e.g., VEPO response level 1) is sufficient (e.g., exceeds the change required to reduce the current aggregate power consumption (identified and provided by power management service 402) to a value below the aggregate power threshold identified by demand manager 710). If sufficient, in some embodiments, enforcement manager 708 may evaluate the estimated impact according to a predefined scheme to determine whether the estimated impact is sufficient and / or acceptable. As a non-limiting example, even if the estimated power reduction is sufficient to reduce the current aggregate power consumption below the current aggregate power threshold, the response level may be rejected by enforcement manager 708 if the estimated impact exceeds the conditions of the predefined scheme (e.g., the number of affected hosts exceeds a threshold specified in the predefined scheme).If either the estimated power consumption and / or estimated impact of a response level does not pass the requirements / conditions of the predefined scheme, the Implementation Manager 708 may be configured to invoke the Impact Identification Manager 706 to determine an estimated impact and / or estimated power reduction corresponding to a higher VEPO response level (e.g., VEPO Response Level 2).
[0148] The impact and / or power reduction estimates may be invoked by the implementation manager 708 and determined by the impact identification manager 706, one level at a time, until the implementation manager 708 finds a first level that meets the requirements / conditions of the predefined scheme with respect to the estimated impact and / or estimated power reduction. The level determined to be the lowest level (e.g., a level occurring higher in table 800) at which all impact conditions of the predefined scheme are met may be referred to as the “least impact resource level.” The level determined to be the lowest level (e.g., a level occurring higher in table 800) at which all power reduction conditions of the predefined scheme meet (e.g., below) the current aggregate power threshold may be referred to as the “least reduced resource level.” The level determined to be the lowest response level at which all conditions with respect to the estimated impact and estimated power reduction are met may be referred to as the “first sufficient response level.” The implementation manager 708 can be configured to select the response level that is first identified as the least impactful level, the least reduced level, or the first sufficient level according to a predefined scheme.
[0149] In some embodiments, enforcement manager 708 may cause the estimated impact and / or estimated power reduction for all levels (or some subset, such as the first five, next five, etc.) to be determined by impact identification manager 706. Enforcement manager 708 may be configured to select a particular response level based, at least in part, on comparing the estimated impact and / or estimated power reduction of each level to determine the one with the least impact (e.g., relatively the smallest impact), the least reduction (e.g., resulting in the smallest amount of power reduction sufficient to reduce the current aggregate power consumption below the current aggregate power threshold), or the one with the least impact and the least reduction (e.g., based at least in part on a weighting algorithm).
[0150] Any suitable data, such as estimated impacts and / or estimated reductions calculated by impact identification manager 706 (possibly by invoking functionality of computer services 306 and / or power management services 304 of FIG. 3), and / or current and / or forecast values based on those estimates, may be presented via interface 11, which will be discussed in more detail below.
[0151] As discussed herein, any suitable determination and / or identification of the demand manager 710, impact identification manager 706, and / or implementation manager 708 (or any suitable service or subroutine called by them) may rely on current data, historical data, and / or forecast data. As a non-limiting example, the demand manager 710 may determine that a change is needed and / or the amount by which the current aggregate power threshold should be changed based on the current operational status of the temperature control units 112 of FIG. 1 or the predicted operational status of those units during a future time period. This predicted status may be identified, at least in part, based on historical data (e.g., indicating one or more partial or complete failures or shutdowns of the temperature control units) and / or based on output provided by a machine learning model (e.g., a model trained in the manner described in connection with FIG. 12). In some embodiments, the amount of change determined by the demand manager 710 for the aggregate power threshold may be based on predefined formulas and / or tables. For example, formulas and / or tables may be available from which the capacity associated with the temperature control units can be determined and / or calculated. If used, the table may indicate the amount of power consumption attributable to or associated with a temperature control unit. Thus, a failure (e.g., partial and / or complete) of that temperature control unit may be associated with the entire, or some proportional, portion of the power consumption attributable to or associated with that temperature control unit (e.g., the amount of power consumption in the form of heat that the temperature control unit is configured to handle). Thus, the demand manager 710 may determine that a total failure of that temperature control unit requires a reduction in the aggregate power threshold by an amount equal to the overall power consumption attributable to or associated with the temperature control unit, as specified by the table.
[0152] In some embodiments, the implementation manager 708 can be configured to present, recommend, or automatically select a given response level from a set of possible response levels based, at least in part, on any suitable combination of estimated impact and / or estimated power reduction likely to occur if a reduction action corresponding to the given level is implemented (e.g., realized). Thus, in some embodiments, a particular level can be recommended and / or selected by the VEPO orchestration service 702 based on any suitable combination of: 1) determining a level that, if implemented, is likely to result in a sufficient, but not excessive, reduction in power consumption in resources of the physical environment (e.g., the smallest amount of power consumption reduction sufficient to reduce the current power consumption value below the current aggregate power threshold); 2) determining a level that includes the least severe set of actions; or 3) identifying the least impactful level or action (e.g., the set of affected resources and / or customers with the smallest number of potentially affected resources / customers, a priority or category indicating the lowest overall degree of importance, etc.). Determining the least severe or least impactful level or action may be determined independently of the estimated power consumption reduction likely to result from implementing the level and / or action, or determining the least severe and / or least impactful level / action may be determined from one or more levels that, if implemented given the current situation, are estimated to result in a power consumption reduction sufficient to bring the power consumption value down to an aggregate value that falls within the current aggregate power threshold. In some embodiments, implementation manager 708 may cause any suitable data corresponding to current, predicted, or estimated power consumption and / or impact to be presented in the interface of FIG. 11. In some embodiments, selection of the particular level to implement may be made by user input provided in the user interface.Thus, in some embodiments, the implementation manager 708 may refrain from performing operations to implement a level reduction action until the user identifies and / or confirms the level to be implemented via user input provided in the user interface.
[0153] In some embodiments, performing these actions may include instructing the power management service 402 and / or the computer service 306 to perform the actions. By way of example, the enforcement manager 708 may provide identifiers for affected resources (e.g., hosts, instances, workloads, etc.) and reduction action data specifying an estimated power reduction and / or action to take for the power management service 402 (e.g., set a power cap for a particular resource, set a power cap for all low-priority hosts / instances / workloads, set a power cap by a particular amount (e.g., 10%, by at least x amount, etc.) for a particular or all resources that meet a particular set of attributes), which may be configured to execute / enforce the corresponding power cap as instructed by the enforcement manager 708, as described above. Similarly, the compute service 306 may be provided with identifiers for the affected resources (e.g., hosts, instances, workloads, etc.) as well as curtailment action data specifying the action to be taken by the compute service 306 (e.g., shut down all idle hosts, shut down these specific hosts, suspend all low priority workloads, migrate and / or compress all low priority workloads to maximize the number of idle / free hosts, etc.), which may be configured to enforce / enforce the corresponding power cap as directed by the enforcement manager 708, as described above.
[0154] The demand manager 710 may be further configured to identify that a previous trigger that resulted in the response level being implemented has ceased to occur and may modify the aggregate power threshold based on identifying that the trigger condition has ceased to occur. For example, the demand manager 710 may identify (e.g., by monitoring the operational data 318) that a previously failed temperature control unit is no longer operational. As a result, the demand manager 710 may modify (e.g., increase) the aggregate power threshold by the amount that was attributed to the temperature control unit when it was fully operational. Following this modification, the implementation manager 706 may be configured to perform any suitable operation to reverse any curtailment actions previously taken, to the extent possible, depending on which response level was selected in response to the original trigger condition. If a power cap was utilized, the implementation manager 706 may instruct the power management service 402 to remove the previously applied power cap. If a workload and / or instance migration was performed, the implementation manager 706 may instruct the compute service 306 to attempt to migrate those resources back to the original host and / or instance that hosted the instance and / or workload, respectively. If resources are suspended and / or shut off, the enforcement manager 706 can instruct the compute service 306 to perform operations to resume instances and / or workloads and / or send signals to power on powered-down hosts. In this manner, the enforcement manager 706 can perform any suitable operations to recover from curtailment actions taken based on a previous selection of a given response level.
[0155] Figure 10 is a block diagram illustrating an example use case in which multiple response levels are applied based on a set of host current states, according to at least one embodiment. The example provided in Figure 10 assumes that the VEPO response levels and corresponding actions discussed in connection with Figure 8 are the current configuration data utilized by the VEPO Orchestration Service 702 of Figure 7. As shown, the following conditions are assumed:
[0156] Server 1004A and server 1004C are free. Server 1004D hosts resources (eg, instances) associated with customers in category level 1.
[0157] Server 1004I hosts a BM instance associated with a customer associated with priority level 2.
[0158] Server 1004J hosts the BM for a customer associated with priority level 3.
[0159] Server 1004K hosts VMs for a customer associated with priority level 1.
[0160] Server 1004B hosts one or more instances that are considered critical to recovery. In some embodiments, these instances may be associated with services 912 offered by a cloud provider.
[0161] As described above, the VEPO orchestration service 702 can incrementally or in parallel identify estimated impacts and / or estimated power reductions corresponding to each response level. The VEPO orchestration service 702 can identify (e.g., through functionality provided by the power management service 304 and / or the compute service 306) applicable responses to a given level (e.g., reduction actions corresponding to a given response level). By way of example, in the current example, the VEPO orchestration service 702 (via the power management service 304) can identify that VEPO Response Level 1 can be estimated to affect two servers, server 1004A and server 1004C, based on identifying servers 1004A and 1004C as a set of resources to which the redundancy actions of VEPO Response Level 1 apply (e.g., from non-cloud provider-based resources, such as resources associated with a user enclave, such as user enclave 902). In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount by which an estimated power reduction for a given level, if implemented / realized, would affect the aggregate power consumption of resources in the physical environment (e.g., a 10% reduction with no impact to customers, free BM capacity reduced to 0%, free CM capacity severely curtailed, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0162] The VEPO orchestration service 702 (via the power management service 304) can determine that VEPO Response Level 2 can be estimated to impact server 1004A and server 1004C based on identifying that server 1004D hosts resources to which VEPO Response Level 3 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with a user enclave, such as user enclave 902), as well as including VEPO Response Level 1 reduction actions. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount to which the estimated power reductions for VEPO Response Level 2, if implemented / implemented, would affect the aggregate power consumption of the resources of the physical environment (e.g., a 20% reduction with minimal impact to customers, up to 400 low-priority VMs affected, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO Response Levels.
[0163] The VEPO orchestration service 702 (via the power management service 304) can determine that VEPO Response Level 3 can be presumed to affect servers 1004A, 1004C, and 1004D based on identifying that the BM hosted by server 1004I and the VM hosted by server 1004J are resources to which VEPO Response Level 3 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with user enclaves, such as user enclave 902), as well as including VEPO Response Level 3 reduction actions. Server 1004K, and / or resources hosted by server 1004K can be excluded based on meeting the exclusions associated with the response level being evaluated. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount by which the estimated power reductions for VEPO response level 3, if implemented / realized, would affect the aggregate power consumption of resources in the physical environment (e.g., medium impact - 40% reduction, up to 200 category 1 or 2 BMs / VMs affected, etc.). Any suitable portion of this information can be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0164] In addition to identifying that the BM hosted by server 1004I and the VM hosted by server 1004J are resources to which VEPO Response Level 4 redundancy actions apply (e.g., from non-cloud provider-based resources, such as resources associated with user enclaves, such as user enclave 902), the VEPO orchestration service 702 (via power management service 304) can determine that VEPO Response 4 can be estimated to affect all servers (e.g., the set of servers including servers 1004A and 1004C-1004L) except for server 1004B based on including the reduction actions of VEPO Response Levels 1-3. In some embodiments, the VEPO orchestration service 702 (e.g., impact identification manager 706) can determine the amount to which the estimated power reduction for VEPO Response Level 4, if executed / realized, would affect the aggregate power consumption of the resources of the physical environment (e.g., impact is significant, 80% reduction, region down, etc.). Any suitable portion of this information may be presented in the user interface 1100, with or without the presentation of similar data associated with other VEPO response levels.
[0165] FIG. 11 is a schematic diagram of an example user interface 1100, according to at least one embodiment. User interface 1100 may include any suitable number and types of graphical interface elements (e.g., drop-down boxes 1102-1108) for selecting and / or displaying VEPO data corresponding to a given region, availability domain (AD), building, and / or room, respectively. Other interface elements, such as radio buttons and / or edit boxes, may also be employed. The VEPO data presented in user interface 1100 may correspond to the region, AD, building, and room selected via graphical interface elements 1102-1108 (also referred to as "options" 1102-1108). As shown, VEPO data 1109 is displayed according to the values selected via options 1102-1108.
[0166] As described above, any suitable information utilized by the VEPO orchestration service 702, the power management service 304, and / or the compute service 306 may be presented via the user interface 1100. As a non-limiting example, any suitable combination of power data 310, account data 312, environment data 314, host / instance data 316, and / or operational data 318 may be presented in the user interface 1100.
[0167] While the specific content of the VEPO data 1109 displayed in the user interface 1100 may vary, as shown, the VEPO data 1109 includes VEPO level data (e.g., identifiers corresponding to the VEPO response levels of FIG. 8 presented in column 1110), current power consumption values (e.g., current power consumption values presented in column 1112), estimated reductions associated with each VEPO level (e.g., estimated VEPO reductions presented in column 1114), and estimated impacts associated with each VEPO level (e.g., estimated VEPO impacts presented in column 1116).
[0168] At 1118, a current aggregate power threshold (e.g., the current aggregate power threshold determined by demand engine 710 of FIG. 7) may be displayed. At 1120, a current power reduction required to meet the current aggregate power threshold (e.g., indicating a 5 kW reduction is required) may be calculated (e.g., by demand engine 710) based, at least in part, on determining the difference between the current aggregate power threshold (e.g., 11 kW) and the current power consumption (e.g., 16 kW) presented at 1112.
[0169] User interface 1100 can be configured to present estimated impact data corresponding to each of the VEPO response levels ("VEPO Levels" for brevity) in column 1116. In some embodiments, areas corresponding to rows in column 1116 may be selectable to present an expanded view of estimated impact data corresponding to a particular VEPO level. Similarly, areas corresponding to rows in column 1114 may be selectable to present an expanded view of estimated VEPO reduction corresponding to a particular VEPO level.
[0170] User interface 1100 may include an indicator 1122 that indicates a recommended VEPO response level (as shown, VEPO Level 2 is recommended by the system). As discussed above, the recommended VEPO level may be identified by power orchestration service 702 based on the factors discussed above. In some embodiments, the recommended VEPO level (e.g., VEPO Level 2) may be selected based on determining that VEPO Level 2 is associated with the smallest amount of estimated VEPO reduction that meets the required reduction, as shown at 1120. In some embodiments, VEPO Level 2 may also or alternatively be recommended based, at least in part, on determining that the estimated impact of applying VEPO Level 2 provides the smallest impact relative to the estimated impact of VEPO Level 104.
[0171] In some embodiments, the area corresponding to each row of VEPO data 1109 may be selectable to select a given VEPO response level. Once selected, the user may be provided with an additional menu or option to instruct the system to proceed with the selected VEPO response according to the reduction action associated with that level.
[0172] FIG. 12 shows a flow illustrating an example method 1200 for training one or more machine learning models according to at least one embodiment. In some embodiments, the models 1202 can be trained (e.g., by the power management service 304 of FIG. 3 , the power orchestration service 302 of FIG. 3 , the compute service 306 of FIG. 3 , or a different device or system) using any suitable machine learning algorithm (e.g., supervised, unsupervised, etc.) and any suitable number of training datasets (e.g., training data 1208). A supervised machine learning algorithm refers to a machine learning task that involves learning an inferred function that maps inputs to outputs based on a labeled training dataset in which example input / output pairs are known. An unsupervised machine learning algorithm refers to a set of algorithms used to analyze and cluster unlabeled datasets (e.g., unlabeled data 1210). These algorithms are configured to identify patterns or groupings of data without requiring human intervention. In some embodiments, any suitable number of models 1202 can be trained during the training phase 1204.
[0173] At least one of the models can be trained to identify a likelihood or confidence that the aggregate power consumption of downstream devices (e.g., hosts 324 of FIG. 3 ) exceeds a budget threshold corresponding to upstream devices (e.g., string-level devices such as string PDUs 204 of FIG. 2 ) from which power is distributed to any suitable combination of hosts 324 and / or servers 104, according to at least one embodiment. The training data 1208 for training one or more of the models 1202 may include any suitable combination of individual and / or aggregate power consumption values (current or historical) of power data 310 corresponding to one or more hosts (e.g., hosts 324 of FIG. 3 , etc., versus servers 104 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the model may be trained to identify a predicted amount corresponding to the aggregate power consumption of any suitable combination of hosts 324 and / or servers 104. In some embodiments, the training data 1208 may include current or historical power caps applied to each host, and the output provided by such model 1208 may identify a predicted power cap value for any suitable combination of hosts 324 and / or servers 104.
[0174] At least one of the models can be trained to identify predicted workloads for any suitable combination of hosts 324 and / or servers 104 according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify predicted workloads may include any suitable combination of current and / or historical host / instance data 316 corresponding to one or more hosts (e.g., host 324 of FIG. 3 for server 104 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models can be trained to identify a likelihood or confidence that a set of predicted workloads will be assigned to any suitable combination of hosts 324 and / or servers 104.
[0175] At least one of the models can be trained to identify predicted partial or complete failures of one or more temperature control units according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify such failures may include any suitable combination of current and / or historical environmental data 314 and / or operational data 318 corresponding to the physical location where the host 324 and / or server 104 are located, current and / or historical external temperatures occurring currently and / or in the past for areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or historical ambient locations associated with the physical location, weather forecasts corresponding to future time periods, and current and / or historical failures that have occurred to one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ). In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the model can be trained to identify the likelihood or confidence that a particular failure (partial or complete) of the temperature control system 112 will occur in a future period. In some embodiments, the model can be trained to identify the degree of failure for the temperature control system 112 (e.g., 80% failure, 100% failure, etc.).
[0176] At least one of the models can be trained to provide an estimated aggregate power reduction required due to partial and / or complete failure of one or more temperature control units (e.g., temperature control system 112) according to at least one embodiment. Training data 1208 for training one or more of the models 1202 to identify an estimated aggregate power reduction required may include any suitable combination of current and / or historical data of environmental data 314 and / or operational data 318 corresponding to the physical locations where the hosts 324 and / or servers 104 are located, current and / or historical external temperatures occurring currently and / or in the past for areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or historical ambient locations associated with the physical locations, weather forecasts corresponding to future time periods, current and / or historical failures that have occurred to one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ), current and / or historical aggregate power thresholds, and / or current and / or historical power consumption levels of any suitable combination of hosts 324 and / or servers 104. In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models may be trained to identify the likelihood or confidence that estimated aggregate power reductions will be needed during a future time period.
[0177] At least one of the models can be trained to predict future (e.g., ambient and / or external) temperatures according to at least one embodiment. The training data 1208 for training one or more of the models 1202 to identify future temperatures may include any suitable combination of environmental data 314 and / or operational data 318 corresponding to the physical location where the host 324 and / or server 104 are located, current and / or past external temperatures occurring currently and / or in the past with respect to areas outside the physical environment (e.g., in the geographic area where the physical environment is located), current and / or past ambient locations associated with the physical location, weather forecasts corresponding to future time periods, current and / or past faults experienced by one or more temperature control units (e.g., temperature control system 112 of FIG. 1 ), current and / or past aggregate power thresholds, and / or current and / or past power consumption levels of any suitable combination of the host 324 and / or server 104. In some embodiments, the training data 1208 for training these particular models 1202 may include any suitable combination of power data 310, account data 312, environmental data 314, host / instance data 316, and / or operational data 318. In some embodiments, the models may be trained to identify the likelihood or confidence that a predicted temperature will occur during a future time period.
[0178] In general, models 1202 may include any suitable number of models. Models 1202 may be individually trained from training data 1208 described above to determine the likelihood or confidence or accuracy of the output provided by model 1202. In some embodiments, models 1202 may be configured to determine values corresponding to the examples provided above, where the likelihood and / or confidence values may indicate the degree of likelihood or confidence that the output provided by the model is accurate.
[0179] Generally, model 1202 can be trained during training phase 1204 using a supervised learning algorithm and labeled data 1206 to identify outputs described in the examples above. The likelihood value can be a binary indicator indicating the degree of likelihood, a percentage, a confidence value, or the like. The likelihood value can be a binary indicator indicating whether a particular budget amount is likely or unlikely to be breached, or the likelihood value can indicate the likelihood and / or confidence that a predicted output will become reality in the future. Labeled data 1206 can be any suitable portion of potential training data (e.g., training data 1208) that can be used to train the model to generate the outputs described above. Labeled data 1206 can include any suitable number of examples of current and / or historical data corresponding to any suitable number of devices in a physical environment (e.g., data center 102, server 104, temperature control system 112, PDUs of FIG. 2, etc.). In some embodiments, labeled data 1206 can include labels identifying known likelihoods and / or actual values. Using the labeled data 1206, a model (e.g., an inferred function) can be learned that maps an input (e.g., one or more instances of training data) to an output (e.g., a predicted value, or a likelihood / confidence value that the predicted value or output will occur in a future time period). In some embodiments, any suitable combination of the VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3, or separate services or systems, can train any suitable combination of the models 1202. In some embodiments, any suitable combination of the VEPO orchestration service 302, power management service 304, and / or compute service 306 can obtain a pre-trained version of the models 1202.
[0180] The models 1202, and the various types of those models described above, may include any suitable number of models trained using unsupervised learning techniques to identify likelihoods / confidences and / or predictive quantities corresponding to the examples provided above. Unsupervised machine learning algorithms are configured to learn patterns from untagged data. In some embodiments, the training phase 1204 may utilize an unsupervised machine learning algorithm to generate one or more of the models 1202. For example, the training data 1208 may include unlabeled data 1210 (e.g., any suitable combination of current and / or historical values corresponding to power data 310, account data 312, environmental data 314, host / instance data 316, operational data 318, etc.). The unlabeled data 1210 may be utilized with an unsupervised learning algorithm that segments entries in the unlabeled data 1210 into groups. The unsupervised learning algorithm may be configured to cluster similar entries into common groups. Examples of unsupervised learning algorithms may include clustering methods such as k-means clustering, DBScan, etc. In some embodiments, the unlabeled data 1210 can be clustered with the labeled data 1206 so that unlabeled instances in a given group can be assigned the same label as other labeled instances in that group.
[0181] In some embodiments, any suitable portion of the training data 1208 may be utilized to train the model 1202 during the training phase 1204. For example, 70% of the labeled data 1206 and / or unlabeled data 1210 may be utilized to train the model 1202. Once trained, or at any suitable point in time, the model 1202 may be evaluated to assess its quality (e.g., the accuracy of the output 1212 with respect to the label corresponding to the labeled data 1206). By way of example, a portion of the example labeled data 1206 and / or unlabeled data 1210 may be utilized as input to the model 1202 to generate the output 1212. By way of example, an example of the labeled data 1206 may be provided as input, and the corresponding output (e.g., output 1212) may be compared with a label already known to be associated with that example. If a portion of the output (e.g., a label) matches the example label, that portion of the output may be deemed accurate. Any suitable number of labeled examples can be utilized, and the number of correct labels can be compared to the total number of examples provided (and / or the total number of previously identified labels) to determine an accuracy value for a given model that quantifies the model's accuracy. For example, if 90 out of 100 input examples produce output labels that match previously known label examples, the model being evaluated can be determined to be 90% accurate.
[0182] In some embodiments, as the model 1202 is utilized with subsequent inputs, subsequent outputs generated by the model 1202 can be added to the corresponding inputs and used to retrain and / or update the model 1202 at 1216. In some embodiments, an example may not be used to retrain or update the model until a feedback procedure 1214 is performed. In the feedback procedure 1214, an example (e.g., an example including one or more historical consumption data instances corresponding to one or more devices and / or racks) and the corresponding output generated for that example by one of the models 1202 are presented to a user, who identifies whether the generated output (e.g., quantity and / or likelihood confidence value) is correct for a given example.
[0183] The training process shown in FIG. 12 (e.g., method 1200) can be performed any suitable number of times, at any suitable intervals and / or according to any suitable schedule, so that the accuracy of model 1202 improves over time.
[0184] In some embodiments, any suitable number and / or combination of models 1202 may be used to determine the output. In some embodiments, the power management service 402 may utilize any suitable combination of outputs provided by the models 1202 to determine whether a budget threshold for a given component (e.g., a column-level PDU) is likely to be breached (and / or likely to be breached by a certain amount). Thus, in some embodiments, models trained with any suitable combination of supervised and unsupervised learning algorithms may be utilized by the power management service 402.
[0185] FIG. 13 is a block diagram illustrating an example method 1300 for implementing a response level from a plurality of response levels based, at least in part, on estimated power reductions, according to at least one embodiment. Method 1300 may be performed by one or more components of power orchestration system 300 of FIG. 3 or subcomponents thereof discussed in connection with FIGS. 4 and 7. By way of example, method 1300 may be performed, at least in part, by any suitable combination of VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3. The operations of method 1300 may be performed in any suitable order. Method 1300 may include more or fewer operations than those shown in FIG. 13.
[0186] Method 1300 may begin at 1302, where a respective set of workloads executing on each of a plurality of hosts may be identified (e.g., by impact identification manager 706 of FIG. 7). In some embodiments, the respective set of workloads executing on each of the plurality of hosts may be identified based, at least in part, on host / instance data 316 of FIG. 3.
[0187] At 1304, a plurality of response levels may be identified (e.g., by impact identification manager 706 of FIG. 7 ), specifying the applicability of a respective set of reduction actions to a plurality of hosts. In some embodiments, the plurality of response levels may be identified from configuration data (e.g., configuration data 712 of FIG. 7 ). In some embodiments, the plurality of response levels (e.g., VEPO response levels of FIG. 8 ) specify the applicability of a respective set of reduction actions to a plurality of hosts (e.g., via one or more reduction actions corresponding to each response level). In some embodiments, a first response level of the plurality of response levels specifies the applicability of a first set of reduction actions to a plurality of hosts (e.g., based at least in part on one or more reduction actions associated with that response level). By way of example, the VEPO response level of FIG. 8 may specify that all free / idle hosts and hypervisors that can be fully evacuated are applicable to the action (e.g., powering off). In this manner, the reduction action indicates that a subset of resources (e.g., free / idle hosts and hypervisors that can be fully evacuated) are applicable to a given reduction action (e.g., powering off) and, through association, the corresponding response level.
[0188] At 1306, a first estimate for power reduction resulting from the first response level may be determined (e.g., by impact identification manager 706). In some embodiments, the first estimate may be determined based at least on (a) a respective set of workloads executing on each of the multiple hosts and (b) the applicability of a first set of reduction actions on the multiple hosts according to the first response level.
[0189] At 1308, a first response level may be selected from the plurality of response levels (e.g., by enforcement manager 708 of FIG. 7 ) based at least on a first estimate of power reduction resulting from the first response level. In some embodiments, enforcement manager 708 may select the first response level based, at least in part, on identifying that a respective set of reduction actions for the plurality of hosts in the first response level are sufficient to reduce the power consumption levels of the plurality of hosts to a value below an aggregate power threshold (e.g., an aggregate power threshold determined by demand manager 710 of FIG. 7 ). In some embodiments, corresponding estimates (e.g., VEPO response levels of FIG. 8 ) for each of the response levels may be identified (e.g., by impact identification manager 706), and the first response level may be selected, at least in part, based on being sufficient to reduce the power consumption levels of the plurality of hosts to a value below the aggregate power threshold while having the least impact (e.g., impacting fewer workloads, hosts, customers, or any suitable combination of the above).
[0190] At 1310, a first set of curtailment actions may be applied to the plurality of hosts according to the selected first response level. By way of example, instructions may be sent (e.g., by enforcement manager 708 of FIG. 7 ) that cause the first set of curtailment actions to be applied to the plurality of hosts. In some embodiments, the instructions may be sent to power management service 304 of FIG. 3 and / or compute service 306 of FIG. 3 . At least some of the instructions may result in power caps being identified / applied, workloads and / or virtual machines being migrated, virtual machines and / or hosts being shut off, workloads being migrated, or any suitable combination of the above.
[0191] FIG. 14 is a block diagram illustrating an example method 1400 for implementing a response level from a plurality of response levels based, at least in part, on predicted impact, according to at least one embodiment. Method 1400 may be performed by one or more components of power orchestration system 300 of FIG. 3 or subcomponents thereof discussed in connection with FIGS. 4 and 7. By way of example, method 1400 may be performed, at least in part, by any suitable combination of VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3. The operations of method 1400 may be performed in any suitable order. Method 1400 may include more or fewer operations than those shown in FIG. 14.
[0192] Method 1400 may begin at 1402, where a respective set of workloads executing on each of a plurality of hosts may be identified (e.g., by impact identification manager 706 of FIG. 7). In some embodiments, the respective set of workloads executing on each of the plurality of hosts may be identified based at least in part on host / instance data 316 of FIG. 3.
[0193] At 1404, a plurality of response levels may be identified (e.g., by impact identification manager 706 of FIG. 7 ), specifying the applicability of a respective set of reduction actions to a plurality of hosts. In some embodiments, the plurality of response levels may be identified from configuration data (e.g., configuration data 712 of FIG. 7 ). In some embodiments, the plurality of response levels (e.g., VEPO response levels of FIG. 8 ) specify the applicability of a respective set of reduction actions to a plurality of hosts (e.g., via one or more reduction actions corresponding to each response level). In some embodiments, a first response level of the plurality of response levels specifies the applicability of a first set of reduction actions to a plurality of hosts (e.g., based at least in part on one or more reduction actions associated with that response level). By way of example, the VEPO response level of FIG. 8 may specify that all free / idle hosts and hypervisors that can be fully evacuated are applicable to the action (e.g., powering off). In this manner, the reduction action indicates that a subset of resources (e.g., free / idle hosts and hypervisors that can be fully evacuated) are applicable to a given reduction action (e.g., powering off) and, through association, the corresponding response level.
[0194] At 1406, a first estimate for a predicted impact resulting from the first response level may be determined (e.g., by impact identification manager 706). In some embodiments, the first estimate may be determined based at least on (a) a respective set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of reduction actions on the plurality of hosts according to the first response level. In some embodiments, the predicted workload impact may be determined based on at least one of a host priority, or a number of affected hosts, a number of affected workloads, a priority level of the affected workloads, a number of affected customers, or a priority level of the affected customers. In some embodiments, the predicted workload impact may be predicted based at least in part on historical host / instance data (e.g., host / instance data 316 of FIG. 3), account data (e.g., account data 312), and / or any suitable data discussed in connection with FIG. 3. In some embodiments, the predicted workload impact may be predicted based, at least in part, on providing input data (e.g., host / instance data 316, account data 312, etc., of FIG. 3) to a machine learning model (e.g., one or more of models 1202 of FIG. 12).
[0195] At 1408, a first response level may be selected from the plurality of response levels (e.g., by enforcement manager 708 of FIG. 7 ) based at least on a first estimate of the impact to the workload resulting from the first response level. In some embodiments, enforcement manager 708 may select the first response level based, at least in part, on determining that the first estimate of the impact to the workload indicates a lesser impact to the workload than an estimated impact to the workload corresponding to at least one other response level of the plurality of response levels. In some embodiments, the impact to the workload may indicate a lesser impact to the workload, host, and / or customer. In some embodiments, enforcement manager 708 may select the first response level further based on determining that applying the set of reduction actions corresponding to the first response level is expected to reduce power consumption levels of the plurality of hosts to a value below the aggregate power threshold.
[0196] At 1410, a first set of curtailment actions may be applied to the plurality of hosts according to the selected first response level. By way of example, instructions may be sent (e.g., by enforcement manager 708 of FIG. 7 ) that cause the first set of curtailment actions to be applied to the plurality of hosts. In some embodiments, the instructions may be sent to power management service 304 of FIG. 3 and / or compute service 306 of FIG. 3 . At least some of the instructions may result in power caps being identified / applied, workloads and / or virtual machines being migrated, virtual machines and / or hosts being shut off, workloads being migrated, or any suitable combination of the above.
[0197] FIG. 15 is a block diagram illustrating an exemplary method for preemptively migrating workloads affected by a selected response level, according to at least one embodiment. Method 1500 may be performed by one or more components of power orchestration system 300 of FIG. 3 or subcomponents thereof discussed in connection with FIGS. 4 and 7. By way of example, method 1500 may be performed, at least in part, by any suitable combination of VEPO orchestration service 302, power management service 304, and / or compute service 306 of FIG. 3. The operations of method 1500 may be performed in any suitable order. Method 1500 may include more or fewer operations than those shown in FIG. 15.
[0198] Method 1500 may begin at 1502, where a respective predicted set of workloads executing on each of a plurality of hosts during a future time may be determined (e.g., by impact identification manager 706 of FIG. 7). In some embodiments, the respective predicted set of workloads executing on each of the plurality of hosts during a future time may be identified based at least in part on historical data (e.g., included in host / instance data 316 of FIG. 3). In some embodiments, the predicted set of workloads may be predicted based at least in part on providing input data (e.g., host / instance data 316, account data 312, etc. of FIG. 3) to a machine learning model (e.g., one or more of models 1202 of FIG. 12).
[0199] At 1504, a plurality of response levels may be identified (e.g., by impact identification manager 706 of FIG. 7 ) that specify the applicability of a respective set of reduction actions to a plurality of hosts. In some embodiments, the plurality of response levels may be identified from configuration data (e.g., configuration data 712 of FIG. 7 ). In some embodiments, the plurality of response levels (e.g., VEPO response levels of FIG. 8 ) specify the applicability of a respective set of reduction actions to a plurality of hosts (e.g., via one or more reduction actions corresponding to each response level). In some embodiments, a first response level of the plurality of response levels specifies the applicability of a first set of reduction actions to a plurality of hosts (e.g., based at least in part on one or more reduction actions associated with that response level). By way of example, the VEPO response level of FIG. 8 may specify that all free / idle hosts and hypervisors that can be fully evacuated are applicable to the action (e.g., powering off). In this manner, the reduction action indicates that a subset of resources (e.g., free / idle hosts and hypervisors that can be fully evacuated) are applicable to a given reduction action (e.g., powering off) and, through association, to the corresponding response level.
[0200] At 1506, a first estimate for power reduction resulting from the first response level may be determined (e.g., by impact identification manager 706). In some embodiments, the first estimate may be determined based at least on (a) a respective predicted set of workloads executing on each of the plurality of hosts and (b) applicability of the first set of reduction actions at the plurality of hosts according to the first response level. In some embodiments, the power reduction may be determined based on at least one of a predicted set of workloads, a predicted power consumption value, a predicted priority, etc. In some embodiments, the predicted workload, predicted power consumption value, predicted priority, etc. may be predicted based, at least in part, on historical / current power data (e.g., included in power data 310 of FIG. 3 , historical / current host / instance data (e.g., included in host / instance data 316 of FIG. 3 ), historical / current account data (e.g., included in account data 312), and / or any suitable historical and / or current data discussed in connection with FIG. 3 . In some embodiments, the power reduction resulting from the first response level may be determined, at least in part, based on providing input data (e.g., power data 310, host / instance data 316, account data 312, environmental data 314, operational data 318, any suitable combination of the above) to a machine learning model (e.g., one or more of models 1202 of FIG. 12 ).
[0201] At 1508, a first response level may be selected from the plurality of response levels (e.g., by enforcement manager 708 of FIG. 7 ) based at least on a first estimate of power reduction resulting from the first response level. In some embodiments, enforcement manager 708 may select the first response level based, at least in part, on identifying that a respective set of reduction actions for the plurality of hosts in the first response level are sufficient to reduce the power consumption levels of the plurality of hosts to a value below an aggregate power threshold (e.g., an aggregate power threshold determined by demand manager 710 of FIG. 7 ). In some embodiments, corresponding estimates for each of the response levels (e.g., VEPO response levels of FIG. 8 ) may be identified (e.g., by impact identification manager 706), and the first response level may be selected, at least in part, based on being sufficient to reduce the power consumption levels of the plurality of hosts to a value below the aggregate power threshold while having the least impact (e.g., impacting fewer workloads, hosts, customers, or any suitable combination of the above).
[0202] At 1510, one or more workloads may be identified (e.g., by impact identification manager 706 of FIG. 7 ) that (a) are currently executing on the plurality of hosts and (b) would be affected by applying a first set of reduction actions to the plurality of hosts according to a selected first response level. In some embodiments, the one or more workloads may be predicted workloads. Impact identification manager 706 may predict the one or more workloads based, at least in part, on historical data (e.g., included in host / instance data 316 of FIG. 3 ) and / or, at least in part, on output provided by at least one of models 1202 of FIG. 12 (e.g., output provided in response to providing input data including any suitable historical and / or current data for host / instance data 316).
[0203] At 1512, the affected workloads may be preemptively migrated from multiple hosts to one or more other hosts ahead of a future time period. As a non-limiting example, the affected workloads (and / or the virtual machines they execute) may be migrated to other hosts to reduce the total number of hosts used and / or to free up hosts and / or virtual machines so that they can be shut down. In some embodiments, these migrations may be initiated, at least in part, by the implementation manager 708 based on sending instructions to the compute service 306 of FIG. 3 .
[0204] 16-19 illustrate a number of example environments that may be hosted by host 324 of FIG. 3. The environments illustrated in FIGS. 16-19 illustrate cloud computing, multi-tenant environments. As discussed above, cloud computing and / or multi-tenant environments, as well as other environments, may benefit from utilizing the power management techniques disclosed herein. These techniques enable data centers, including components that implement the environments described in connection with FIGS. 1-4 and 7, among others, to utilize the data center's power resources more efficiently than conventional techniques that leave large amounts of power unused.
[0205] IaaS Infrastructure Example Infrastructure as a Service (IaaS) is a specific type of cloud computing. IaaS can be configured to provide virtualized computing resources over a public network (e.g., the Internet). In the IaaS model, a cloud computing provider can host infrastructure components (e.g., servers, storage devices, network nodes (e.g., hardware), deployment software, platform virtualization (e.g., hypervisor layer), etc.). In some cases, an IaaS provider can also offer various services pertaining to those infrastructure components (example services include billing software, monitoring software, logging software, load balancing software, clustering software, etc.). Accordingly, these services can be policy-driven, allowing IaaS users to implement policies that drive load balancing to maintain application availability and performance.
[0206] In some cases, IaaS customers can access resources and services over a wide area network (WAN) such as the Internet and use the cloud provider's services to install the remaining elements of their application stack. For example, a user can log into an IaaS platform to create virtual machines (VMs), install an operating system (OS) on each VM, deploy middleware such as databases, create storage buckets for workloads and backups, and even install enterprise software on the VMs. The customer can then use the provider's services to perform a variety of functions, including distributing network traffic, troubleshooting application issues, monitoring performance, and managing disaster recovery.
[0207] In most cases, the cloud computing model requires the participation of a cloud provider, which can be, but is not required to be, a third-party service that specializes in providing (e.g., offering, renting, selling) IaaS. An entity can also choose to deploy a private cloud and become its own provider of infrastructure services.
[0208] In some examples, IaaS deployment is the process of placing a new application, or a new version of an application, onto a prepared application server, etc. IaaS deployment may also include the process of preparing the server (e.g., installing libraries, daemons, etc.), which is often managed by the cloud provider below the hypervisor layer (e.g., server, storage, network hardware, and virtualization). Thus, the customer may be responsible for handling the deployment of the (OS), middleware, and / or application (e.g., in self-service virtual machines (e.g., that can be spun up on demand) etc.
[0209] In some instances, IaaS provisioning may refer to obtaining the computers or virtual hosts to be used and also installing the necessary libraries or services on those computers or virtual hosts. In most cases, deployment does not include provisioning, which must be performed first.
[0210] In some cases, IaaS provisioning presents two distinct challenges. First, there is the initial challenge of provisioning the initial set of infrastructure before anything is running. Second, there is the challenge of evolving the existing infrastructure (e.g., adding new services, modifying services, removing services, etc.) once everything is provisioned. In some cases, these two challenges can be addressed by allowing the configuration of the infrastructure to be defined declaratively. In other words, the infrastructure (e.g., what components are needed and how they interact) can be defined by one or more configuration files. Thus, the overall topology of the infrastructure (e.g., what resources depend on what and how each works together) can be described declaratively. In some cases, once the topology is defined, workflows can be generated to create and / or manage the different components described in the configuration files.
[0211] In some examples, the infrastructure can have many interconnected elements. For example, there can be one or more virtual private clouds (VPCs) (e.g., potential on-demand pools of configurable and / or shared computing resources), also known as a core network. In some examples, there can also be one or more inbound / outbound traffic group rules and one or more virtual machines (VMs) provisioned to define how inbound and / or outbound traffic on the network is configured. Other infrastructure elements, such as load balancers, databases, etc., can also be provisioned. The infrastructure can evolve over time as more infrastructure elements are desired and / or added.
[0212] In some cases, continuous deployment techniques can be employed to enable deployment of infrastructure code across various virtual computing environments. Additionally, the described techniques can enable infrastructure management within these environments. In some examples, a service team may write code that is desired to be deployed to one or more, but often many, different production environments (e.g., across various different geographic locations, sometimes across the globe). However, in some examples, the infrastructure onto which the code will be deployed must first be set up. In some cases, provisioning can be done manually, utilizing provisioning tools to provision resources, and / or utilizing deployment tools to deploy the code once the infrastructure has been provisioned.
[0213] 16 is a block diagram 1600 illustrating an example IaaS architecture pattern according to at least one embodiment. A service operator 1602 may be communicatively coupled to a secure host tenancy 1604, which may include a virtual cloud network (VCN) 1606 and a secure host subnet 1608. In some examples, the service operator 1602 may employ one or more client computing devices, which may be portable handheld devices (e.g., iPhone®, mobile phone, iPad®, computing tablet, personal digital assistant (PDA)), or wearable devices (e.g., Google® Glass head-mounted display) running software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 8, Palm OS, and supporting the Internet, email, short message service (SMS), Blackberry®, or other communication protocols. Alternatively, the client computing devices may be general-purpose personal computers, including, by way of example, personal and / or laptop computers running various versions of the Microsoft Windows operating system, the Apple Macintosh operating system, and / or the Linux operating system. The client computing devices may be workstation computers running any of a variety of commercially available UNIX or UNIX-like operating systems, including, without limitation, various GNU / Linux operating systems such as Google Chrome OS.Alternatively or additionally, the client computing devices may be any other electronic devices, such as thin client computers, Internet-enabled gaming systems (e.g., Microsoft Xbox gaming consoles with or without Kinect® gesture input devices), and / or personal messaging devices, that can communicate over a network that has access to VCN 1606 and / or the Internet.
[0214] VCN 1606 may include a local peering gateway (LPG) 1610, which may be communicatively coupled to a secure shell (SSH) VCN 1612 via the LPG 1610 included in the SSH VCN 1612. The SSH VCN 1612 may include an SSH subnet 1614, which may be communicatively coupled to a control plane VCN 1616 via the LPG 1610 included in the control plane VCN 1616. The SSH VCN 1612 may also be communicatively coupled to a data plane VCN 1618 via the LPG 1610. The control plane VCN 1616 and the data plane VCN 1618 may be included in a service tenancy 1619, which may be owned and / or operated by the IaaS provider.
[0215] The control plane VCN 1616 may include a control plane demilitarized zone (DMZ) tier 1620 that acts as a perimeter network (e.g., the portion of the enterprise network between the enterprise intranet and external networks). DMZ-based servers have limited responsibility and can help mitigate contained breaches. Additionally, the DMZ tier 1620 may include one or more load balancer (LB) subnets 1622, a control plane app tier 1624 that may include an app subnet 1626, and a control plane data tier 1628 that may include a database (DB) subnet 1630 (e.g., a front-end DB subnet and / or a back-end DB subnet). LB subnet 1622 included in control plane DMZ tier 1620 can be communicatively coupled to app subnet 1626 included in control plane app tier 1624 and to an Internet gateway 1634 that may be included in control plane VCN 1616, and app subnet 1626 can be communicatively coupled to DB subnet 1630 included in control plane data tier 1628 and to a service gateway 1636 and a network address translation (NAT) gateway 1638. Control plane VCN 1616 can include service gateway 1636 and NAT gateway 1638.
[0216] The control plane VCN 1616 may include a data plane mirrored app tier 1640, which may include an app subnet 1626. The app subnet 1626 included in the data plane mirrored app tier 1640 may include a virtual network interface controller (VNIC) 1642 on which a compute instance 1644 can run. The compute instance 1644 may communicatively couple the app subnet 1626 of the data plane mirrored app tier 1640 to the app subnet 1626, which may be included in the data plane app tier 1646.
[0217] The data plane VCN 1618 may include a data plane app layer 1646, a data plane DMZ layer 1648, and a data plane data layer 1650. The data plane DMZ layer 1648 may include a LB subnet 1622 that can be communicatively coupled to an app subnet 1626 of the data plane app layer 1646 and an internet gateway 1634 of the data plane VCN 1618. The app subnet 1626 can be communicatively coupled to a service gateway 1636 of the data plane VCN 1618 and a NAT gateway 1638 of the data plane VCN 1618. The data plane data layer 1650 may also include a DB subnet 1630 that can be communicatively coupled to the app subnet 1626 of the data plane app layer 1646.
[0218] The internet gateway 1634 of the control plane VCN 1616 and the internet gateway 1634 of the data plane VCN 1618 can be communicatively coupled to a metadata management service 1652, which can be communicatively coupled to the public internet 1654. The public internet 1654 can be communicatively coupled to a NAT gateway 1638 of the control plane VCN 1616 and the NAT gateway 1638 of the data plane VCN 1618. The service gateway 1636 of the control plane VCN 1616 and the service gateway 1636 of the data plane VCN 1618 can be communicatively coupled to cloud services 1656.
[0219] In some examples, the service gateway 1636 of the control plane VCN 1616 or the service gateway 1636 of the data plane VCN 1618 can make application programming interface (API) calls to the cloud service 1656 without traversing the public Internet 1654. The API calls from the service gateway 1636 to the cloud service 1656 can be one-way. That is, the service gateway 1636 can make an API call to the cloud service 1656, and the cloud service 1656 can send the requested data to the service gateway 1636. However, the cloud service 1656 cannot initiate the API call to the service gateway 1636.
[0220] In some examples, secure host tenancy 1604 can be directly connected to service tenancy 1619, which may otherwise be isolated. Secure host subnet 1608 can communicate with SSH subnet 1614 through LPG 1610, which may enable bidirectional communication in an otherwise isolated system. By connecting secure host subnet 1608 to SSH subnet 1614, secure host subnet 1608 can access other entities in service tenancy 1619.
[0221] The control plane VCN 1616 can enable users of the service tenancy 1619 to set up or otherwise provision desired resources. The desired resources provisioned in the control plane VCN 1616 can be deployed or otherwise used in the data plane VCN 1618. In some examples, the control plane VCN 1616 can be isolated from the data plane VCN 1618, and the data plane mirror app layer 1640 of the control plane VCN 1616 can communicate with the data plane app layer 1646 of the data plane VCN 1618 via a VNIC 1642, which can be included in the data plane mirror app layer 1640 and the data plane app layer 1646.
[0222] In some examples, a user or customer of the system can make a request, for example, a request for a create, read, update, or delete (CRUD) operation, via the public internet 1654, which can communicate such a request to the metadata management service 1652. The metadata management service 1652 can communicate the request to the control plane VCN 1616 via the internet gateway 1634. The request can be received by the LB subnet 1622 included in the control plane DMZ tier 1620. The LB subnet 1622 may determine that the request is valid, and in response to this determination, the LB subnet 1622 can send the request to the app subnet 1626 included in the control plane app tier 1624. If the request is validated and requires a call to the public internet 1654, the call to the public internet 1654 can be sent to the NAT gateway 1638, which can make the call to the public internet 1654. Metadata that may be desired to be stored by the request can be stored in the DB subnet 1630.
[0223] In some examples, the data plane mirror app layer 1640 can facilitate direct communication between the control plane VCN 1616 and the data plane VCN 1618. For example, it may be desired that a change, update, or other suitable modification to the configuration be applied to resources included in the data plane VCN 1618. The control plane VCN 1616 can communicate directly with the resources included in the data plane VCN 1618 via the VNIC 1642, thereby performing the change, update, or other suitable modification to the configuration on the resources.
[0224] In some embodiments, the control plane VCN 1616 and the data plane VCN 1618 may be included in the service tenancy 1619. In this case, a user or customer of the system may not own or operate either the control plane VCN 1616 or the data plane VCN 1618. Instead, an IaaS provider may own or operate the control plane VCN 1616 and the data plane VCN 1618, both of which may be included in the service tenancy 1619. This embodiment may enable network isolation, thereby preventing users or customers from interacting with other users' or customers' resources. This embodiment may also enable users or customers of the system to store databases privately without having to rely on the public internet 1654, which may not have the desired level of threat protection for storage.
[0225] In another embodiment, the LB subnet 1622 included in the control plane VCN 1616 can be configured to receive signals from the service gateway 1636. In this embodiment, the control plane VCN 1616 and the data plane VCN 1618 can be configured to be called by the IaaS provider's customers without calling the public internet 1654. Customers of the IaaS provider may desire this embodiment because databases used by the customers can be stored in the service tenancy 1619, which can be controlled by the IaaS provider and isolated from the public internet 1654.
[0226] Figure 17 is a block diagram 1700 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1702 (e.g., service operator 1602 of Figure 16) can be communicatively coupled to a secure host tenancy 1704 (e.g., secure host tenancy 1604 of Figure 16), which can include a virtual cloud network (VCN) 1706 (e.g., VCN 1606 of Figure 16) and a secure host subnet 1708 (e.g., secure host subnet 1608 of Figure 16). VCN 1706 can include a local peering gateway (LPG) 1710 (e.g., LPG 1610 of Figure 16), which can be communicatively coupled to a secure shell (SSH) VCN 1712 (e.g., SSH VCN 1612 of Figure 16) via the LPG 1610 included in the SSH VCN 1712. SSH VCN 1712 can include SSH subnet 1714 (e.g., SSH subnet 1614 in FIG. 16 ), and SSH VCN 1712 can be communicatively coupled to control plane VCN 1716 (e.g., control plane VCN 1616 in FIG. 16 ) via LPG 1710 included in control plane VCN 1716. Control plane VCN 1716 can be included in service tenancy 1719 (e.g., service tenancy 1619 in FIG. 16 ), and data plane VCN 1718 (e.g., data plane VCN 1618 in FIG. 16 ) can be included in customer tenancy 1721, which may be owned or operated by a user or customer of the system.
[0227] The control plane VCN 1716 may include a control plane DMZ tier 1720 (e.g., control plane DMZ tier 1620 in FIG. 16 ) that may include a LB subnet 1722 (e.g., LB subnet 1622 in FIG. 16 ), a control plane app tier 1724 (e.g., control plane app tier 1624 in FIG. 16 ) that may include an app subnet 1726 (e.g., app subnet 1626 in FIG. 16 ), and a control plane data tier 1728 (e.g., control plane data tier 1628 in FIG. 16 ) that may include a database (DB) subnet 1730 (e.g., similar to DB subnet 1630 in FIG. 16 ). LB subnet 1722 included in control plane DMZ tier 1720 can be communicatively coupled to app subnet 1726 included in control plane app tier 1724 and to an Internet gateway 1734 (e.g., Internet gateway 1634 in FIG. 16 ) that may be included in control plane VCN 1716, and app subnet 1726 can be communicatively coupled to DB subnet 1730 included in control plane data tier 1728 and to a service gateway 1736 (e.g., service gateway 1636 in FIG. 16 ) and a network address translation (NAT) gateway 1738 (e.g., NAT gateway 1638 in FIG. 16 ). Control plane VCN 1716 can include service gateway 1736 and NAT gateway 1738.
[0228] The control plane VCN 1716 may include a data plane mirrored app layer 1740 (e.g., data plane mirrored app layer 1640 of FIG. 16 ), which may include an app subnet 1726. The app subnet 1726 included in the data plane mirrored app layer 1740 may include a virtual network interface controller (VNIC) 1742 (e.g., VNIC 1642) on which a compute instance 1744 (e.g., similar to compute instance 1644 of FIG. 16 ) can run. The compute instance 1744 can facilitate communication between the app subnet 1726 of the data plane mirrored app layer 1740 and the app subnet 1726, which may be included in the data plane app layer 1746 (e.g., data plane app layer 1646 of FIG. 16 ), via the VNIC 1742 included in the data plane mirrored app layer 1740 and the VNIC 1742 included in the data plane app layer 1746.
[0229] The internet gateway 1734 included in the control plane VCN 1716 can be communicatively coupled to a metadata management service 1752 (e.g., metadata management service 1652 of FIG. 16 ), which can be communicatively coupled to the public internet 1754 (e.g., public internet 1654 of FIG. 16 ). The public internet 1754 can be communicatively coupled to a NAT gateway 1738 included in the control plane VCN 1716. The service gateway 1736 included in the control plane VCN 1716 can be communicatively coupled to cloud services 1756 (e.g., cloud services 1656 of FIG. 16 ).
[0230] In some examples, data plane VCN 1718 may be included in customer tenancy 1721. In this case, the IaaS provider may provide a control plane VCN 1716 for each customer, and the IaaS provider may set up a unique compute instance 1744 for each customer, which is included in service tenancy 1719. Each compute instance 1744 may enable communication between the control plane VCN 1716, which is included in service tenancy 1719, and the data plane VCN 1718, which is included in customer tenancy 1721. The compute instance 1744 may enable resources provisioned in the control plane VCN 1716, which is included in service tenancy 1719, to be deployed to or otherwise used in the data plane VCN 1718, which is included in customer tenancy 1721.
[0231] In another example, a customer of the IaaS provider may have a database that resides in customer tenancy 1721. In this example, control plane VCN 1716 may include data plane mirror app tier 1740, which may include app subnet 1726. Data plane mirror app tier 1740 may reside in data plane VCN 1718, but data plane mirror app tier 1740 may not reside in data plane VCN 1718. That is, data plane mirror app tier 1740 may have access to customer tenancy 1721, but data plane mirror app tier 1740 may not reside in data plane VCN 1718 or be owned and operated by the IaaS provider customer. Data plane mirror app tier 1740 may be configured to make calls to data plane VCN 1718, but may not be configured to make calls to any entities included in control plane VCN 1716. A customer may wish to deploy or otherwise use resources in data plane VCN 1718 that have been provisioned in control plane VCN 1716, and data plane mirror app layer 1740 can facilitate the desired deployment or other use of the customer's resources.
[0232] In some embodiments, the IaaS provider's customer can apply filters to the data plane VCN 1718. In this embodiment, the customer can determine what the data plane VCN 1718 can access, and the customer can restrict access from the data plane VCN 1718 to the public internet 1754. The IaaS provider may not be able to apply filters or otherwise control access from the data plane VCN 1718 to any external networks or databases. Applying filters and controls to the data plane VCN 1718 that the customer includes in their customer tenancy 1721 can help to isolate the data plane VCN 1718 from other customers and from the public internet 1754.
[0233] In some embodiments, service gateway 1736 can call cloud services 1756 to access services that may not reside on the public internet 1754, on the control plane VCN 1716, or on the data plane VCN 1718. The connection between cloud service 1756 and control plane VCN 1716 or data plane VCN 1718 may not be live or continuous. Cloud service 1756 may reside on different networks owned or operated by the IaaS provider. Cloud service 1756 may be configured to receive calls from service gateway 1736 and may not be configured to receive calls from the public internet 1754. Some cloud services 1756 may be isolated from other cloud services 1756, and control plane VCN 1716 may be isolated from cloud services 1756 that may not be in the same region as control plane VCN 1716. For example, control plane VCN 1716 may be located in “Region 1,” and cloud service “Deployment 15” may be located in Region 1 and “Region 2.” If a call to deployment 15 is made by service gateway 1736 included in control plane VCN 1716 located in region 1, the call may be sent to deployment 15 in region 1. In this example, control plane VCN 1716 or deployment 15 in region 1 may not be communicatively coupled or otherwise in communication with deployment 15 in region 2.
[0234] Figure 18 is a block diagram 1800 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1802 (e.g., service operator 1602 of Figure 16) can be communicatively coupled to a secure host tenancy 1804 (e.g., secure host tenancy 1604 of Figure 16), which can include a virtual cloud network (VCN) 1806 (e.g., VCN 1606 of Figure 16) and a secure host subnet 1808 (e.g., secure host subnet 1608 of Figure 16). VCN 1806 can include an LPG 1810 (e.g., LPG 1610 of Figure 16), which can be communicatively coupled to an SSH VCN 1812 (e.g., SSH VCN 1612 of Figure 16) via the LPG 1810 included in the SSH VCN 1812. SSH VCN 1812 can include SSH subnet 1814 (e.g., SSH subnet 1614 in FIG. 16 ), and SSH VCN 1812 can be communicatively coupled to control plane VCN 1816 (e.g., control plane VCN 1616 in FIG. 16 ) via LPG 1810 included in control plane VCN 1816 and to data plane VCN 1818 (e.g., data plane 1618 in FIG. 16 ) via LPG 1810 included in data plane VCN 1818. Control plane VCN 1816 and data plane VCN 1818 can be included in service tenancy 1819 (e.g., service tenancy 1619 in FIG. 16 ).
[0235] The control plane VCN 1816 may include a control plane DMZ tier 1820 (e.g., the control plane DMZ tier 1620 of FIG. 16 ) that may include a load balancer (LB) subnet 1822 (e.g., the LB subnet 1622 of FIG. 16 ), a control plane app tier 1824 (e.g., the control plane app tier 1624 of FIG. 16 ) that may include an app subnet 1826 (e.g., similar to the app subnet 1626 of FIG. 16 ), and a control plane data tier 1828 (e.g., the control plane data tier 1628 of FIG. 16 ) that may include a DB subnet 1830. LB subnet 1822 included in control plane DMZ tier 1820 can be communicatively coupled to app subnet 1826 included in control plane app tier 1824 and to an Internet gateway 1834 (e.g., Internet gateway 1634 in FIG. 16 ) that may be included in control plane VCN 1816, and app subnet 1826 can be communicatively coupled to DB subnet 1830 included in control plane data tier 1828 and to service gateway 1836 (e.g., service gateway in FIG. 16 ) and network address translation (NAT) gateway 1838 (e.g., NAT gateway 1638 in FIG. 16 ). Control plane VCN 1816 may include service gateway 1836 and NAT gateway 1838.
[0236] Data plane VCN 1818 may include a data plane app layer 1846 (e.g., data plane app layer 1646 in FIG. 16 ), a data plane DMZ layer 1848 (e.g., data plane DMZ layer 1648 in FIG. 16 ), and a data plane data layer 1850 (e.g., data plane data layer 1650 in FIG. 16 ). Data plane DMZ layer 1848 may include LB subnet 1822, which may be communicatively coupled to trusted app subnet 1860 and untrusted app subnet 1862 of data plane app layer 1846, which are included in data plane VCN 1818, as well as to Internet gateway 1834. Trusted app subnet 1860 may be communicatively coupled to service gateway 1836, which is included in data plane VCN 1818, NAT gateway 1838, which is included in data plane VCN 1818, and DB subnet 1830, which is included in data plane data layer 1850. The untrusted app subnet 1862 may be communicatively coupled to a service gateway 1836 included in the data plane VCN 1818 and to a DB subnet 1830 included in the data plane data layer 1850. The data plane data layer 1850 may include a DB subnet 1830 that may be communicatively coupled to a service gateway 1836 included in the data plane VCN 1818.
[0237] The untrusted app subnet 1862 may include one or more primary VNICs 1864(1)-(N), which may be communicatively coupled to tenant virtual machines (VMs) 1866(1)-(N). Each tenant VM 1866(1)-(N) may be communicatively coupled to a respective app subnet 1867(1)-(N), which may be included in a respective container egress VCN 1868(1)-(N), which may be included in a respective customer tenancy 1870(1)-(N). Each secondary VNIC 1872(1)-(N) may facilitate communication between the untrusted app subnet 1862 included in the data plane VCN 1818 and the app subnet included in the container egress VCN 1868(1)-(N). Each container egress VCN(1)-(N) may include a NAT gateway 1838, which may be communicatively coupled to the public internet 1854 (e.g., public internet 1654 in FIG. 16 ).
[0238] The internet gateway 1834 included in the control plane VCN 1816 and the internet gateway 1834 included in the data plane VCN 1818 can be communicatively coupled to a metadata management service 1852 (e.g., metadata management system 1652 of FIG. 16 ), which can be communicatively coupled to the public internet 1854. The public internet 1854 can be communicatively coupled to a NAT gateway 1838 included in the control plane VCN 1816 and the NAT gateway 1838 included in the data plane VCN 1818. The service gateway 1836 included in the control plane VCN 1816 and the service gateway 1836 included in the data plane VCN 1818 can be communicatively coupled to cloud services 1856.
[0239] In some embodiments, data plane VCN 1818 can be integrated with customer tenancy 1870. This integration can be useful or desirable for an IaaS provider's customer in some cases, such as when they may want support when running code. A customer may provide code to run that may be disruptive, may communicate with other customer resources, or may otherwise have undesirable effects. In response, the IaaS provider can determine whether to run the code that the customer has provided to the IaaS provider.
[0240] In some examples, a customer of an IaaS provider can grant temporary network access to the IaaS provider and request a function to be attached to the data plane app layer 1846. The code that performs the function can run in VMs 1866(1)-(N), and the code may not be configured to run anywhere else on the data plane VCN 1818. Each VM 1866(1)-(N) can be connected to one customer tenancy 1870. Each container 1871(1)-(N) contained in a VM 1866(1)-(N) can be configured to run code. In this case, there can be double isolation (e.g., containers 1871(1)-(N) can run code, and containers 1871(1)-(N) can be contained in VMs 1866(1)-(N) that are at least in the untrusted app subnet 1862). This can help prevent erroneous or otherwise unwanted code from damaging the IaaS provider's network or damaging a different customer's network. Containers 1871(1)-(N) can be communicatively coupled to customer tenancy 1870 and can be configured to send or receive data from customer tenancy 1870. Containers 1871(1)-(N) may not be configured to send or receive data from any other entity in data plane VCN 1818. When the code execution is complete, the IaaS provider can kill or otherwise discard containers 1871(1)-(N).
[0241] In some embodiments, trusted app subnet 1860 can execute code that can be owned or operated by the IaaS provider. In this embodiment, trusted app subnet 1860 can be communicatively coupled to DB subnet 1830 and configured to perform CRUD operations on DB subnet 1830. Untrusted app subnet 1862 can be communicatively coupled to DB subnet 1830, but in this embodiment, the untrusted app subnet can be configured to perform read operations within DB subnet 1830. Containers 1871(1)-(N) that can be included in each customer's VMs 1866(1)-(N) and that can execute code from the customer may not be communicatively coupled to DB subnet 1830.
[0242] In other embodiments, the control plane VCN 1816 and the data plane VCN 1818 may not be directly communicatively coupled. In this embodiment, there may be no direct communication between the control plane VCN 1816 and the data plane VCN 1818. However, communication may occur indirectly in at least one manner. An IaaS provider may establish an LPG 1810 that can facilitate communication between the control plane VCN 1816 and the data plane VCN 1818. In another example, the control plane VCN 1816 or the data plane VCN 1818 can make a call to a cloud service 1856 through a service gateway 1836. For example, a call from the control plane VCN 1816 to the cloud service 1856 may include a request for a service that can communicate with the data plane VCN 1818.
[0243] Figure 19 is a block diagram 1900 illustrating another example pattern of an IaaS architecture, according to at least one embodiment. A service operator 1902 (e.g., service operator 1602 of Figure 16) can be communicatively coupled to a secure host tenancy 1904 (e.g., secure host tenancy 1604 of Figure 16), which can include a virtual cloud network (VCN) 1906 (e.g., VCN 1606 of Figure 16) and a secure host subnet 1908 (e.g., secure host subnet 1608 of Figure 16). VCN 1906 can include an LPG 1910 (e.g., LPG 1610 of Figure 16), which can be communicatively coupled to an SSH VCN 1912 (e.g., SSH VCN 1612 of Figure 16) via the LPG 1910 included in the SSH VCN 1912. SSH VCN 1912 can include SSH subnet 1914 (e.g., SSH subnet 1614 in FIG. 16 ), and SSH VCN 1912 can be communicatively coupled to control plane VCN 1916 (e.g., control plane VCN 1616 in FIG. 16 ) via LPG 1910 included in control plane VCN 1916 and to data plane VCN 1918 (e.g., data plane 1618 in FIG. 16 ) via LPG 1910 included in data plane VCN 1918. Control plane VCN 1916 and data plane VCN 1918 can be included in service tenancy 1919 (e.g., service tenancy 1619 in FIG. 16 ).
[0244] The control plane VCN 1916 may include a control plane DMZ layer 1920 (e.g., control plane DMZ layer 1620 of FIG. 16 ) that may include a LB subnet 1922 (e.g., LB subnet 1622 of FIG. 16 ), a control plane app layer 1924 (e.g., control plane app layer 1624 of FIG. 16 ) that may include an app subnet 1926 (e.g., app subnet 1626 of FIG. 16 ), and a control plane data layer 1928 (e.g., control plane data layer 1628 of FIG. 16 ) that may include a DB subnet 1930 (e.g., DB subnet 1830 of FIG. 18 ). LB subnet 1922 included in control plane DMZ tier 1920 can be communicatively coupled to app subnet 1926 included in control plane app tier 1924 and to an Internet gateway 1934 (e.g., Internet gateway 1634 in FIG. 16 ) that may be included in control plane VCN 1916, and app subnet 1926 can be communicatively coupled to DB subnet 1930 included in control plane data tier 1928 and to service gateway 1936 (e.g., service gateway in FIG. 16 ) and network address translation (NAT) gateway 1938 (e.g., NAT gateway 1638 in FIG. 16 ). Control plane VCN 1916 can include service gateway 1936 and NAT gateway 1938.
[0245] Data plane VCN 1918 may include a data plane app layer 1946 (e.g., data plane app layer 1646 in FIG. 16 ), a data plane DMZ layer 1948 (e.g., data plane DMZ layer 1648 in FIG. 16 ), and a data plane data layer 1950 (e.g., data plane data layer 1650 in FIG. 16 ). Data plane DMZ layer 1948 may include a LB subnet 1922, which may be communicatively coupled to a trusted app subnet 1960 (e.g., trusted app subnet 1860 in FIG. 18 ) and an untrusted app subnet 1962 (e.g., untrusted app subnet 1862 in FIG. 18 ) of data plane app layer 1946 included in data plane VCN 1918, as well as to an Internet gateway 1934. Trusted app subnet 1960 may be communicatively coupled to service gateway 1936 included in data plane VCN 1918, NAT gateway 1938 included in data plane VCN 1918, and DB subnet 1930 included in data plane data layer 1950. Untrusted app subnet 1962 may be communicatively coupled to service gateway 1936 included in data plane VCN 1918 and DB subnet 1930 included in data plane data layer 1950. Data plane data layer 1950 may include DB subnet 1930 that may be communicatively coupled to service gateway 1936 included in data plane VCN 1918.
[0246] The untrusted app subnet 1962 may include primary VNICs 1964(1)-(N) that can be communicatively coupled to tenant virtual machines (VMs) 1966(1)-(N) that reside within the untrusted app subnet 1962. Each tenant VM 1966(1)-(N) can execute code in a respective container 1967(1)-(N) and can be communicatively coupled to an app subnet 1926 that can be included in a data plane app layer 1946 that can be included in a container egress VCN 1968. Each secondary VNIC 1972(1)-(N) can facilitate communication between the untrusted app subnet 1962 included in the data plane VCN 1918 and the app subnet included in the container egress VCN 1968. The container egress VCN may include a NAT gateway 1938 that can be communicatively coupled to the public internet 1954 (e.g., public internet 1654 in FIG. 16 ).
[0247] The internet gateway 1934 included in the control plane VCN 1916 and the internet gateway 1934 included in the data plane VCN 1918 can be communicatively coupled to a metadata management service 1952 (e.g., metadata management system 1652 of FIG. 16 ), which can be communicatively coupled to the public internet 1954. The public internet 1954 can be communicatively coupled to a NAT gateway 1938 included in the control plane VCN 1916 and the NAT gateway 1938 included in the data plane VCN 1918. The service gateway 1936 included in the control plane VCN 1916 and the service gateway 1936 included in the data plane VCN 1918 can be communicatively coupled to cloud services 1956.
[0248] In some examples, the pattern illustrated by the architecture of block diagram 1900 in FIG. 19 may be considered an exception to the pattern illustrated by the architecture of block diagram 1800 in FIG. 18 and may be desirable for an IaaS provider's customers when the IaaS provider cannot communicate directly with the customer (e.g., in a disconnected region). The customer can access each of the containers 1967(1)-(N) contained in each customer's VMs 1966(1)-(N) in real time. The containers 1967(1)-(N) can be configured to call each of the secondary VNICs 1972(1)-(N) contained in the app subnet 1926 of the data plane app tier 1946, which can be contained in the container egress VCN 1968. The secondary VNICs 1972(1)-(N) can send the call to a NAT gateway 1938, which can send the call to the public Internet 1954. In this example, containers 1967(1)-(N) that a customer can access in real time can be isolated from control plane VCN 1916 and can be isolated from other entities included in data plane VCN 1918. Containers 1967(1)-(N) can also be isolated from resources of other customers.
[0249] In another example, a customer can use containers 1967(1)-(N) to invoke cloud service 1956. In this example, the customer can execute code in containers 1967(1)-(N) that requests a service from cloud service 1956. Containers 1967(1)-(N) can send the request to secondary VNICs 1972(1)-(N), which can send the request to a NAT gateway that can send the request to public internet 1954. Public internet 1954 can send the request to LB subnet 1922, which is included in control plane VCN 1916, via internet gateway 1934. In response to determining that the request is valid, the LB subnet can send the request to app subnet 1926, which can send the request to cloud service 1956 via service gateway 1936.
[0250] It should be understood that the IaaS architectures 1600, 1700, 1800, 1900 shown in the figures may include components other than those shown. Additionally, the illustrated embodiments are only some examples of cloud infrastructure systems that may incorporate embodiments of the present disclosure. In other embodiments, the IaaS system may have more or fewer components than those shown in the figures, may combine two or more components, or may have a different configuration or arrangement of components.
[0251] In some embodiments, the IaaS systems described herein may include a suite of application, middleware, and database service offerings that are self-service, subscription-based, elastically scalable, reliable, highly available, and securely delivered to customers. One example of such an IaaS system is Oracle Cloud Infrastructure (OCI), offered by the present assignee.
[0252] 20 illustrates an exemplary computer system 2000 upon which various embodiments may be implemented. System 2000 may be used to implement any of the computer systems described above. As shown, computer system 2000 includes a processing unit 2004 that communicates with multiple peripheral subsystems via a bus subsystem 2002. These peripheral subsystems may include a processing acceleration unit 2006, an I / O subsystem 2008, a storage subsystem 2018, and a communication subsystem 2024. Storage subsystem 2018 includes a tangible computer-readable storage medium 2022 and a system memory 2010.
[0253] Bus subsystem 2002 provides a mechanism for allowing the various components and subsystems of computer system 2000 to communicate with each other as intended. While bus subsystem 2002 is shown schematically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 2002 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus, which may be implemented as a Mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0254] Processing unit 2004, which may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers), controls the operation of computer system 2000. One or more processors may be included in processing unit 2004. These processors may include single-core or multi-core processors. In some embodiments, processing unit 2004 may be implemented as one or more independent processing units 2032 and / or 2034, with a single-core or multi-core processor included in each processing unit. In other embodiments, processing unit 2004 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0255] In various embodiments, processing unit 2004 may execute various programs according to program code and may maintain multiple simultaneously executing programs or processes. At any given time, some or all of the program code to be executed may reside in processor 2004 and / or in storage subsystem 2018. Through suitable programming, processor 2004 may provide the various functions described above. Computer system 2000 may further include a processing acceleration unit 2006, which may include a digital signal processor (DSP), a special purpose processor, etc.
[0256] The I / O subsystem 2008 may include user interface input devices and user interface output devices. User interface input devices may include a keyboard, a pointing device such as a mouse or trackball, a touchpad or touchscreen integrated into a display, a scroll wheel, a click wheel, a dial, buttons, switches, a keypad, a voice input device with a voice command recognition system, a microphone, and other types of input devices. User interface input devices may include, for example, a motion sensing and / or gesture recognition device such as a Microsoft Kinect® motion sensor that enables a user to control and interact with an input device such as a Microsoft Xbox® 360 game controller through a natural user interface using gestures and spoken commands. User interface input devices may also include an eye gesture recognition device such as a Google Glass® blink detector that detects eye activity from a user (e.g., "blinking" during picture taking and / or menu selection) and translates the eye gesture as input to an input device (e.g., Google Glass®). Additionally, the user interface input devices may include a voice recognition sensing device that allows a user to interact with a voice recognition system (e.g., Siri® Navigator) via voice commands.
[0257] User interface input devices may also include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads, and graphic tablets, as well as audio / visual devices such as speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser distance measuring devices, and eye-tracking devices. Furthermore, user interface input devices may include medical imaging input devices, such as computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, positron emission tomography (PET) scanners, and medical ultrasound scanners. User interface input devices may also include audio input devices, such as MIDI keyboards, digital musical instruments, and the like.
[0258] User interface output devices may include display subsystems, indicator lights, or non-visual displays such as audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), a liquid crystal display (LCD), or a plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 2000 to a user or another computer. For example, user interface output devices may include various display devices that visually convey text, graphics, and audio / video information, such as, but not limited to, monitors, printers, speakers, headphones, automobile navigation systems, plotters, audio output devices, and modems.
[0259] Computer system 2000 may include a storage subsystem 2018 that provides a tangible, non-transitory, computer-readable storage medium for storing software and data structures that provide the functionality of embodiments described in this disclosure. The software may include programs, code modules, instructions, scripts, etc., that, when executed by one or more cores or processors of processing unit 2004, provide the functionality described above. Storage subsystem 2018 may also provide a repository for storing data used in accordance with the present disclosure.
[0260] 20 , storage subsystem 2018 may include various components including a system memory 2010, a computer-readable storage medium 2022, and a computer-readable storage medium reader 2020. The system memory 2010 may store program instructions that are loadable and executable by the processing unit 2004. The system memory 2010 may also store data used during the execution of the instructions and / or data generated during the execution of the program instructions. A variety of different types of programs may be loaded into the system memory 2010, including, but not limited to, client applications, web browsers, mid-tier applications, relational database management systems (RDBMS), virtual machines, containers, etc.
[0261] The system memory 2010 may also store an operating system 2016. Examples of operating systems 2016 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including, but not limited to, various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS operating systems. In some implementations in which the computer system 2000 runs one or more virtual machines, the virtual machines, along with their guest operating systems (GOS), may be loaded into the system memory 2010 and executed by one or more processors or cores of the processing unit 2004.
[0262] The system memory 2010 may be configured differently depending on the type of computer system 2000. For example, the system memory 2010 may be volatile memory (such as random access memory (RAM)) and / or non-volatile memory (such as read-only memory (ROM), flash memory, etc.). Different types of RAM configurations may be provided, including static random access memory (SRAM), dynamic random access memory (DRAM), etc. In some implementations, the system memory 2010 may include a basic input / output system (BIOS), which contains the basic routines that help to transfer information between elements within the computer system 2000, such as during start-up.
[0263] Computer-readable storage medium 2022 can represent a remote, local, fixed, and / or removable storage device, plus storage medium, for temporarily and / or more permanently containing and storing computer-readable information used by computer system 2000, including instructions executable by processing unit 2004 of computer system 2000.
[0264] The computer-readable storage medium 2022 may include any suitable medium known or used in the art, including, but not limited to, storage media and communication media such as volatile and nonvolatile, removable and non-removable media, implemented in any method or technology for storing and / or transmitting information. This may include tangible computer-readable storage media such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, Digital Versatile Disk (DVD) or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage device, or other tangible computer-readable medium.
[0265] By way of example, the computer-readable storage medium 2022 may include a hard disk drive that reads from or writes to non-removable, non-volatile magnetic media, a magnetic disk drive that reads from or writes to removable, non-volatile magnetic disks, and an optical disk drive that reads from or writes to removable, non-volatile optical disks such as CD-ROMs, DVDs, and Blu-Ray® disks or other optical media. The computer-readable storage medium 2022 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD disks, digital video tapes, and the like. The computer-readable storage media 2022 can also include solid-state drives (SSDs) based on non-volatile memory such as flash memory-based SSDs, enterprise flash drives, solid-state ROM, volatile memory-based SSDs such as solid-state RAM, dynamic RAM, static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. Disk drives and their associated computer-readable media can provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for the computer system 2000.
[0266] Machine-readable instructions executable by one or more processors or cores of the processing unit 2004 may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium may include physically tangible memory or storage devices, including volatile and / or non-volatile memory storage devices. Examples of non-transitory computer-readable storage media include magnetic storage media (e.g., disks or tapes), optical storage media (e.g., DVDs, CDs), various types of RAM, ROM, or flash memory, hard drives, floppy drives, removable memory drives (e.g., USB drives), or other types of storage devices.
[0267] The communications subsystem 2024 provides an interface to other computer systems and networks. The communications subsystem 2024 serves as an interface for receiving data from the computer system 2000 and transmitting data from the computer system 2000 to other systems. For example, the communications subsystem 2024 may enable the computer system 2000 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 2024 may include a radio frequency (RF) transceiver component for accessing wireless voice and / or data networks (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (enhanced data rates for global evolution), WiFi (IEEE 802.11 family of standards, or other mobile communications technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some embodiments, the communications subsystem 2024 may provide wired network connectivity (e.g., Ethernet) in addition to or instead of a wireless interface.
[0268] In some embodiments, the communications subsystem 2024 may also receive incoming communications in the form of structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc., on behalf of one or more users who may use the computer system 2000.
[0269] By way of example, the communications subsystem 2024 may be configured to receive data feeds 2026 in real time from users of social networks and / or other communications services, such as web feeds such as Twitter® feeds, Facebook® updates, Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party sources.
[0270] Additionally, the communications subsystem 2024 may be configured to receive data in the form of continuous data streams, which may include event streams 2028 of real-time events and / or event updates 2030, which may be continuous or infinite in nature with no apparent end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0271] The communications subsystem 2024 may also be configured to output structured and / or unstructured data feeds 2026, event streams 2028, event updates 2030, etc. to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 2000.
[0272] The computer system 2000 can be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head-mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or any other data processing system.
[0273] Because the nature of computers and networks is constantly changing, the description of computer system 2000 shown in the figure is intended merely as a specific example. Many other configurations are possible, having more or fewer components than the system shown in the figure. For example, customized hardware may be used, and / or particular elements may be implemented in hardware, firmware, software (including applets), or a combination. Furthermore, connections to other computing devices, such as network input / output devices, may be employed. Based on the disclosure and teachings provided herein, one of ordinary skill in the art will recognize other manners and / or methods for implementing various embodiments.
[0274] The embodiments may be implemented by using a computer program product comprising a computer program / instructions which, when executed by a processor, cause the processor to perform any of the methods described in this disclosure.
[0275] While specific embodiments have been described, various modifications, variations, alternative constructions, and equivalents are encompassed within the scope of the present disclosure. The embodiments are not limited to operating in one specific data processing environment, but can freely operate in multiple data processing environments. Furthermore, while the embodiments have been described using a particular sequence of transactions and steps, it should be apparent to those skilled in the art that the scope of the present disclosure is not limited to the described sequence of transactions and steps. Various features and aspects of the above-described embodiments may be used individually or jointly.
[0276] Furthermore, while embodiments have been described using particular combinations of hardware and software, it should be recognized that other combinations of hardware and software are within the scope of the present disclosure. Embodiments may be implemented exclusively in hardware, exclusively in software, or using a combination thereof. Various processes described herein may be implemented on the same processor or on any combination of different processors. Thus, when a component or service is described as being configured to perform certain operations, such configuration may be achieved, for example, by designing electronic circuitry to perform the operations, by programming a programmable electronic circuit (such as a microprocessor) to perform the operations, or any combination thereof. Processes may communicate using various techniques, including, but not limited to, conventional techniques for inter-process communication, and different pairs of processes may use different technologies, or the same pair of processes may use different technologies at different times.
[0277] Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. It will be apparent, however, that additions, subtractions, deletions, and other modifications and changes may be made without departing from the broader spirit and scope as set forth in the appended claims. Accordingly, while specific disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are within the scope of the following claims.
[0278] In the context of describing the disclosed embodiments (particularly in the context of the claims which follow), use of the terms "a," "an," and "the" and similar referents should be construed to encompass both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The terms "comprising," "having," "including," and "containing" should be construed as open-ended terms (i.e., meaning "including, but not limited to"), unless otherwise noted. The term "connected" should be construed as contained within, attached to, or joined to one another, either partially or as a whole, even if there is intervening material. The recitation of ranges of values herein, unless otherwise indicated herein, is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, and each separate value is incorporated herein as if it were individually recited herein. Unless otherwise indicated herein or clearly contradicted by context, all methods described herein can be performed in any suitable order. Any and all examples provided herein, or the use of exemplary language (e.g., "etc.") are intended merely to clarify the embodiments and do not limit the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure.
[0279] Disjunctive language, such as the phrase "at least one of X, Y, or Z," is intended to be understood in context as generally used to indicate that an item, term, etc. may be either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z), unless specifically stated otherwise. Thus, such disjunctive language is not intended to, and should not, generally imply that some embodiments require that at least one of X, at least one of Y, or at least one of Z, respectively, be present.
[0280] Preferred embodiments of the present disclosure are described herein, including the best mode known for carrying out the disclosure. Variations of these preferred embodiments will become apparent to those skilled in the art upon reading the foregoing description. Such variations will be readily apparent to those skilled in the art, and the present disclosure may be practiced in ways other than as specifically described herein. Accordingly, this disclosure includes all modifications and equivalents of the subject matter recited in the claims appended hereto as permitted by applicable law. Moreover, unless otherwise indicated herein, any combination of the above-described elements in all possible variations thereof is encompassed by the present disclosure.
[0281] Example embodiments of the present disclosure can be described in light of the following provisions. Clause 1. A method is disclosed. The method may include a computer system identifying a respective set of workloads executing on each of a plurality of hosts. The method may include the computer system identifying a plurality of response levels that specify applicability of a respective set of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of a first set of curtailment actions to a plurality of resources. The method may include determining a first estimate for power reduction resulting from the first response level based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) applicability of the first set of curtailment actions on the plurality of hosts according to the first response level. The method may include selecting a first response level from the plurality of response levels based at least on the first estimate for power reduction resulting from the first response level. The method may include applying the first set of curtailment actions to the plurality of hosts according to the selected first response level.
[0282] Clause 2. The method of clause 1 may further include determining a second estimate of a power reduction resulting from a second response level of the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of a second set of curtailment actions at the plurality of hosts according to the second response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on the first estimate of a power reduction resulting from the first response level.
[0283] Clause 3. The method of clause 1 or clause 2, wherein determining a first estimate of power reduction resulting from the first response level includes 1) determining that a first set of reduction actions is applicable to a first subset of the plurality of hosts; 2) identifying respective sets of workloads executing on each of the first subset of hosts; and 3) determining respective power consumptions of the respective sets of workloads executing on each of the first subset of hosts.
[0284] Clause 4. The method of clause 3, wherein determining a first estimate of a power reduction resulting from the first response level further includes determining a sum of the power consumptions of the respective sets of workloads executing on each of the first subset of hosts as the first estimate of a power reduction resulting from the first response level.
[0285] Clause 5. The method of clause 3, wherein determining a first estimate for power reduction resulting from the first response level further includes: 1) determining an estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts after application of the first set of reduction actions to the first subset of hosts; and 2) determining the first estimate for power reduction resulting from the first response level based on (a) the estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts and (b) the estimated power consumption for each of the respective sets of workloads executing on each of the first subset of hosts after application of the first set of reduction actions to the first subset of hosts.
[0286] Clause 6. The method of any of clauses 1-5, further comprising: 1) determining a difference between current values of aggregate power consumption of the plurality of hosts and current values of aggregate power thresholds of the plurality of hosts; and 2) determining that a first value for a power reduction resulting from a first response level is greater than the difference. In some embodiments, selecting the first response level is based at least on determining that the first value for a power reduction resulting from the first response level is greater than the difference.
[0287] Clause 7. The method of any of clauses 1 through 6, wherein the respective sets of workloads executing on each of the plurality of hosts and the plurality of response levels specifying applicability of the respective sets of curtailment actions to the plurality of hosts are identified based, at least in part, on at least one of: 1) identifying a degradation or failure of a thermal control system for a physical environment including the plurality of hosts; or 2) identifying a governmental power supply curtailment; or 3) identifying an increase in external temperature.
[0288] Clause 8. A method is disclosed. The method may include a computer system identifying a respective set of workloads executing on each of a plurality of hosts. The method may include the computer system identifying a plurality of response levels that specify applicability of a respective set of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of a first set of curtailment actions to a plurality of resources such that a first curtailment action of the first set of curtailment actions is applicable to a first resource of the plurality of resources. The method may include determining a first estimate of a predicted impact resulting from the first response level based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions on the plurality of hosts according to the first response level. In some embodiments, the impact on the workloads is determined based on at least one of a priority of the hosts, or a number of affected hosts, a number of affected workloads, a priority level of the affected workloads, a number of affected customers, or a priority level of the affected customers. The method may include selecting a first response level from a plurality of response levels based at least on a first estimate of an impact on the workload resulting from the first response level, and applying a first set of reduction actions to the plurality of hosts according to the selected first response level.
[0289] Clause 9. The method of clause 8 further includes determining a second estimate of the impact on the workload resulting from a second response level of the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of a second set of reduction actions at the plurality of hosts according to the second response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on the first estimate of the impact on the workload resulting from the first response level and the second estimate of the impact on the workload resulting from the second response level.
[0290] Clause 10. The method of clause 9, wherein the first set of abatement actions is more severe than the second set of abatement actions, and the first estimate of the workload impact resulting from the first response level is less than the second estimate of the workload impact resulting from the second response level.
[0291] Clause 11. The method of any of clauses 8 through 10, wherein determining a first estimate of an impact on workloads resulting from a first response level includes: 1) determining that a first set of reduction actions is applicable to a first subset of the plurality of hosts; 2) identifying a respective set of workloads executing on each of the first subset of hosts as workloads affected by the first response level; and 3) determining at least one of (a) a number of workloads affected by the first response level or (b) a priority type of the workloads affected by the first response level.
[0292] Clause 12. In any of the methods of clauses 8 to 11, determining a first estimate of an impact on workloads resulting from the first response level includes: 1) determining that a first set of reduction actions is applicable to a first subset of the plurality of hosts; 2) identifying respective sets of workloads executing on each of the first subset of hosts; 3) identifying respective customers for the respective sets of workloads executing on each of the first subset of hosts as customers affected by the first response level; and 4) determining at least one of (a) a number of customers affected by the first response level, or (b) a priority type of the customers affected by the first response level.
[0293] Clause 13. The method of any of clauses 8 through 12, further including determining a second estimate of a power reduction resulting from the first response level based at least on (a) a respective set of workloads executing on each of the plurality of hosts and (b) applicability of the first set of reduction actions at the plurality of hosts according to the first response level. In some embodiments, selecting a first response level from the plurality of response levels is based at least on a first estimate of an impact on the workload resulting from the first response level and a second estimate of a power reduction resulting from the first response level.
[0294] Clause 14. The method of clause 13, wherein the selected first response level is a response level among the plurality of response levels associated with the smallest workload impact that achieves a power reduction equal to or greater than the difference between the current value of the aggregate power consumption of the plurality of hosts and the current value of the aggregate power threshold of the plurality of hosts.
[0295] Clause 15. A method is disclosed. The method may include a computer system determining a respective predicted set of workloads executing on each of a plurality of hosts during a future time period. The method may include the computer system identifying a plurality of response levels specifying applicability of a respective set of curtailment actions to the plurality of hosts. In some embodiments, a first response level of the plurality of response levels specifies applicability of the first set of curtailment actions to a plurality of resources. The method may include determining a first estimate for power reduction resulting from the first response level based at least on (a) the respective predicted set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions on the plurality of hosts according to the first response level. The method may include selecting a first response level from the plurality of response levels based at least on the first estimate for power reduction resulting from the first response level. The method may include identifying one or more workloads (a) currently executing on the plurality of hosts and (b) that will be affected by applying the first set of curtailment actions to the plurality of hosts according to the selected first response level. The method may include preemptively migrating affected workloads from a plurality of hosts to one or more other hosts in advance of a future time period.
[0296] Clause 16. The method of clause 15, wherein determining the respective predicted sets of workloads executing on each of the plurality of hosts during the future time period is based on historical patterns of workloads executing on the plurality of hosts.
[0297] Clause 17. The method of clause 15 or clause 16, wherein determining the respective predicted set of workloads executing on each of the plurality of hosts during the future time period is performed at least in part using a machine learning model pre-trained using a supervised learning algorithm to predict the set of workloads based, at least in part, on historical workload data provided as training data.
[0298] Clause 18. The method of clauses 15-17 further includes the computer system obtaining a forecasted value of aggregate power consumption of the plurality of hosts over a future time period. The method may further include the computer system obtaining a forecasted value of an aggregate power threshold for the plurality of hosts over a future time period. The method may further include selecting from a plurality of response levels in response to determining that the forecasted value of the aggregate power consumption exceeds the forecasted value of the aggregate power threshold.
[0299] Clause 19. The method of clauses 15-18 further includes the computer system identifying predicted failures of temperature control systems associated with the plurality of hosts over a future time period. The method may further include the computer system obtaining a predicted value of an aggregate power threshold for the plurality of hosts over a future time period based at least in part on identifying the predicted failures. The method may further include selecting from a plurality of response levels in response to identifying the predicted failures.
[0300] Clause 20. The method of clause 19, wherein identifying predicted faults in the temperature control system utilizes a machine learning model pre-trained using training data including historical data associated with a plurality of hosts, and the machine learning model is trained using a supervised learning algorithm to identify predicted faults from the input data.
[0301] Clause 21. The method of clauses 15-20, wherein the historical data in the training data includes temperature control system data, historical power consumption data corresponding to the set of hosts, and historical aggregate power thresholds.
[0302] Clause 22. A system is disclosed. The system may include a memory configured to store instructions and one or more processors configured to execute the instructions to perform a method according to any of clauses 1 to 21.
[0303] Clause 23. A non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium can store instructions that, when executed by a processor, cause the processor to perform the method of any of clauses 11 to 21.
[0304] All references cited herein, including publications, patent applications, and patents, are incorporated by reference to the same extent as if each individual reference were individually and specifically indicated to be incorporated by reference and were set forth in its entirety herein.
[0305] While the foregoing specification has described aspects of the disclosure with reference to specific embodiments thereof, those skilled in the art will recognize that the disclosure is not limited thereto. Various features and aspects of the above-described disclosure can be used individually or jointly. Moreover, the embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings should be regarded as illustrative rather than restrictive.
Claims
1. 1. A method comprising: a computer system identifying a respective set of workloads executing on each of a plurality of hosts; identifying a plurality of response levels that specify applicability of respective sets of abatement actions to a plurality of hosts; a first response level of the plurality of response levels specifying applicability of a first set of curtailment actions to the plurality of resources, the method further comprising: determining a first estimate for power reduction resulting from the first response level based at least on (a) the respective sets of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions at the plurality of hosts according to the first response level; selecting the first response level from the plurality of response levels based at least on the first estimate of the power reduction resulting from the first response level; applying the first set of reduction actions to the plurality of hosts according to the selected first response level.
2. determining a second estimate for the power reduction resulting from a second response level of the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) applicability of a second set of curtailment actions at the plurality of hosts according to the second response level; The method of claim 1 , wherein selecting the first response level from the plurality of response levels is based at least on the first estimate of the power reduction resulting from the first response level.
3. Determining the first estimate for the power reduction resulting from the first response level includes: determining that the first set of reduction actions is applicable to a first subset of the plurality of hosts; identifying a respective set of the workloads executing on each of a first subset of the hosts; and determining a respective power consumption of a respective set of the workloads executing on each of the first subset of hosts.
4. Determining the first estimate for the power reduction resulting from the first response level includes:
4. The method of claim 3, further comprising determining a sum of the respective power consumptions of the respective sets of the workloads executing on each of the first subset of the hosts as the first estimate for the power reduction resulting from the first responsiveness level.
5. Determining the first estimate for the power reduction resulting from the first response level includes: determining an estimated power consumption for each of the respective set of workloads executing on each of the first subset of hosts after application of the first set of curtailment actions to the first subset of hosts; 4. The method of claim 3, further comprising: determining the first estimate for the power reduction resulting from the first response level based on: (a) the respective power consumption of the respective set of workloads executing on each of the first subset of hosts; and (b) the respective estimated power consumption of the respective set of workloads executing on each of the first subset of hosts after application of the first set of reduction actions to the first subset of hosts.
6. determining a difference between a current value of an aggregate power consumption of the plurality of hosts and a current value of an aggregate power threshold of the plurality of hosts; determining that the first value for the power reduction resulting from the first response level is greater than the difference; 6. The method of claim 1, wherein selecting the first response level is based at least on determining that the first value for the power reduction resulting from the first response level is greater than the difference.
7. The respective sets of workloads executing on each of the plurality of hosts and the plurality of response levels specifying applicability of the respective sets of reduction actions to the plurality of hosts are at least partially determined by: Identifying a degradation or failure of a temperature control system for a physical environment containing said plurality of hosts; or Identifying government cuts in electricity supply; or The method according to any one of claims 1 to 6, wherein the determination is based on at least one of: determining an increase in external temperature.
8. 1. A method comprising: a computer system identifying a respective set of workloads executing on each of a plurality of hosts; the computer system identifying a plurality of response levels that specify applicability of respective sets of abatement actions to the plurality of hosts; a first response level of the plurality of response levels specifying applicability of a first set of abatement actions to a plurality of resources such that a first abatement action of the first set of abatement actions is applicable to a first resource of the plurality of resources, the method further comprising: determining a first estimate of a predicted impact resulting from the first response level based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of reduction actions on the plurality of hosts according to the first response level; The impact on the workloads is determined based on at least one of a host priority, or a number of affected hosts, a number of affected workloads, a priority level of the affected workloads, a number of affected customers, or a priority level of the affected customers, and the method further comprises: selecting the first response level from the plurality of response levels based at least on the first estimate of an impact on the workload resulting from the first response level; applying the first set of reduction actions to the plurality of hosts according to the selected first response level.
9. determining a second estimate of an impact on the workloads resulting from a second response level of the plurality of response levels based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of a second set of curtailment actions at the plurality of hosts according to the second response level; 9. The method of claim 8, wherein selecting the first response level from the plurality of response levels is based at least on a first estimate of an impact on the workload resulting from the first response level and a second estimate of an impact on the workload resulting from the second response level.
10. 10. The method of claim 9, wherein the first set of abatement actions is more severe than the second set of abatement actions, and the first estimate of the impact to the workload resulting from the first response level is less than the second estimate of the impact to the workload resulting from the second response level.
11. Determining the first estimate of an impact on the workload resulting from the first response level includes: determining that the first set of reduction actions is applicable to a first subset of the plurality of hosts; identifying a respective set of the workloads executing on each of a first subset of the hosts as the affected workloads resulting from the first response level; and determining at least one of (a) the number of the affected workloads resulting from the first response level, or (b) the priority type of the affected workloads resulting from the first response level.
12. Determining the first estimate of an impact on the workload resulting from the first response level includes: determining that the first set of reduction actions is applicable to a first subset of the plurality of hosts; identifying a respective set of the workloads executing on each of a first subset of the hosts; identifying each customer for each set of workloads executing on each of the first subset of hosts as the affected customers resulting from the first response level; and determining at least one of: (a) the number of affected customers resulting from the first response level; or (b) the priority type of the affected customers resulting from the first response level.
13. determining a second estimate for power reduction resulting from the first response level based at least on (a) the respective set of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions at the plurality of hosts according to the first response level; 13. The method of claim 8, wherein selecting the first response level from the plurality of response levels is based at least on the first estimate of an impact on the workload resulting from the first response level and the second estimate of the power reduction resulting from the first response level.
14. 14. The method of claim 13, wherein the selected first response level is a response level among the plurality of response levels associated with the smallest workload impact that achieves a power reduction equal to or greater than a difference between a current value of aggregate power consumption of the plurality of hosts and a current value of an aggregate power threshold of the plurality of hosts.
15. 1. A method comprising: determining a respective projected set of workloads executing on each of a plurality of hosts during a future time period; the computer system identifying a plurality of response levels that specify applicability of respective sets of abatement actions to the plurality of hosts; a first response level of the plurality of response levels specifying applicability of a first set of curtailment actions to a plurality of resources, the method further comprising: determining a first estimate for power reduction resulting from the first response level based at least on (a) the respective projected sets of workloads executing on each of the plurality of hosts and (b) the applicability of the first set of curtailment actions at the plurality of hosts according to the first response level; selecting the first response level from the plurality of response levels based at least on the first estimate of the power reduction resulting from the first response level; (a) identifying one or more workloads currently executing on the plurality of hosts and (b) that would be affected by applying the first set of curtailment actions to the plurality of hosts in accordance with the selected first response level; preemptively migrating the affected workload from the plurality of hosts to one or more other hosts prior to the future period of time.
16. 16. The method of claim 15, wherein determining the respective projected set of workloads executing on each of a plurality of hosts during the future time period is based on historical patterns of workloads executing on the plurality of hosts.
17. Determining a predicted set of each of the workloads executing on each of the plurality of hosts during the future time period is performed, at least in part, using a machine learning model that has been pre-trained using a supervised learning algorithm to predict a set of workloads based on historical workload data provided as training data.
17. The method of claim 15 or claim 16.
18. obtaining a forecast of aggregate power consumption of the plurality of hosts over the future time period; obtaining a forecast of aggregate power thresholds of the plurality of hosts over the future time period; 18. The method of claim 15, further comprising selecting from the plurality of response levels in response to determining that the predicted value of the aggregate power consumption exceeds the predicted value of the aggregate power threshold.
19. the computer system identifying predicted failures of temperature control systems associated with the plurality of hosts over a future time period; obtaining, by the computer system, a forecast of aggregate power thresholds of the plurality of hosts for the future time period based at least in part on identifying the predicted failure; The method of any of claims 15 to 18, further comprising selecting from the plurality of response levels in response to identifying the predicted fault.
20. 20. The method of claim 19, wherein identifying the predicted fault of the temperature control system utilizes a machine learning model that is pre-trained using training data that includes historical data associated with the plurality of hosts, the machine learning model being trained using a supervised learning algorithm to identify the predicted fault from input data.
21. The method of any of claims 15 to 20, wherein the historical data of the training data includes temperature control system data, historical power consumption data corresponding to a set of hosts, and historical aggregate power thresholds.
22. a memory configured to store instructions; and one or more processors configured to execute the instructions to perform the method of any of claims 1 to 21.
23. A non-transitory computer readable medium having stored thereon instructions that, when executed by a processor, cause the processor to perform the method of any of claims 1 to 21.