Throughput Optimized Service Quality Warning Power Capping System
By receiving and comparing the power measurements of multiple machines in the data center, determining whether to remove the consumed power, and commanding the machine to reduce the power consumption, the problem of difficult to achieve fine-grained power capping in the co-located cluster of throughput and delay-sensitive jobs in the prior art is solved, and effective power management and performance protection for throughput-oriented jobs and delay-sensitive jobs are achieved.
Patent Information
- Application Number
- CN202110586000.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-04-29
- Filing Date
- 2021-05-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-05-27
AI Technical Summary
The prior art is difficult to effectively achieve power capping in clusters co-located with throughput and latency-sensitive workloads, resulting in a significant impact on throughput-oriented tasks, and latency-sensitive tasks are difficult to avoid power capping.
By receiving power measurements from multiple machines in the data center, identifying the power limit of the machine, and determining whether to remove the consumed power based on the comparison results, the machine is ordered to operate according to the limit to reduce power consumption, achieving a fine-grained power cap.
This enables CPU share throttling for throughput-oriented workloads without affecting latency sensitive jobs, reducing power consumption, avoiding power overloads, while ensuring minimal system availability and performance impact.
Smart Images

Figure CN113312235B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of the filing date of U.S. Provisional Patent Application No. 63 / 030,639, filed May 27, 2020, the disclosure of which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to a throughput optimized quality of service early warning power capping system. Background Art
[0004] Data centers form the backbone of popular online services such as search, streaming video, email, social networking, online shopping, and the cloud. The growing demand for online services forces hyperscale providers to invest large amounts of money to continuously expand their data center fleets. Much of this investment is allocated to purchasing and building infrastructure such as buildings, power delivery, and cooling to host the servers that make up warehouse-scale computers.
[0005] Power oversubscription is the practice of deploying more servers in a data center than the data center nominally supports if all servers are 100% utilized. It can increase the power capacity of existing data centers in situ and reduce future data center builds. On the other hand, power oversubscription carries the risk of overload during power spikes, so it is usually accompanied by protection systems such as power capping. Power capping systems enable safe power oversubscription by preventing overload during power emergencies. Power capping actions include pausing low-priority tasks, throttling CPU voltage and frequency using techniques such as dynamic voltage and frequency scaling (DVFS) and running average power limit (RAPL), or packing threads in a subset of available cores. The action needs to be compatible with the workload and meet service level objectives (SLOs). However, this is a challenge for clusters where throughput-oriented workloads are co-located with latency-sensitive workloads.
[0006] Throughput-oriented tasks represent an important class of computational workloads. Examples include web indexing, log processing, and machine learning model training. These workloads have deadlines for when the computation needs to be completed, often measured in hours, which makes them ideal candidates for performance throttling when the cluster faces power contingencies due to power oversubscription. However, missing deadlines can result in serious consequences, such as lost revenue and degraded quality, making them less susceptible to interruption.
[0007] Latency-sensitive workloads are another category. They need to complete the requested computation within milliseconds to seconds. A typical example is a job processing user requests. High latency leads to a poor user experience, ultimately resulting in loss of users and revenue. Unlike throughput-oriented pipelines, such jobs are not suitable for performance throttling. They are usually considered high priority and need to be exempt from power capping. Throughput-oriented and latency-sensitive jobs are often co-located on the same server to improve resource utilization. This poses a huge challenge to power capping as a fine-grained power capping mechanism is required. Summary of the invention
[0008] One aspect of the present disclosure provides a method, including: receiving, by one or more processors, power measurements of a plurality of machines in a data center, the plurality of machines performing one or more tasks; identifying power limits of the plurality of machines; comparing, by the one or more processors, the received power measurements with the power limits of the plurality of machines; determining, based on the comparison, whether to shed power consumed by the plurality of machines; and commanding, by the one or more processors, the plurality of machines to operate according to the one or more limits to reduce power consumption.
[0009] According to some examples, the method may further include: identifying a first threshold; and identifying a second threshold that is higher than the first threshold; wherein determining whether to remove power includes determining whether the received power measurement meets or exceeds the second threshold, and wherein commanding the plurality of machines to operate according to the one or more limits includes sending a command to reduce power consumption by a first predetermined percentage. Determining whether to remove power may include determining whether the received power measurement exceeds the first threshold only during a predetermined time period after the second threshold has been met or exceeded. When the first threshold is met or exceeded during the predetermined time period, the method may further include sending a second command to the plurality of machines to reduce power consumption by a second predetermined percentage that is lower than the first predetermined percentage.
[0010] According to some examples, measurements may be received from one or more power meters coupled to a plurality of machines.
[0011] Commanding the plurality of machines to operate in accordance with the one or more limits may include: for all machines in the power domain, limiting the processing time of tasks within a machine-level scheduler. Limiting the processing time may include applying a first multiplier to low priority tasks, wherein the first multiplier is selected to prevent the low priority tasks from running and consuming power. In some examples, the method may further include applying a second multiplier to the low priority tasks. One or more high priority tasks may be exempted from the limit.
[0012] According to some examples, the method may further include limiting a number of available schedulable processor entities to control power when the power measurement becomes unavailable.
[0013] Commanding the plurality of machines to operate in accordance with the one or more constraints may include making a portion of a processing component of each of the plurality of machines unavailable for tasks.
[0014] Another aspect of the present disclosure provides a system including one or more processors in communication with a plurality of machines performing one or more tasks. The one or more processors may be configured to receive power measurements of the plurality of machines; identify power limits of the plurality of machines; compare the received power measurements with the power limits of the plurality of machines; determine whether to remove power consumed by the plurality of machines based on the comparison; and command the plurality of machines, by the one or more processors, to operate according to one or more limits to reduce power consumption.
[0015] The one or more processors may be further configured to: identify a first threshold; and identify a second threshold that is higher than the first threshold; wherein determining whether to remove power includes determining whether the received power measurement meets or exceeds the second threshold, and wherein commanding the plurality of machines to operate according to the one or more limits includes sending a command to reduce power consumption by a first predetermined percentage. Determining whether to remove power may include determining whether the received power measurement exceeds the first threshold only during a predetermined time period after the second threshold has been met or exceeded. When the first threshold is met or exceeded during the predetermined time period, the one or more processors may be further configured to send a second command to the plurality of machines to reduce power consumption by a second predetermined percentage that is lower than the first predetermined percentage.
[0016] According to some examples, measurements may be received from one or more power meters coupled to a plurality of machines.
[0017] Commanding the plurality of machines to operate in accordance with the one or more limits may include: for all machines in the power domain, limiting a processing time for a task within a machine-level scheduler. In limiting the processing time, the one or more processors may be configured to apply a first multiplier to low priority tasks, wherein the first multiplier is selected to prevent the low priority tasks from running and consuming power. The one or more processors may be configured to exempt one or more high priority tasks from the limit.
[0018] In some examples, commanding the plurality of machines to operate in accordance with the one or more constraints may include making a portion of a processing component of each of the plurality of machines unavailable for the task.
[0019] Another aspect of the present disclosure provides a non-transitory computer-readable medium storing instructions executable by one or more processors to perform a method, including: receiving power measurements of multiple machines in a data center, the multiple machines performing one or more tasks; identifying power limits of the multiple machines; comparing the received power measurements with the power limits of the multiple machines; based on the comparison, determining whether to remove power consumed by the multiple machines; and commanding the multiple machines to operate according to one or more limits to reduce power consumption. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1A is a block diagram illustrating an exemplary software architecture of a system according to aspects of the present disclosure.
[0021] Figure 1B is a block diagram illustrating an exemplary system according to aspects of the present disclosure.
[0022] Figure 2A -B is a diagram illustrating an exemplary load shaping strategy according to aspects of the present disclosure.
[0023] Figure 3A -C illustrates exemplary results of running a workload according to aspects of the present disclosure.
[0024] Figure 4 A graph is provided that illustrates throttling triggered by various combinations of parameters in accordance with aspects of the present disclosure.
[0025] Figure 5 Graphs illustrating total power consumption and CPU usage according to aspects of the present disclosure are provided.
[0026] Figure 6 is a flow chart illustrating an exemplary method according to aspects of the present disclosure. DETAILED DESCRIPTION
[0027] The present disclosure provides a system and method for throttling the CPU shares of throughput-oriented workloads so that they are slowed down enough to keep power under a specified budget without affecting latency-sensitive jobs. The system architecture enables oversubscription over large power domains. Power pooling and statistical multiplexing across machines in large power domains maximize the possibility of power oversubscription. Moreover, the system implements mechanisms and strategies for power throttling with broad applicability. The system can be designed to rely only on established Linux kernel functions and is therefore independent of the hardware platform. This allows the flexible introduction of various platforms into the data center without compromising the efficiency of power capping.
[0028] Task-level mechanisms allow for differentiated Quality of Service (QoS). Specifically, the system does not impact the ability to serve latency-sensitive workloads collocated with throughput-oriented workloads, and has the ability to apply different CPU caps to workloads with different SLOs. Platform independence and QoS alerting features allow the system to be customized for a wide range of hardware platforms and software applications.
[0029] The advantage of this system is that it achieves power safety while minimizing performance impact. The two-threshold scheme enables minimum performance impact while ensuring that a large amount of power can be removed to avoid power overload in emergency situations. The system focuses on availability and introduces a failover subsystem to maintain power safety guarantees in the face of power telemetry failures.
[0030] The system is able to perform two end-to-end power capping actuations tailored for throughput-oriented workloads: a primary mechanism called reactive capping, and a failover mechanism called proactive capping. Reactive capping monitors real-time power signals read from power meters and reacts to high power measurements by throttling workloads. When the power signal becomes unavailable (e.g., due to meter downtime), proactive capping takes over and assesses the risk of a circuit breaker trip. The assessment is based on factors such as recent power and how long the signal has been unavailable. If the risk is deemed high, it proactively throttles tasks.
[0031] Architecture and implementation
[0032] Figure 1A An exemplary software architecture of the system is illustrated. As shown, the architecture includes a meter observer module 124, a power notifier module 126, a risk assessor module 122, and a machine manager module 128. The meter observer module 144 polls the power readings from the meters 114. For example, the meter observer module 144 may poll at a rate of one reading per second, multiple readings per second, one reading every few seconds, etc. The meter observer module 144 passes the readings to the power notifier module 126 and also stores a copy in the power history database 112.
[0033] The power notifier module 126 is the central module that implements the control logic for reactive and proactive capping. When power readings are available, it uses the readings for reactive capping logic. When readings are not available, it queries the risk assessor module 122 for proactive capping logic.
[0034] The risk assessor module 122 uses the historical power information from the power history database 112 to assess the risk of the circuit breaker tripping. If any of the logic determines capping, the power notifier module 126 passes the appropriate capping parameters to the machine manager module 128. For example, the power capping parameters may be received from the power limit database 116. In response, the machine manager module 128 sends a load shaping request to one or more node controllers 130 associated with each machine. The load shaping request may be sent in the form of a remote procedure call (RPC) or any other message or communication format. According to some examples, the RPC may be sent to multiple node controllers 130 simultaneously. The RPC may include a command to cause each machine to reduce power.
[0035] According to some examples, a load shaping request may include fields such as a maximum priority of the job to be throttled, a multiplier for hardcapping the usage of lower priority jobs, and a duration for capping jobs. The request may also include instructions whether the multiplier should be applied only to the CPU limit or to the sustained CPU rate.
[0036] When a new load shaping request is received, the node controller 130 checks whether there is already an ongoing load shaping event. If there is no ongoing event, the new request may be accepted. If there is already an ongoing event, the new request may be accepted if it is more aggressive in terms of its power reduction effect. For example, the new RPC may be more aggressive if it has a higher maximum throttling priority or a lower hard cap multiplier, or if a switch is made from applying the multiplier only from the CPU limit to applying to the sustained CPU rate.
[0037] The size of a power domain may vary from a few megawatts to tens of megawatts, depending on the power architecture of the data center. One instance of the system can be deployed for each protected power domain. The instance can be replicated for fault tolerance. For example, there are 4 replicas in a 2-master, 2-slave configuration. The master replica can read the power meter and issue a power removal RPC. The power removal RPC service is designed to be idempotent and can handle duplicate RPCs from different master replicas. Two identical master replicas can be used to ensure that power removal is available even during master election. When the master becomes unavailable, the slave takes over.
[0038] Figure 1BAn exemplary system is illustrated that includes a server device (such as controller 190) on which the system can be implemented. Controller 190 may include hardware configured to manage load shaping and throttling of devices in data center 180. According to one example, controller 190 may reside within and control a particular data center. According to other examples, controller 190 may be coupled to one or more data centers 180, such as through a network, and may manage the operations of multiple data centers.
[0039] Data center 180 may be located at a considerable distance from controller 190 and / or other data centers (not shown). Data center 180 may include one or more computing devices, such as processors, servers, shards, units, etc. In some examples, the computing devices in the data center may have different capabilities. For example, different computing devices may have different processing speeds, workloads, etc. Although only a few of these computing devices are shown, it should be understood that each data center 180 may include any number of computing devices, and the number of computing devices in a first data center may be different from the number of computing devices in a second data center. In addition, it should be understood that the number of computing devices in each data center 180 may change over time, for example, as hardware is removed, replaced, upgraded, or expanded.
[0040] In some examples, controller 190 can communicate with computing devices in data center 180 and can facilitate execution of programs. For example, controller 190 can track the capabilities, status, workload, or other information of each computing device and use such information to assign tasks. Controller 190 can include processor 198 and memory 192, including data 194 and instructions 196. In other examples, such operations can be performed by one or more computing devices in data center 180, and a separate controller can be omitted from the system.
[0041] Controller 190 may include processor 198, memory 192, and other components typically found in server computing devices. Memory 192 may store information accessible by processor 198, including instructions 196 that may be executed by processor 198. Memory may also include data 194 that may be retrieved, manipulated, or stored by processor 198. Memory 192 may be a non-transitory computer-readable medium capable of storing information accessible by processor 198, such as a hard drive, solid-state drive, tape drive, optical storage, memory card, ROM, RAM, DVD, CD-ROM, writeable, and read-only memory. Processor 198 may be a well-known processor or other lesser-known types of processors. Alternatively, processor 198 may be a dedicated controller, such as an ASIC.
[0042] The instructions 196 may be a set of instructions (such as machine code) directly executed by the processor 198 or a set of instructions (such as a script) indirectly executed by the processor 198. In this regard, the terms "instructions," "steps," and "programs" may be used interchangeably herein. The instructions 196 may be stored in an object code format for direct processing by the processor 198, or in other types of computer languages, including scripts or collections of independent source code modules that are interpreted on demand or compiled in advance.
[0043] Data 194 may be retrieved, stored, or modified by processor 198 in accordance with instructions 196. For example, although the systems and methods are not limited to a particular data structure, data 194 may be stored in a computer register, a relational database, as a table with a plurality of different fields and records, or an XML document. Data 194 may also be formatted in a computer readable format, such as, but not limited to, binary values, ASCII, or Unicode. In addition, data 194 may include information sufficient to identify the relevant information, such as numbers, descriptive text, proprietary codes, pointers, references to data stored in other storage, including other network locations, or information used by functions that compute the relevant data.
[0044] although Figure 1B The processor 198 and memory 192 are functionally shown as being within the same block, but the processor 198 and memory 192 may actually include multiple processors and memories that may or may not be stored within the same physical housing. For example, some instructions 196 and data 194 may be stored on a removable CD-ROM, while other instructions may be stored within a read-only computer chip. Some or all of the instructions and data may be stored in a location physically remote from the processor 198 but still accessible to the processor 198. Similarly, the processor 198 may actually include a collection of processors that may or may not operate in parallel.
[0045] Reactive capping using CPU bandwidth control
[0046] CPU utilization is a good proxy for the CPU power consumed by running tasks. According to some examples, the CPU bandwidth control function of the Linux(R) Completely Fair Scheduler (CFS) kernel scheduler can be used to precisely control the CPU utilization of tasks running in a node, thereby allowing control of the power consumed by the node.
[0047] Individual tasks may be running in their own control group (cgroup). The scheduler provides two parameters in the cgroup, namely quota and period. The quota controls the amount of CPU time a workload gets to run during a period of time. It is shared among the CPUs in the system. The quota and period can be set for each cgroup and are usually specified at millisecond granularity. A separate (per cgroup and per CPU), cumulative runtime_remaining variable is kept in the kernel. When a thread is running, the accumulated runtime_remaining is consumed. When it reaches zero, the running thread is descheduled and no thread in the cgroup can run until the runtime is replenished at the beginning of the next period. The historical CPU utilization of all workloads running in the machine can be tracked. During power events, each node in the power domain may receive RPC calls to throttle throughput-oriented workloads.
[0048] The RPC call contains parameters about how much the task's CPU usage should be reduced. The node controller that receives the RPC call uses the historical CPU usage of all throughput-oriented workloads to determine the throttled CPU time. Then, new quota and period values are calculated for each cgroup and written to the cgroup in the machine. QoS differentiation is achieved by grouping tasks of different priorities into different cgroups. In our implementation, throughput-oriented tasks are assigned lower priorities, while latency-sensitive tasks are assigned higher priorities. CPU bandwidth control is applied to cgroups of low-priority tasks, exempting cgroups of high-priority tasks.
[0049] CPU bandwidth control is advantageous over DVFS and RAPL in that it is platform independent and has per-cgroup control. To ensure that CPU bandwidth control achieves the expected power reduction, it can be used in conjunction with power metering and negative feedback. The relationship between CPU bandwidth throttling and power consumption remains monotonic, so the feedback loop can be stabilized. CPU jailing is another mechanism that can be used to limit power consumption during power metering. CPU jailing has a more predictable throttling-power relationship suitable for open-loop control.
[0050] Control strategy: load shaping
[0051] A variety of different load shaping control strategies are available. Figure 2AThe first strategy shown applies a multiplier when both the low and high thresholds are reached. For example, when the power reaches the low threshold, a soft multiplier will be applied. If the power continues to increase and reaches the high threshold, a hard multiplier will be applied.
[0052] The multiplier may be, for example, a number between 0 and 1 that is applied to reduce the power consumed by a single machine, such as by reducing the number of jobs executed by a single machine at a given time. For example, a multiplier close to 0 (such as 0.01) may quickly reduce power by reducing 99% of the tasks on a single machine at a given time. Such a multiplier may be referred to as a "hard" or "high-impact" multiplier. In contrast, a multiplier close to 1 (such as 0.9 or 0.75) may reduce power and also minimize performance impact by reducing a lower percentage of tasks executed by a single machine. Such a multiplier may be referred to as a "soft" or "low-impact" multiplier.
[0053] Multipliers may be applied to CPU limits, sustained CPU rates, the number of low priority jobs, etc. For example, a multiplier may limit the processing time of a task within a machine-level scheduler for all machines in a power domain. For example, a multiplier may limit the time during which a task may access the CPU. A multiplier may be applied only to lower priority tasks, while higher priority tasks may be exempted.
[0054] Figure 2B A second load shaping control strategy is illustrated that uses the power meter readings as input to determine when and how much to throttle the workload. A multiplier parameter is predetermined for the throttling action. When throttling begins, the node controller receives an RPC call containing the multiplier and sets a CPU cap for each affected task, which is calculated by the following formula:
[0055] Cap = multiplier * CPUavg
[0056] Where CPUavg is the average CPU usage of the task over the past 10 seconds. This cap is updated once per second throughout the period that throttling is active. The effect is that if a task consumes CPU up to the cap, the task's CPU cap decreases exponentially over time at a rate roughly equal to this multiplier. This is proportional to their CPU usage, fairly disadvantaging tasks.
[0057] This strategy aims to strike a balance between power safety and performance impact. It does this by maintaining two power thresholds, with the higher threshold close to the device limit for system protection, such as 98% of the limit, and the lower threshold, such as at 96% of the limit. The higher threshold is associated with a "hard" or "high impact" multiplier close to 0 (such as 0.01), which is intended to quickly reduce power consumption to ensure safety. The lower threshold is associated with a "soft" or "low impact" multiplier close to 1, such as 0.9 or 0.75, which is intended to minimize performance impact even if power may temporarily exceed the lower threshold.
[0058] According to some examples, the lower threshold is not activated until throttling is triggered, and is deactivated when throttling is active for a period of time. In this way, power is allowed to reach a range between the two thresholds without throttling. More specifically, before any throttling occurs, the higher threshold is active and the lower threshold is not active. The first throttling event occurs only when power reaches the higher threshold. Once power reaches the higher threshold, the lower threshold is activated and expires after a specified duration when power is below the lower threshold and no throttling occurs.
[0059] When power drops below the minimum activity threshold, throttling stops. If all tasks are unthrottled at the same time, power may surge and quickly reach the threshold again, causing unnecessary oscillations, or even reach dangerous levels before load shaping takes effect. To avoid these undesirable effects, the system is designed to stop in a gradual manner. Each machine is assigned a random throttling timeout in the range [timeout_min, timeout_max]. When throttling stops, the machines will be gradually unthrottled according to their timeouts. This makes the power rise curve smooth. In an exemplary embodiment, an RPC call can be sent every second to extend the timeout until the power is below the threshold. A repeatedly refreshed timeout can be used instead of a stop throttling RPC, because the RPC may be lost and fail to reach the machine.
[0060] Load shaping parameters (such as a higher threshold and associated hard multiplier, a lower threshold and associated soft multiplier, a lower threshold expiration, and a throttling timeout range) can be determined in advance. The thresholds and multipliers can be selected to balance power safety and performance impact. The higher threshold may be close to the protected limit such that throttling is not triggered too frequently, but not so close that the system does not have enough time to react to events. The hard multiplier can be close to 0 to quickly reduce a large amount of power. The lower threshold can be similarly set to have a balance between the throttling frequency and sufficient guard band for possible power peaks. The soft multiplier can be close to 1 to minimize the performance impact. The lower threshold expiration can be long enough to make the lower threshold useful, but not so long as to significantly increase the throttling frequency. The throttling timeout range can be set for a reasonable oscillation pattern. For example, the throttling timeout range can be about 1 - 20 seconds.
[0061] Thus, in an embodiment, as power increases, the lower threshold will be crossed more frequently. Since a multiplier is applied to the measured CPU utilization for tasks in a moving window, frequent activations are exacerbated. Staying above the lower threshold will result in the CPU time being limited to a percentage corresponding to the soft / low impact multiplier. In this way, as the power of the load fluctuates, the soft multiplier will converge to the impact of the high multiplier as needed. The high threshold serves as the activation point for the soft multiplier. As long as the load does not exceed the high threshold, the load is allowed to remain in or near the soft multiplier region indefinitely. The high threshold is further used to eliminate sudden power spikes that may occur when the low threshold is in effect.
[0062] The load shaping control strategy determines when the actuator should throttle the CPU usage and by how much to control power. Formally, the power consumption of the power domain can be written as:
[0063]
[0064] where t is (discrete) time, p is the total power consumption, N is the number of machines, f i is the power consumed by machine i (in the range [0,1]) as a monotonic function of the normalized machine CPU utilization, c i is the CPU used by the controllable tasks, u i is the uncontrollable CPU used by the exempt tasks and the Linux kernel, and n is the power consumed by non - machine devices. The CPU used by the controllable tasks (c i ) can be capped such that for a power limit l, p < l. Overload (p > l) can be preferably prevented, while keeping p close to l when p < l may improve efficiency.
[0065] According to some examples, a Random Unthrottled / Multiplicative Decrease (RUMD) algorithm can be used. If p(t) > l, a ceiling is applied to the CPU usage of each controllable task. The ceiling is equal to the previous usage of the task multiplied by a multiplier m in the range (0, 1). Then, the power consumption at the next time step is:
[0066]
[0067] This ceiling can be updated frequently, such as once per second, and c i will exponentially decay over time until p < l. Since u i and n terms, it is not possible to ensure that p(t + 1) < p(t). However, the implementation can provide a high confidence that the non-removable power is less than the power limit, i.e.,
[0068]
[0069] Therefore, the power will eventually drop below this limit. For example, if the system is configured to fully unthrottle all machines within 5 seconds, then 20% of the machines in a random non-overlapping set will be unthrottled per second.
[0070] In some examples, the ceiling can be incrementally increased for each machine in an accumulative manner simultaneously, resulting in an Additive Increase / Multiplicative Decrease (AIMD) algorithm. Similar to AIMD, the RUMD algorithm also has some desired properties of partial distribution. There are a central component and a distributed component that mainly act independently. Besides the total power, the central policy controller does not require detailed system states such as the CPU usage and task allocation of each machine. The distributed node controllers can make independent decisions based only on some parameters sent by the policy controller to all node controllers.
[0071] Failover mechanism: Active capping
[0072] The reactive ceiling system depends on the power signal provided by a power meter installed near the protected power device (such as a circuit breaker). This is different from the more widely adopted method of collecting power measurements from individual computing nodes and aggregating them at a higher level. It has the advantage of simplicity, avoiding aggregation and related challenges such as time misalignment and partial collection failures, and avoiding the need to estimate the power consumption of non-computing devices (such as data center air conditioners) that do not provide power measurements.
[0073] Transient network problems may cause power signal interruptions from several seconds to several minutes, while the downtime of the meter may be as long as several days to several weeks before being repaired. Without a power signal, the load shaping unit cannot determine how much each task should be throttled under a strong power guarantee. A fallback method (referred to as CPU imprisonment in this article) can operate without power measurements.
[0074] According to some examples, signals may also be collected from auxiliary sources, such as a power supply unit of the machine, or from a power model.
[0075] Node-level mechanism for CPU jailing
[0076] CPU jailing reduces the number of CPUs of a machine available to a task by modifying the CPU affinity mask. A parameter jailing_fraction is predetermined, which is the fraction of CPUs that are "jailed" so that they are unavailable to tasks. jailing_fraction can be the fraction of each machine's CPUs that is made unavailable to tasks.
[0077] It may happen that cluster operators intentionally overcommit resources and drive up machine utilization. When resource overcommitment is compounded with CPU confinement, intensive CPU resource contention is expected. Each task has a CPU limit request, which is converted to a share value in Linux CFS. CFS can be leveraged to maintain CPU ratios between tasks when available CPU decreases due to confinement. Certain privileged processes, such as critical system daemons, are explicitly exempted from confinement and can still run on confined CPUs. This is because their CPU usage is very low compared to regular tasks, but the risks and consequences of privileged system processes being starved of CPU are very high. For example, the risks and consequences could be that the machine will not function correctly.
[0078] According to some examples, each jail request carries a duration and can be renewed by request. Upon expiration of the jail period, the previously unavailable CPU immediately becomes available to all tasks.
[0079] CPU jailing immediately caps peak power consumption, as it effectively limits the maximum CPU utilization on each individual machine to (1-jailing_fraction). It thus places an upper limit on power consumption, allowing safe operation for long periods of time without power signals.
[0080] jailing_fraction can be applied uniformly to individual machines, regardless of their CPU utilization. Thus, machines with low utilization are less affected, while machines with high utilization are greatly affected. When machine CPU utilization is significantly below (1-jailing_fraction), CPU jailing has virtually no impact on tasks on those machines. A secondary effect is that jailed CPUs are more likely to enter deep sleep due to increased idleness, which helps to further reduce power. Note that not all jailed CPUs will enter deep sleep all the time, as exempted privileged processes and kernel threads may still use them from time to time.
[0081] In some instances, CPU confinement may result in a relaxed ability to differentiate QoS. For example, the latency of service tasks may be disproportionately affected compared to the throughput of batch tasks. This effect may be mitigated by the fact that latency-sensitive tasks run at a higher priority and can preempt lower priority throughput-oriented tasks. In some implementations, CPU confinement may be used only when load shaping is not applicable.
[0082] The jailing_fraction may be determined based on one or more factors, such as the performance SLO of the workload, CPU power relationship, and power oversubscription. Because CPU jailing does not distinguish between the QoS of the workload, there is a possibility that high-priority, latency-sensitive jobs may suffer performance degradation under high CPU contention. Applying the jailing_fraction at the expected frequency specified by the control policy should not compromise the SLO of these jobs.
[0083] For power safety, the power should be reduced to a safe level after jailing some parts of the core.The value of the jailing_fraction can be calculated according to the power oversubscription level and the CPU-power relationship of a given set of hardware in the power domain.
[0084] J=1-U cpu =1-f power cpu (1 / (1+osr))
[0085] Where J is jailing_fraction, U cpu is the maximum allowable CPU utilization, f power cpu is the function that converts power utilization to CPU utilization, while osr is the oversubscription ratio defined by the additional oversubscribed power capacity as a fraction of the nominal capacity. 1 / (1+osr) gives the maximum safe power utilization, which can be converted to U assuming the CPU power relationship is monotonic. cpu An exemplary jailing_fraction may be 20%-50%.
[0086] As a fallback, when power measurements from the meter are lost and the risk of power overload is high, CPU jailing may be triggered. The risk is determined by two factors, the predicted power consumption and the duration of meter unavailability. Higher predicted power consumption and longer meter unavailability mean higher risk. Given recent power consumption, the likelihood of power reaching the protected device limit during some meter outage can be predicted based on historical data. If the likelihood is high due to recent high power consumption and a sufficiently long outage, CPU jailing may be triggered.
[0087] Comparison of node-level mechanisms
[0088] Figure 3A -C illustrates exemplary results of running a workload that stresses the CPU and memory to maximize power consumption. Figure 3A The figure shows the results of using CPU bandwidth control. Figure 3B shows the results of using DVFS, and Figure 3C The results of using RAPL to limit CPU power are shown. The CPU power is normalized to the highest power observed when no power management mechanism is enabled. Figure 3A It is shown that by using CPU bandwidth control, the CPU power can be reduced to 34% of the maximum power due to the considerable CPU idle time of bandwidth control. Figure 3B It is shown that with DVFS, power consumption is still relatively high at 57% when the lowest frequency limit is applied. The clock frequency is normalized to the base frequency of the processor. The normalized frequency can be higher than 1.0 because the CPU clock frequency can be higher than the base frequency when thermal and other conditions permit. Only a portion of the possible frequency range is shown in this chart. Due to the stress testing nature of the workload, the highest clock frequency observed without applying limits is close to the base frequency. The upper right corner of the chart reflects that the frequency limit must be reduced below the actual clock frequency to reduce power consumption.
[0089] Figure 3C It is shown that of the three, RAPL has the widest power reduction range. It is able to reduce power to 22% of the maximum power. However, we noticed that system management tasks became unresponsive when RAPL approached the lowest power limits, suggesting that the risk of machine timeouts would be higher if these limits were actually used. In contrast, CPU bandwidth control only throttles throughput-oriented jobs and does not affect system management tasks.
[0090] Example Application: Load Shaping Results
[0091] Figure 4 The graphs show the results of experiments performed in a warehouse-sized data center running production workloads. Throttling was triggered manually using various combinations of parameters. Power data was collected from data center power meters, which happened to be what the system also reads. Other metrics were sampled from individual machines and aggregated at the same power domain level as the power readings. Power measurements were normalized to the device limits of the power domain. Machine metrics such as CPU utilization were normalized to the total capacity of all machines in the power domain unless otherwise stated. Task failures were normalized to the total number of tasks affected.
[0092] Load shaping is triggered by manually lowering the upper power threshold to just below the ongoing power consumption of the power domain. Figure 4 (a1) shows a typical load shaping pattern where the power oscillates around the lower threshold. In the seconds after the throttling is triggered, the power is greatly reduced due to the hard multiplier. At the same time, the lower threshold is activated. When the power drops below the lower threshold, the throttling is gradually increased and the power rises back until it reaches the lower threshold. Then, the power is reduced again, but with less margin due to the soft multiplier. The process continues as the throttle is repeatedly turned on and off, causing the power to oscillate around the lower threshold. Compared to (a1), Figure 4 (b1) shows that a soft multiplier close to 1.0 results in smaller oscillations as expected. The delay from load shaping triggering to significant power reduction is less than 5 seconds. Figure 4 (a2) and (b2) show the CPU utilization corresponding to (a1) and (b1), respectively. At the CPU utilization level shown, a 10% reduction in CPU utilization is required to achieve a 2% power reduction.
[0093] When tasks slow down, they should not fail due to CPU starvation or unexpected side effects.
[0094] A major advantage of load shaping over DVFS is that it can differentiate quality of service at the cgroup level, allowing jobs with different SLOs to run on the same machine. Jobs can be divided into groups based on their priority, and load shaping can be performed on low-priority groups while exempting high-priority groups.
[0095] Figure 5 The total power consumption and CPU usage of the two groups of jobs during the event are shown. The CPU usage of the shaped group and the exempt group decreased by about 10% and 3%, respectively. The exempt group was indirectly affected because the jobs in both groups are production jobs with complex interactions. An example is that the high priority master job in the exempt group coordinates the low priority workers in the shaped group, while the master job has less work and consumes lower CPU when the workers are throttled. However, the ability of load shaping to differentiate between jobs is obvious.
[0096] Load shaping reduces power to a safe level just below the threshold, and power is allowed to oscillate around that threshold. However, in extreme cases where power still exceeds the threshold, the system will need to continually reduce CPU bandwidth to jobs, eventually rendering the jobs inoperable. For example, power may remain high after throttling is triggered due to continuous scheduling of new compute-intensive jobs, or because many high-priority jobs that are exempt from this mechanism experience spikes in their CPU usage. In this case, stopping the affected jobs is the right trade-off to prevent power overload.
[0097] Figure 6 6 is a flow chart illustrating an exemplary method 600 for load shaping. Although the operations are illustrated and described in a particular order, it should be understood that the order may be modified or the operations may be performed simultaneously. In addition, operations may be added or omitted.
[0098] In block 610, power measurements are received for a plurality of machines performing one or more tasks. The measurements may be received, for example, at a controller or other processing unit or collection of processing units. The measurements may be received directly from the plurality of machines, or through one or more intermediate devices, such as power meters coupled to the plurality of machines. The plurality of machines may be, for example, computing devices in a data center. The plurality of machines may be performing one or more tasks or jobs.
[0099] In block 620, power limits for the plurality of machines are identified. For example, the power limits for each machine may be stored in a database, and the database may be accessed to identify the limits for one or more specific machines.
[0100] In block 630, the received power measurement is compared to the identified power limit. For example, the controller can determine whether the power measurement is close to the limit, such as whether the power measurement is within a predetermined range of the limit. According to some examples, one or more predefined thresholds can be set, wherein the comparison takes into account whether the one or more thresholds have been reached. For example, a first threshold can be set to a lower level, such as a first percentage of the power limit, and a second threshold can be set to a higher level compared to the first threshold, such as a second percentage of the power limit that is higher than the first percentage.
[0101] In block 640, based on the comparison of the power measurement with the identified power limit, it is determined whether to remove power consumed by the plurality of machines. For example, if one or more of the plurality of machines exceeds one or more predetermined thresholds, it may be determined that power should be removed. According to a first example, such as in conjunction with Figure 2A As shown, both a high threshold and a low threshold are set, and it can be determined whether either threshold has been reached. According to a second example, such as in combination with Figure 2B As shown, where both a high threshold and a low threshold are set, it is possible to first determine whether only the high threshold is reached. Once the high threshold is reached, triggering a response action, the low threshold can be activated for a predetermined time period. During this time period, it can be determined whether the low threshold has been reached, thereby triggering a second response action. Once the time period expires, the low threshold can be deactivated and is therefore no longer considered until the high threshold is reached again.
[0102] In block 650, if it is determined that power should be removed, a command may be sent to one or more of the plurality of machines to cause the one or more machines to operate in accordance with the one or more limits to reduce power. For example, an RPC may be sent that includes a multiplier for reducing the workload. For example, the multiplier may be a number between 0 and 1 that may be applied to a CPU limit, a sustained CPU rate, or a number of tasks, such as low priority tasks. Thus, when the multiplier is applied, the workload is reduced by a percentage corresponding to the multiplier.
[0103] In examples where multiple thresholds are set, different limits may be implemented based on which threshold is triggered. For example, triggering a higher threshold may result in sending a hard multiplier (e.g., a number close to 1), resulting in a significant reduction in workload. By the same example, triggering a lower threshold may result in sending a soft multiplier (such as a number closer to 0), resulting in a less significant reduction in workload than when a hard multiplier is applied.
[0104] Unless otherwise stated, the above alternative examples are not mutually exclusive, but can be implemented in various combinations to obtain unique advantages. Since these and other variations and combinations of the features discussed above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the embodiments should be made by way of example rather than by limitation of the subject matter defined by the claims. In addition, the examples described herein and the phrases expressed as "such as", "including", etc. should not be interpreted as limiting the subject matter of the claims to specific examples; on the contrary, the examples are intended to illustrate only one of many possible embodiments. In addition, the same reference numerals in different figures may identify the same or similar elements.
Claims
1. A method for power capping, comprising: Receiving, by one or more processors, power measurements for a plurality of machines in a data center, the plurality of machines performing one or more tasks; identifying power limits of the plurality of machines, including identifying a first threshold and identifying a second threshold that is higher than the first threshold; comparing, by the one or more processors, the received power measurements to power limits for the plurality of machines; determining whether to remove power consumed by the plurality of machines based on the comparison, including determining whether the received power measurement meets or exceeds the second threshold; commanding, by the one or more processors, the plurality of machines to operate according to one or more constraints to reduce power consumption below the first threshold when the received power measurement meets or exceeds the second threshold; determining, by the one or more processors, that the received power measurement exceeds the first threshold during a predetermined period of time after reducing power consumption below the first threshold; as well as When the first threshold is met or exceeded during the predetermined time period, the plurality of machines are commanded to reduce power below the first threshold.
2. The method according to claim 1, in, Commanding the plurality of machines to operate in accordance with one or more constraints to reduce power consumption below the first threshold includes sending a command to reduce the power consumption by a first predetermined percentage.
3. The method according to claim 2, wherein: When the first threshold is met or exceeded during the predetermined time period, the method further includes sending a second command to the plurality of machines to reduce power consumption by a second predetermined percentage, the second predetermined percentage being lower than the first predetermined percentage.
4. The method according to claim 1, wherein: The measurements are received from one or more power meters coupled to the plurality of machines.
5. The method according to claim 1, wherein: Instructing the plurality of machines to operate according to the one or more constraints includes limiting, for all machines in the power domain, processing times of tasks within a machine-level scheduler.
6. The method according to claim 5, wherein: Limiting the processing time includes applying a first multiplier to a low priority task, wherein the first multiplier is selected to prevent the low priority task from running and consuming power. The method of claim 6 , further comprising applying a second multiplier to the low priority task. The method of claim 6 , further comprising exempting one or more high priority tasks from the limit.
9. The method of claim 1, further comprising limiting a number of available schedulable processor entities to control power when the power measurement becomes unavailable.
10. The method according to claim 1, wherein: Commanding the plurality of machines to operate in accordance with the one or more constraints includes making a portion of a processing component of each of the plurality of machines unavailable for tasks.
11. A system for power capping, comprising: One or more processors in communication with a plurality of machines performing one or more tasks, the one or more processors being configured to: receiving power measurements for a plurality of machines; identifying power limits of the plurality of machines, including identifying a first threshold and identifying a second threshold that is higher than the first threshold; comparing the received power measurements to power limits for the plurality of machines; determining whether to remove power consumed by the plurality of machines based on the comparison, including determining whether the received power measurement meets or exceeds the second threshold; commanding, by the one or more processors, the plurality of machines to operate according to one or more constraints to reduce power consumption below the first threshold when the received power measurement meets or exceeds the second threshold; determining that the received power measurement exceeds the first threshold during a predetermined time period after reducing power consumption below the first threshold; as well as When the first threshold is met or exceeded during the predetermined time period, the plurality of machines are commanded to reduce power below the first threshold.
12. The system according to claim 11, in, Commanding the plurality of machines to operate in accordance with one or more constraints to reduce power consumption below the first threshold includes sending a command to reduce the power consumption by a first predetermined percentage.
13. The system according to claim 12, wherein: When the first threshold is met or exceeded during the predetermined time period, the one or more processors are further configured to send a second command to the plurality of machines to reduce power consumption by a second predetermined percentage, the second predetermined percentage being lower than the first predetermined percentage.
14. The system according to claim 11, wherein: The measurements are received from one or more power meters coupled to the plurality of machines.
15. The system according to claim 11, wherein: Instructing the plurality of machines to operate according to the one or more constraints includes limiting, for all machines in the power domain, processing times of tasks within a machine-level scheduler.
16. The system of claim 15, wherein: In limiting the processing time, the one or more processors are configured to apply a first multiplier to a low priority task, wherein the first multiplier is selected to prevent the low priority task from running and consuming power.
17. The system of claim 16, wherein: In limiting the processing time, the one or more processors are configured to exempt one or more high priority tasks from the limit.
18. The system of claim 11, wherein: Commanding the plurality of machines to operate in accordance with the one or more constraints includes making a portion of a processing component of each of the plurality of machines unavailable for tasks.
19. A non-transitory computer-readable medium storing instructions executable by one or more processors to perform operations comprising: receiving power measurements for a plurality of machines in a data center, the plurality of machines performing one or more tasks; identifying power limits of the plurality of machines, including identifying a first threshold and identifying a second threshold that is higher than the first threshold; comparing the received power measurements to power limits for the plurality of machines; determining whether to remove power consumed by the plurality of machines based on the comparison, including determining whether the received power measurement meets or exceeds the second threshold; commanding the plurality of machines to operate according to one or more constraints to reduce power consumption below the first threshold when the received power measurement meets or exceeds the second threshold; determining that the received power measurement exceeds the first threshold during a predetermined time period after reducing power consumption below the first threshold; as well as When the first threshold is met or exceeded during the predetermined time period, the plurality of machines are commanded to reduce power below the first threshold.
Citation Information
Patent Citations
Rack resource utilization
US20170255243A1
Dynamic power capping of a subset of servers when a power consumption threshold is reached and allotting an amount of discretionary power to the servers that have power capping enabled
US9250684B1