Power capacity planning for computing systems

By using machine learning algorithms to evaluate and predict the activity and power consumption of data center computing systems, the inaccuracy of power capacity planning in existing technologies is solved, enabling precise power capacity planning and real-time management of virtual systems, thereby improving the power management efficiency and reliability of data centers.

CN116830125BActive Publication Date: 2026-04-14EATON INTELLIGENT POWER LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
EATON INTELLIGENT POWER LTD
Filing Date
2021-11-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing data center computing system power capacity planning mainly relies on the maximum power consumption fixed derating value of the PS, which fails to effectively predict dynamic load changes, resulting in inaccurate power management. This is especially true in virtual systems, where existing software such as VMware vRealize Operation Manager 7 only considers computing and storage capacity while ignoring power capacity.

Method used

Machine learning algorithms are used to evaluate computing system activities and predict activity evolution. Combined with power consumption, UPS autonomy, and redundancy levels, accurate power capacity planning data is generated. The first, second, and third machine learning algorithms are used to predict the number of virtual machines, processing load, UPS autonomy, and redundancy levels, respectively, and user interface displays and warnings are generated.

Benefits of technology

It enables precise planning of power capacity for data center computing systems, can predict dynamic load changes, provide real-time alerts, and help data center managers optimize power architecture to avoid underpowering or overpowering, thereby improving system reliability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116830125B_ABST
    Figure CN116830125B_ABST
Patent Text Reader

Abstract

A computer-implemented method for power capacity planning of a computing system is disclosed, wherein the method comprises the following steps: a) assessing (S10) an activity of the computing system, wherein the activity comprises a number of program instances executed by the computing system and a load caused by the executed program instances; b) predicting (S12) an evolution of the activity based on the assessed activity by using a first machine learning algorithm (10) configured for activity evolution prediction; c) predicting (S14) a power consumption of the computing system as a first power metric based on the predicted activity evolution and by using a second machine learning algorithm (12) configured for power consumption prediction of the computing system; d) predicting (S16) an autonomy of one or more uninterruptible power supplies of the computing system as a second power metric based on the predicted power consumption and by using a third machine learning algorithm (12') configured for uninterruptible power supply autonomy prediction and receiving the power consumption prediction of the computing system as input; e) predicting (S18) a redundancy level of the computing system as a third power metric based on the predicted power consumption and a power architecture of the computing system; f) generating (S20) output data related to power capacity planning by processing the first, second and third power metrics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to power capacity planning for power devices (such as UPS or ePDU) of computing systems that can be used in data centers, and in particular to the prediction of power metrics and the evolution of the predicted power metrics based on one or more user-defined scenarios for use in power capacity planning of computing systems. Background Technology

[0002] A typical data center comprises a building or group of buildings with one or more rooms. Each room in such a data center typically contains one or more rows, in which one or more racks can be arranged, containing IT (information technology) system equipment, such as physical servers (PS) or server computers. This IT system equipment is typically powered by power equipment including power devices such as (but not limited to) electronic power distribution units (ePDUs) or uninterruptible power supplies (UPSs) or combinations thereof.

[0003] An example of a computing system is a virtual system comprising several virtual machines (VMs) hosted by two or more Service Providers (PSs). Such virtual systems can be used, for example, in a data center with PSs hosting the VMs. Typically, each PS hosts one hypervisor. Each hypervisor hosts one or more VMs.

[0004] Power capacity planning for computing systems used in data centers (such as the virtual systems mentioned earlier) is typically based on a fixed derating of the maximum power consumption of the computing system's power supply (PS). For example, in the case of 10 PS, each PS includes two power supply units, each with a maximum power consumption of 500 watts. A 60% derating would result in a derating of 6000 watts of power consumption.

[0005] When VMware vRealize Operations Manager 7 software from VMware is used to plan virtual systems, compute and storage metrics can be considered for capacity planning; however, this results in compute-only and storage-only capacity planning.

[0006] GB1919009.9 relates to power management of computing systems used in data centers, and specifically to managing actions on such computing systems, particularly actions to be performed through a "demand response" mechanism in response to power events or grid instability. The power management described in GB1919009.9 provides a method for predicting the impact of these actions on the power consumption of the computing system.

[0007] JP2011160596A relates to a power supply system comprising IT equipment, represented by a server, and a power supply device for supplying power to the IT equipment. The number of operating power supply units is controlled such that the power supply efficiency of the power supply device is maximized based on the load current flowing in the multiple operating servers, and an uninterruptible power supply (UPS) is arranged on the output side of each power supply unit. Furthermore, the number of operating power supply units is controlled by utilizing operational information or measuring power consumption. Even in the event of a prediction failure, insufficient current is compensated by power supply from the UPS installed at each output of the power supply unit to maintain stable operation of the server device and other equipment, avoiding momentary power interruptions in the power supply bus. Summary of the Invention

[0008] This specification describes a computer-implemented method and system for power capacity planning of a computing system that can be used in a data center.

[0009] According to one aspect of this specification, a computer-implemented method for power capacity planning of a computing system is provided, wherein the method includes the following steps:

[0010] a) Evaluate the activities of the computing system, including the number of program instances executed by the computing system and the load caused by the executed program instances;

[0011] b) Predict the evolution of the activity based on the evaluated activity by using a first machine learning algorithm configured for activity evolution prediction;

[0012] c) Based on the predicted activity evolution and by using a second machine learning algorithm configured for power prediction of the computing system, predict the power consumption of the computing system as a first power metric.

[0013] d) Based on the predicted power consumption and by using a third machine learning algorithm configured for predicting the autonomy of uninterruptible power supplies and receiving the power consumption prediction of the computing system as input, predict the autonomy of one or more uninterruptible power supplies of the computing system as a second power metric.

[0014] e) Based on the predicted power consumption and the power architecture of the computing system, predict the redundancy level of the computing system as a third power metric;

[0015] f) Generate output data related to power capacity planning by processing the first, second, and third power metrics.

[0016] The program instance may include virtual machines and / or containers, and the assessment of the activities of the computing system may include at least one of the following: determining the number of virtual machines executed by the computing system; determining the number of containers executed by the computing system; determining the evolution pattern of the number of virtual machines and / or containers executed by the computing system; and determining the evolution pattern of the processing load and / or storage load of each virtual machine and / or container.

[0017] The prediction of the evolution of the activity may include using the first machine learning algorithm to predict the evolution of the number of virtual machines and / or containers and / or the evolution of the processing load and / or storage load based on the determined number of virtual machines and / or containers, the evolution pattern of the determined number of virtual machines and / or containers, and / or the evolution pattern of the processing load and / or storage load of each determined virtual machine and / or container.

[0018] The second machine learning algorithm for predicting the power consumption of the computing system may be based on the power consumption of the physical server executing the program instance of the computing system, and / or the third machine learning algorithm for predicting the autonomy of one or more uninterruptible power supplies of the computing system may include an uninterruptible power supply autonomy model.

[0019] This prediction of the redundancy level of the computing system may include receiving data about the power architecture from a power manager program configured to manage the power requirements of the computing system.

[0020] The generation of output data related to power capacity planning may include generating data for displaying the first, second, and third power metrics (in particular, the time evolution of the first, second, and third power metrics) on a user interface.

[0021] The method may also include generating data for displaying warnings related to the first, second, and third power metrics on the user interface.

[0022] The method may further include receiving a user-defined scenario related to the power capacity planning, performing the prediction action c)-e) based on the received user-defined scenario to obtain the first, second, and third power metrics of the received user-defined scenario, and generating output data related to the power capacity planning by processing the first, second, and third power metrics of the received user-defined scenario.

[0023] Another aspect of this specification relates to a computer-implemented system for power capacity planning of a computing system, wherein the system is particularly configured to perform the methods according to any embodiment of this application, and wherein the system includes

[0024] An evaluation module is configured to evaluate the activities of the computing system, wherein the activities include the number of program instances executed by the computing system and the load caused by the executed program instances;

[0025] A first prediction module is configured to predict the evolution of an activity based on the evaluated activity by using a first machine learning algorithm configured for activity evolution prediction.

[0026] The second prediction module is configured to predict the power consumption of the computing system as a first power metric based on the predicted activity evolution and by using a second machine learning algorithm configured for power consumption prediction of the computing system, and is configured to predict the autonomy of one or more uninterruptible power supplies of the computing system as a second power metric based on the predicted power consumption and by using a third machine learning algorithm configured for uninterruptible power supply autonomy prediction and receiving the power consumption prediction of the computing system as input.

[0027] A third prediction module is configured to predict the redundancy level of the computing system as a third power metric based on the predicted power consumption and the power architecture of the computing system; and

[0028] An output data generation module is configured to generate output data related to power capacity planning by processing the first, second, and third power metrics.

[0029] Another aspect of this specification relates to a non-transitory computer-readable storage device for storing software, the software including instructions executable by a processor of a computing device, which, upon such execution, cause the computing device to perform the methods disclosed in this specification.

[0030] Details of one or more embodiments are set forth below in the accompanying drawings and description. Further features and advantages will be apparent from the description and drawings, and from the claims. Attached Figure Description

[0031] Figure 1A and Figure 1B An example of a method for calculating the power capacity planning of a system is shown;

[0032] Figure 2 An example of a system for calculating the power capacity planning of a system is shown;

[0033] Figure 3 Examples of representations of output data generated for three different user-defined scenarios using a method for calculating the power capacity planning of a system are shown. Detailed Implementation

[0034] In the following text, functionally similar or identical elements may have the same reference numerals. Absolute values ​​are shown below by way of example only and should not be construed as limiting.

[0035] The term "virtual machine" (VM) used herein describes the emulation of a particular computer system. In the context of this invention, a VM is a special case of a computer program with an operating system. This solution also applies to "lightweight" VMs, also known as "containers." VMs and containers can be considered as encapsulating computing environments, combining disparate IT components and isolating them from the underlying system (particularly the computing system on which the VM or container runs). The term "physical server" (PS) used herein describes an entity comprising a physical computer. A PS may include hypervisor software that configures the physical computer to host one or more virtual machines. In the context of this invention, a PS is a special case of a computing device. The term "virtual system" used herein refers to a system comprising two or more PSs, each hosting at least one VM. The term "computing system," as used herein, generally describes a system comprising software and hardware such as those used in data centers. In the context of this invention, a virtual system is a special case of a computing system. A computing system may include one or more virtual systems.

[0036] This specification relates to power capacity planning for computing systems that can be used in data centers, and more particularly to predicting power metrics and calculating the evolution of those predicted power metrics based on one or more user-defined scenarios for power capacity planning of computing systems. Furthermore, the methods and systems described herein allow data center managers to predict data center power metrics based on user-defined scenarios, such as adding 10 servers and 100 VMs to an existing data center computing system. Therefore, the methods and systems described herein allow data center managers to perform power capacity planning in data centers.

[0037] The methods and systems described herein may use some of the features described in GB 1919009.9, which is incorporated herein by reference and describes how machine learning can be used, in particular, to predict VM-level power consumption in computing systems, such as virtual systems. The machine learning algorithms described in GB 1919009.9 may be applied to at least some of the methods and systems described herein. In particular, the methods and systems from GB 1919009.9 may be extended to predict power capacity metrics according to this specification, which can then be used by the methods and systems as described herein.

[0038] Existing software configured for capacity planning in data centers (such as VMware vRealize Operations Manager 7 from VMware) performs capacity planning in data centers based solely on compute and storage metrics. This specification proposes extending power capacity planning to one or more power metrics such as:

[0039] - Power consumption, which means the electrical energy per unit time and the amount of energy supplied to operate components such as computing systems (especially PS).

[0040] - UPS autonomy, which means the duration for which a UPS can continue operating in the event of a power failure. This value depends on the load level of the computing system being powered by the UPS.

[0041] - Redundancy levels (N, N+1, 2N, ...) refer to replicating one or more critical components of a computing system within a data center to increase system reliability. This redundancy level can depend on the load level of the computing system.

[0042] Power metric predictions can be specifically based on:

[0043] - IT and power data acquisition (VM resource consumption, PS consumption, etc.).

[0044] - An artificial intelligence model for predicting power consumption in data centers.

[0045] The methods and systems described in this article can specifically solve customer problems in the following ways:

[0046] - Based on "Data Center IT Load Trends", it is possible to estimate when three key metrics are at predefined warning or critical thresholds.

[0047] - Based on "Data Center IT Load Trends" and specific IT load increases, it is possible to estimate when three key metrics are at specific warning or critical thresholds.

[0048] Figure 1A and Figure 1B The actions performed by the method used to calculate the power capacity planning of the system are shown:

[0049] Perform the following tasks a) and b) to obtain data on the activity level in the computing system:

[0050] a) In S10, the activities of the computing system are evaluated. The computing system is represented by data provided, for example, by a program used to plan the computing system in a data center. The data about the computing system may specifically include information about the number of VMs and / or containers. The activities evaluated include the number of program instances (particularly VMs and containers running on the PS) executed by the computing system, and the load caused by the executed program instances. The evaluation specifically includes determining the evolution pattern of the number of VMs and / or containers, such as continuous growth, flattening, decreasing, periodicity on a day / night level, periodicity on a weekend level, periodicity based on a specific annual event (e.g., Black Friday, Christmas, FIFA World Cup, etc.). Furthermore, the evaluation specifically includes determining the evolution pattern of the processing load (CPU load) and / or storage load for each VM and / or container, such as the corresponding load continuous growth, flattening, decreasing, periodicity on a day / night level, periodicity on a weekend level, periodicity based on a specific annual event (e.g., Black Friday, Christmas, FIFA World Cup, etc.).

[0051] b) In S12, the evolution of the activity is predicted based on the evaluated activity. For the prediction, a dedicated machine learning algorithm 10 is applied, which is configured to predict the activity evolution. The machine learning algorithm 10 receives the evaluated activity as input data and outputs a prediction of the activity evolution, particularly the evolution of the number of VMs and / or containers and / or the evolution of processing and / or storage load (especially at individual and global levels).

[0052] For power measurement and capacity planning, perform the following items c) through f):

[0053] c) In S14, based on the predicted activity evolution (item b) and by using another dedicated machine learning algorithm 12 constructed for predicting the power consumption of the computing system, the power consumption of the computing system is predicted as a first power metric (predicted power consumption). The machine learning algorithm 12 may be, for example, the algorithm described in GB1919009.9.

[0054] d) In S16, based on the predicted power consumption and by using yet another dedicated machine learning algorithm 12', the autonomy of one or more UPSs powering the computing system is predicted as a second power metric (predicted UPS autonomy). This dedicated machine learning algorithm is configured to receive the power consumption prediction of the computing system as input and is, for example, the algorithm described in GB1919009.9. In particular, the machine learning algorithm 12' may include a UPS autonomy model for predicting the autonomy of the UPSs of the computing system.

[0055] e) In S18, based on the predicted power consumption and the power architecture of the computing system, the predicted redundancy level of the computing system is used as a third power metric (predicted redundancy level). Information about the data center power architecture can be received, for example, from a program manager program 14 configured to manage the power requirements of the computing system (such as the Eaton Intelligent Power Manager (EIPM) software suite).

[0056] f) In S20, output data related to power capacity planning is generated by processing the first, second, and third power metrics. The processing may specifically include preparing the metrics as output data for display on a user interface (UI) (particularly a graphical UI (GUI) 16). The first to third power metrics obtained as described above can be considered as baseline predictions of the power metrics without additional capacity planning scenarios. The display of graphs representing these baseline predictions of the power metrics may already provide values ​​to the user. Furthermore, warnings 18 and / or alarms may be output, particularly in cases where power metrics change within a predetermined time span, such as when degradation can be foreseen within 6 months.

[0057] Figure 2 A block diagram of a system for calculating power capacity planning is shown. The system can be implemented by a computer program that executes on a computer. The system includes the following modules that implement the functions described above:

[0058] -Evaluation module 102 implements S10;

[0059] - The first prediction module 104 implements S12;

[0060] - The second prediction module 106 implements S14 and S16;

[0061] - The third prediction module 108 implements S18, and

[0062] - Output data generation module 110 implements S20.

[0063] The module can be implemented, for example, as part of a software suite configured for integrated power management in a data center and can extend the functionality of existing software suites, such as the EIPM software suite described above. The module can be considered a separate software module that implements the corresponding functions listed above and receives input data and generates output data, such as... Figure 2 As shown.

[0064] Further functionality can be provided by handling user-defined scenarios. User-defined scenarios can be provided for power capacity planning, such as adding 100 VMs to an existing PS with a given resource usage next month, or adding 30 new PSs hosting 200 VMs next month.

[0065] By taking into account user-defined scenarios to predict the first through third power metrics, the evolution of power metrics from user-defined scenarios will be calculated using S14, S16, and S18 (see c) through e) above. The newly predicted power metrics can then be "added" to the baseline prediction of power metrics without the additional capacity planning scenario obtained in S20 (see f) above. It can also be output for display on a UI, allowing users to see the power metrics for user-defined scenarios and compare them with power metrics without user-defined scenarios.

[0066] User-defined scenarios can be input by the user via a GUI into a computer program or software suite that implements a method and / or system for power capacity planning.

[0067] User-defined scenarios specifically allow users to anticipate and then adjust the power architecture to avoid any degradation in power capacity metrics. For example, users can define different user-defined scenarios, utilize methods and / or systems as described herein to perform power capacity planning, and visualize different power metrics for each scenario on a GUI for comparison. This enables users to detect degradation in power capacity metrics across different user-defined scenarios, thus providing them with the potential to improve power capacity planning for computing systems in their data centers. Furthermore, warnings and / or alerts can be output to users, automatically generated when specific parameters exceed specific thresholds (e.g., when exceeding a certain level of degradation). Users can, for example, modify user-defined scenarios via a GUI to adjust the power architecture regarding power metrics.

[0068] Figure 2 This demonstrates another application of how users can perform power capacity planning using the system shown: Figure 1B Warnings displayed in S20 regarding power metric thresholds, or warnings and / or alarms 18, may be interpreted by the user (in... Figure 2 The arrow from 18 to the user (e.g., because the user can understand under what circumstances the threshold is exceeded). The user can then specifically define a "new power architecture" with additional power capacity based on warnings and / or alarms. For example, the user can interpret an alarm or warning such that further power demands due to increased processing needs could lead to a further reduction in UPS autonomy, and the user plans to increase UPS autonomy by adding further redundancy to the power architecture. The "new power architecture" can be simulated, for example, using the EIPM software suite described above (in... Figure 2 (The arrow in the middle leads from the user to 14). The user will then be able to check whether the "new power architecture" resolves power metric warning "18".

[0069] Figure 3Examples of representations for three different user-defined scenarios 1 through 3 are shown. Each representation shows three graphs over time (in months): the bottom graph represents the redundancy level, the middle graph represents UPS autonomy in minutes (right vertical axis), and the top graph represents power in kW (left vertical axis). Scenario 1 is the predicted evolution without adding a specific load. Note that the trend can be flat or cyclical (daily, weekly, yearly, etc.). Scenario 2 is the predicted evolution when 30 VMs (Activity xxx) are added to an existing 100 PSs. The user can define VM compute characteristics as input parameters. Scenario 3 is the predicted evolution (Model xxx) when 30 VMs are added to a new 20 PSs. The user can define the following information as input parameters: new PS power characteristics (the estimate can be more accurate if the new PS is similar to the existing PSs), and VM compute characteristics.

[0070] The methods and systems described in this paper enable improved power capacity planning for computing systems, particularly virtual systems, especially those used in data centers, at a relatively fine granular level.

Claims

1. A computer-implemented method for power capacity planning of a computing system, the method comprising the following steps: a) Evaluate (S10) the activities of the computing system, wherein the activities include the number of program instances executed by the computing system and the load caused by the executed program instances; b) The evolution of the activity is predicted (S12) based on the evaluated activity by using a first machine learning algorithm (10) configured for activity evolution prediction; c) Based on the predicted activity evolution and by using a second machine learning algorithm (12) configured for power consumption prediction of the computing system, predict (S14) the power consumption of the computing system as a first power metric; d) Based on the predicted power consumption and by using a third machine learning algorithm (12') configured for predicting the autonomy of uninterruptible power supplies and receiving the power consumption prediction of the computing system as input, predict (S16) the autonomy of one or more uninterruptible power supplies of the computing system as a second power metric. e) Based on the predicted power consumption and the power architecture of the computing system, predict (S18) the redundancy level of the computing system as a third power metric. f) Generate (S20) output data related to power capacity planning by processing the first, second and third power metrics.

2. The method of claim 1, wherein the program instance includes virtual machines and / or containers and the evaluation (S10) of the activities of the computing system includes at least one of: determining the number of virtual machines executed by the computing system; determining the number of containers executed by the computing system; determining the evolution pattern of the number of virtual machines and / or containers executed by the computing system; and determining the evolution pattern of the processing load and / or storage load of each virtual machine and / or container.

3. The method according to claim 2, wherein the prediction of the evolution of the activity (S12) comprises predicting the evolution of the number of virtual machines and / or containers and / or the evolution of the processing load and / or storage load by means of the first machine learning algorithm (10) based on the determined number of virtual machines and / or containers, based on the evolution pattern of the determined number of virtual machines and / or containers, and / or based on the evolution pattern of the determined processing load and / or storage load of each virtual machine and / or container.

4. The method according to any one of claims 1 to 3, wherein the second machine learning algorithm (12) for the prediction (S14) of the power consumption of the computing system is based on the power consumption of the physical server executing the program instance of the computing system, and / or wherein the third machine learning algorithm (12') for the prediction (S16) of the autonomy of one or more uninterruptible power supplies of the computing system includes an uninterruptible power supply autonomy model.

5. The method according to any one of claims 1 to 3, wherein the prediction of the redundancy level of the computing system (S18) includes receiving data about the power architecture from a power manager program (14), the power manager program being configured to manage the power requirements of the computing system.

6. The method according to any one of claims 1 to 3, wherein the generation (S20) of the output data related to power capacity planning includes generating data for displaying the first, second and third power metrics on the user interface (16).

7. The method of claim 6, further comprising generating data for displaying warnings (18) related to the first, second and third power metrics on the user interface (16).

8. The method according to any one of claims 1 to 3, wherein the method further comprises - Receive user-defined scenarios (20) related to the power capacity planning. - Perform steps c)-e) based on the received user-defined scenario (20) to obtain the first, second, and third power metrics of the received user-defined scenario (20), and - Output data related to power capacity planning is generated (S20) by processing the first, second and third power metrics of the received user-defined scenario (20).

9. A computer-implemented system (100) for calculating the power capacity planning of a system, wherein the system (100) is particularly configured to perform the method according to any one of claims 1 to 8, and wherein the system (100) comprises - Evaluation module (102), the evaluation module is configured to evaluate the activities of the computing system, wherein the activities include the number of program instances executed by the computing system and the load caused by the executed program instances; - A first prediction module (104) is configured to predict the evolution of an activity based on the evaluated activity by using a first machine learning algorithm (10) configured for activity evolution prediction; - A second prediction module (106) is configured to predict the power consumption of the computing system as a first power metric based on the predicted activity evolution and by using a second machine learning algorithm (12) configured for power consumption prediction of the computing system, and is configured to predict the autonomy of one or more uninterruptible power supplies of the computing system as a second power metric based on the predicted power consumption and by using a third machine learning algorithm (12') configured for uninterruptible power supply autonomy prediction and receiving the power consumption prediction of the computing system as input. - A third prediction module (108) is configured to predict the redundancy level of the computing system as a third power metric based on the predicted power consumption and the power architecture of the computing system. - Output data generation module (110), which is configured to generate output data related to power capacity planning by processing the first, second and third power metrics.

10. A non-transitory computer-readable storage device for storing software, the software comprising instructions executable by a processor of a computing device, the instructions, when executed in such a manner, causing the computing device to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Power management of computing system

    GB201919009D0

  • Techniques for adaptive demand / response energy management of electronic systems

    CN105324902A

  • Power feed system

    JP2011160596A