A cluster power capping configuration method and computing device
By reasonably configuring the power cap value of the server cluster, the problem of inefficient operation of computing nodes is solved, and efficient utilization of power power and smooth execution of important services are achieved.
Patent Information
- Application Number
- CN202411220605.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-08-30
AI Technical Summary
In the prior art, the power cap value configuration of the server cluster is unreasonable, resulting in inefficient operation of the computing node.
By obtaining the power cap value of each cluster, and when the sum exceeds the maximum power output power of the power, the power cap value of the cluster with lower priority is reduced, and at the same time, the power cap value of the calculation node is configured based on the number of CPU cores and the minimum operating power, and the power cap value of the power is reasonably allocated.
The operation efficiency of the computing node and the utilization of power supply are improved, ensuring the smooth execution of important services.
Smart Images

Figure CN119248604B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of server technology, and in particular to a method for configuring a cluster power capping value and a computing device. Background Art
[0002] Server clusters play an important role in processing large-scale scientific problems and huge data sets due to their high computing speed. With the advancement of computer technology, the computing density of server clusters has increased, and the power consumption of server clusters has also increased.
[0003] Related technologies limit the power capping of server clusters to reduce overall power consumption. However, this approach can lead to irrational power capping configurations for compute nodes within the server cluster, which in turn affects the operational efficiency of the compute nodes. Summary of the Invention
[0004] The embodiments of the present application provide a cluster power capping configuration method and computing device, which configures the power capping value of each computing node between clusters, reasonably configures the power capping value of each computing node in the cluster, and improves the operating efficiency of the computing nodes.
[0005] In the first aspect, an embodiment of the present application provides a method for configuring a cluster power cap value, obtaining power cap values corresponding to multiple clusters; when the sum of the power cap values corresponding to multiple clusters is greater than the maximum output power of the power supply, reducing the power cap value corresponding to one or more clusters; wherein the power cap value of the cluster after reduction is greater than or equal to the sum of the minimum operating powers of each computing node in the corresponding cluster; for each cluster in the multiple clusters, configuring the power cap value of the corresponding computing node according to the latest power cap value of the cluster and the number of CPU cores running on each computing node in the cluster; wherein the power cap value of each computing node is greater than or equal to its corresponding minimum operating power.
[0006] In an embodiment of the present application, when the sum of the power capping values of the clusters is greater than the maximum output power of the power supply, the power capping values of one or more clusters are reduced so that the sum of the power capping values of the clusters is less than or equal to the maximum output power of the power supply, thereby enabling each cluster to operate normally. In addition, when reducing the power capping value of the cluster, the embodiment of the present application takes into account the minimum operating power of each computing node in the cluster. While ensuring the operation of each computing node in the cluster, the power capping value of the computing node is further reasonably configured based on the number of CPU cores running on each computing node in the cluster, that is, the operating status of the computing node, thereby improving the operating efficiency of each computing node.
[0007] Optionally, the power capping values corresponding to one or more clusters are reduced in order of priority of each cluster from low to high; wherein the priority of the cluster is determined based on the importance of the services undertaken by the cluster.
[0008] In the embodiment of the present application, the power capping value of the cluster with lower priority is preferentially reduced, and the power capping value of the cluster with higher priority is maintained as much as possible, so that important services in the cluster with higher priority can be smoothly executed.
[0009] Optionally, the power capping value data tables corresponding to the plurality of clusters are queried to obtain the power capping value corresponding to each cluster in the current time period; wherein the power capping value data table includes the power capping value of the corresponding cluster in each time period.
[0010] In the embodiment of the present application, different power capping values can be set for different time periods. For example, during peak hours, a relatively low power capping value can be set for each cluster while ensuring normal operation of each cluster, thereby reducing the operating cost of each cluster.
[0011] Optionally, the average operating power corresponding to each cluster is obtained, and each average operating power is used as the power capping value of the corresponding cluster.
[0012] In the embodiment of the present application, by using the average operating power of the cluster as the power capping value of the cluster, the operating efficiency of the computing nodes in the cluster can be improved while satisfying the premise that the computing nodes in the cluster operate at the minimum power.
[0013] Optionally, an initial power capping value is set for the corresponding computing node according to the minimum operating power corresponding to each computing node in the cluster; when the power capping value of the cluster is greater than the sum of the minimum operating powers of each computing node, the initial power capping value of each computing node is adjusted according to the number of CPU cores running on each computing node and the remaining power index; wherein the remaining power index is the portion of the power capping value of the cluster that exceeds the sum of the minimum operating powers of each computing node.
[0014] In an embodiment of the present application, the power cap value of the cluster is first allocated to each computing node according to the minimum operating power of the computing node to ensure that each computing node can operate normally; then, the power index of the unallocated cluster is dynamically allocated to each computing node according to the number of CPU cores running on each computing node, that is, the operating status of the computing node, to ensure the performance of each computing node.
[0015] Optionally, obtain the power index that can be borrowed or the power index that needs to be borrowed of each cluster; among them, the clusters that can borrow power index are classified as first-category clusters, and the clusters that need to borrow power index are classified as second-category clusters; and distribute the sum of the power index that can be borrowed of each first-category cluster to each second-category cluster.
[0016] In an embodiment of the present application, by borrowing power indicators between clusters, for the first type of cluster, while ensuring the normal operation of the cluster itself, the excess power indicators are loaned to the second type of cluster, thereby improving the power utilization rate; for the second type of cluster, by borrowing power indicators, the performance of the second type of cluster can be improved.
[0017] Optionally, the power index that can be lent or the power index that needs to be borrowed of the running computing nodes in each cluster are obtained; wherein, the power index that can be lent by the computing node is the part of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can borrow the power index is regarded as the first category computing node; the power index that needs to be borrowed by the computing node is the part of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as the second category computing node; for any cluster, the power index that can be lent by the cluster is obtained based on the sum of the power indexes that can be lent by all first category computing nodes in the cluster and the sum of the power indexes that need to be borrowed by all second category computing nodes; wherein, the power index that can be lent by the cluster is the part of the sum of the power indexes that can be lent by all first category computing nodes that exceeds the sum of the power indexes that need to be borrowed by all second category computing nodes.
[0018] Optionally, the power index that can be lent or the power index that needs to be borrowed of the running computing nodes in each cluster are obtained; wherein, the power index that can be lent by the computing node is the part of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can borrow the power index is regarded as the first category computing node; the power index that the computing node needs to borrow is the part of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as the second category computing node; for any cluster, the power index that the cluster needs to borrow is obtained based on the sum of the power index that can be lent by all first category computing nodes in the cluster and the sum of the power index that need to be borrowed by all second category computing nodes; wherein, the power index that the cluster needs to borrow is the part of the sum of the power index that need to be borrowed by all second category computing nodes that exceeds the sum of the power index that can be lent by all first category computing nodes.
[0019] Optionally, the sum of the power indicators that can be borrowed by the first-category clusters is allocated to one or more second-category clusters in descending order of priority of the second-category clusters.
[0020] In the embodiment of the present application, the second-category cluster with a higher priority can obtain the power indicators it needs to borrow first, thereby ensuring that important services within the cluster are executed smoothly.
[0021] Optionally, the power capping value of each corresponding computing node is adjusted according to the power index allocated to the second type cluster and the average operating power of each running computing node.
[0022] In an embodiment of the present application, for a computing node whose power capping value is less than the average operating power, the operating efficiency and reliability of the computing node are improved by adjusting the power capping value to the average operating power.
[0023] Optionally, the power capping value of each corresponding computing node is adjusted according to the average operating power of each running computing node.
[0024] In a second aspect, an embodiment of the present application provides a device for configuring a cluster power cap value, comprising an acquisition module, a reduction module and a configuration module; wherein the acquisition module is used to obtain the power cap values corresponding to multiple clusters respectively; the reduction module is used to reduce the power cap values corresponding to one or more clusters when the sum of the power cap values corresponding to multiple clusters is greater than the maximum output power of the power supply; wherein the power cap value of the cluster after reduction is greater than or equal to the sum of the minimum operating powers of each computing node in the corresponding cluster; the configuration module is used to configure the power cap value of the corresponding computing node for each cluster in the multiple clusters according to the latest power cap value of the cluster and the number of CPU cores running on each computing node in the cluster; wherein the power cap value of each computing node is greater than or equal to its corresponding minimum operating power.
[0025] In a third aspect, an embodiment of the present application provides a computing device, comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory; wherein, when the program stored in the memory is executed, the processor is used to execute the cluster power capping value configuration method as described in any embodiment of the first aspect.
[0026] In a fourth aspect, an embodiment of the present application provides a data center, comprising multiple clusters and management nodes; the management node is used to execute the cluster power capping value configuration method described in any embodiment of the first aspect.
[0027] In a fifth aspect, an embodiment of the present application provides a data center comprising multiple clusters, each cluster comprising multiple computing nodes, one of which serves as a management node; the management node is used to execute the cluster power capping value configuration method described in any embodiment of the first aspect.
[0028] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when running on a computing device, executes the method for configuring the cluster power capping value as described in any embodiment of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following briefly introduces the drawings required for use in the embodiment or the prior art description. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0030] Figure 1 A schematic diagram of an application scenario provided in an embodiment of the present application;
[0031] Figure 2 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0032] Figure 3 A flowchart of a method for configuring a cluster power capping value provided in an embodiment of the present application;
[0033] Figure 4 A flowchart of another method for configuring cluster power capping values provided in an embodiment of the present application;
[0034] Figure 5 A schematic diagram of inter-cluster power indicator transfer provided in an embodiment of the present application;
[0035] Figure 6a A schematic diagram of the structure of a device for configuring cluster power capping values provided in an embodiment of the present application;
[0036] Figure 6b A schematic diagram of the structure of another device for configuring cluster power capping values provided in an embodiment of the present application;
[0037] Figure 7 A structural diagram of a data center provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] It should be noted that the embodiments described in this application are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0039] The terms "first" and "second" in the specification and claims of this application are used to distinguish different objects rather than to describe a specific order of objects. For example, "first data" and "second data" are used to distinguish different data rather than to describe a specific order of data.
[0040] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0041] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0042] In order to make the following description of the embodiments clearer, the technical terms involved in this application are first introduced.
[0043] A server cluster is a group of servers that participate in workload management.
[0044] High performance computing (HPC) is a technology that uses powerful processor clusters to process massive multidimensional data sets in parallel and solve complex problems at extremely high speeds.
[0045] To facilitate understanding of the technical solution of the present application, the application scenarios of the embodiments of the present application are introduced below.
[0046] See also Figure 1 , which shows a schematic diagram of an application scenario of an embodiment of the present application.
[0047] like Figure 1 As shown, this application scenario includes a management node 110 and multiple clusters 120, such as servercluster1, servercluster2, ..., serverclustern. Each cluster 120 includes multiple computing nodes. The number of computing nodes in each cluster can be the same or different. When the number is the same, each cluster includes m computing nodes, such as server1, server2, ..., serverm.
[0048] The management node 110 pre-configures a power cap for each cluster 120. If the power caps of server cluster 1, server cluster 2, ..., server cluster n exceed the maximum output power of the power supply, the power caps of one or more clusters 120 are lowered in descending order of priority. Each cluster 120 then configures a power cap for server 1, server 2, ..., server m based on the current power cap and the number of CPU cores running in each of the clusters. The power caps configured for servers 1, server 2, ..., server m must at least ensure that servers 1, server 2, ..., server m operate at minimum power.
[0049] When the sum of the power capping values of the clusters is greater than the maximum output power of the power supply, the power capping values of one or more clusters are reduced so that the sum of the power capping values of the clusters is less than or equal to the maximum output power of the power supply, thereby enabling the normal operation of each cluster. In addition, when reducing the power capping value of the cluster, the embodiment of the present application considers the minimum operating power of each computing node in the cluster. While ensuring the operation of each computing node in the cluster, the power capping value of the computing node is further reasonably configured based on the number of CPU cores running on each computing node in the cluster, that is, the operating status of the computing node, thereby improving the operating efficiency of each computing node.
[0050] The management node 110 may be any one of a physical server, a cloud server, or a virtual device (such as a virtual machine).
[0051] The computing nodes in the cluster 120 may be any of physical servers, cloud servers, or virtual devices (such as virtual machines).
[0052] The embodiment of the present application does not specifically limit the function of the cluster 120. For example, the cluster 120 can be an AI cluster or an HPC cluster.
[0053] In addition, the management node 110 in the embodiment of the present application can be any computing node in each cluster 120, and the management node 110 can also be a computing device independent of each cluster 120.
[0054] See also Figure 2 , which is a structural diagram of a computing device provided in an embodiment of the present application.
[0055] like Figure 2As shown, the computing device 200 includes a processor 210, a memory 220 and a communication interface 230; wherein the memory 220 is used to store computer instructions; the processor 210 is used to execute the computer instructions, so that the computing device 200 executes the cluster power capping value configuration method.
[0056] In some embodiments, the processor 210 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices. A general-purpose processor may also be a microprocessor or any conventional processor.
[0057] In some embodiments, the memory 220 may be a volatile memory or a non-volatile memory, such as a register. Specifically, a volatile memory refers to a memory in which the data stored therein will be lost when the power supply is interrupted. Among them, the volatile memory is mainly a random access memory (RAM), including a static random access memory (SRAM) and a dynamic random access memory (DRAM). A non-volatile memory refers to a memory in which the data stored therein will not be lost even if the power supply is interrupted. Common non-volatile memories include read-only memory (ROM), optical disks, magnetic disks, solid-state hard disks, and various memory cards based on flash memory technology.
[0058] In some embodiments, the memory 220 stores executable code, and the processor 210 executes the code to implement a testing method for an operating system.
[0059] The communication interface 230 is used to enable the computing device 200 to communicate with the server, and the computing device 200 obtains test cases of the server's operating system through the communication interface 230. The test cases include interpreted language scripts and compiled language scripts.
[0060] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into address bus, data bus and control bus. For ease of understanding, Figure 2 Just one thick line is used, but that does not mean there is only one bus or one type of bus.
[0061] In this embodiment of the present application, when the power capping value configured for each cluster is less than or equal to the maximum output power of the power supply, and the power capping value configured for each cluster satisfies the corresponding computing node's operation at minimum power, the power capping value of each computing node in the cluster is further configured based on the number of CPU cores used by each computing node in the cluster. The following describes the dynamic allocation of power capping values between clusters in conjunction with specific embodiments.
[0062] See also Figure 3 , which is a flowchart of a method for configuring a cluster power capping value provided in an embodiment of the present application.
[0063] like Figure 3 As shown, the method includes:
[0064] S310: The computing device obtains power capping values corresponding to the plurality of clusters.
[0065] The power capping value refers to limiting the maximum operating power of the cluster to prevent it from exceeding a specific value, thereby protecting the power from overload damage or ensuring stable power operation under specific conditions.
[0066] In one possible implementation, each cluster is configured with a corresponding power capping value in different time periods. The computing device queries the power capping value data table corresponding to each cluster according to the current time to obtain the power capping value corresponding to each cluster at the current time.
[0067] For example, the power capping value data table corresponding to server cluster 1 is shown in Table 1 below:
[0068] Table 1
[0069]
[0070] At the current time, in the case of peak power consumption period 1, the power cap value of server cluster 1 obtained by the computing device is 80%*P SC1 ; In the case of the current time for the application level period, the power cap value of servercluster1 obtained by the computing device is PSC1 ; In the case of the current time during the off-peak period, the power cap value of server cluster 1 obtained by the computing device is P SC1 In the case of peak power consumption period 2 at the current time, the power cap value of server cluster 1 obtained by the computing device is 70%*P SC1 Among them, the electricity price during peak period 2 is higher than that during peak period 1, the electricity price during peak period 1 is higher than that during level period, and the electricity price during level period is higher than that during valley period.
[0071] It should be noted that the power capping value configured for the cluster in each time period in the embodiment of the present application should at least meet the condition that all computing nodes in the corresponding cluster run at the minimum power. For example, for server cluster 1, 80%*P SC1 、P SC1 and 70%*P SC1 All computing nodes in server cluster 1 can run at minimum power.
[0072] In addition, the embodiments of the present application can adjust the time corresponding to the peak power consumption period, the time corresponding to the level power consumption period, and the time corresponding to the low power consumption period according to different regions and different seasons, thereby improving the flexibility of configuring the cluster power capping value.
[0073] In the embodiment of the present application, by configuring the power cap value of the cluster according to the electricity prices of different electricity consumption periods (peak electricity consumption period, level electricity consumption period and off-peak electricity consumption period), the operating cost of the cluster can be reduced while ensuring the normal operation of the cluster.
[0074] In a possible implementation, the computing device obtains the average operating power corresponding to each cluster, and uses each average operating power as a power capping value of the corresponding cluster.
[0075] The average operating power of the cluster is calculated by the computing device based on the historical operating power of the corresponding cluster. In the embodiment of the present application, the storage method of the historical operating power of the cluster is specifically limited, for example, the historical operating power of the cluster is stored in a database or memory space.
[0076] It should be understood that the historical operating power in the embodiment of the present application can be updated in real time, and the embodiment of the present application does not specifically limit the frequency of updating the historical operating power. For example, the historical operating power is updated once every hour.
[0077] For example, taking server cluster 1 as an example, the computing device calculates the average operating power of the cluster based on the historical operating power of server cluster 1 in the past 24 hours.
[0078] In the embodiments of the present application, the computing device uses the cluster's average operating power as the cluster's power capping value. This improves the operating efficiency of the computing nodes within the cluster while ensuring that the computing nodes in the cluster operate at minimum power. Furthermore, if the computing device does not have a power capping value for the corresponding time period, using the cluster's average operating power as the power capping value allows the computing device to continue configuring the cluster's power capping value, thereby increasing the robustness of the software program.
[0079] S320: When the sum of the power capping values corresponding to the plurality of clusters is greater than the maximum output power of the power supply, the computing device reduces the power capping value corresponding to one or more clusters.
[0080] The reduced power capping value of the cluster is greater than or equal to the sum of the minimum operating powers of the computing nodes in the corresponding cluster.
[0081] It should be understood that in the embodiments of the present application, multiple clusters can be powered by the same power supply. If the power caps of multiple clusters exceed the maximum output power of the power supply, the power caps of one or more clusters need to be reduced. The following describes methods for reducing the power caps of clusters.
[0082] In one possible implementation, the computing device reduces the power capping value corresponding to one or more clusters in ascending order of priority of the clusters, where the priority of the clusters is determined based on the importance of the services undertaken by the clusters.
[0083] For example, the computing device can assess the importance of each business using a trained importance assessment model. The importance assessment model is trained using data such as the business value, urgency, resource requirements, dependencies, and service level agreements (SLAs) of each business. It should be noted that the present embodiments do not limit the training process of the importance assessment model.
[0084] Among them, business value is reflected in the degree of contribution of the business to the achievement of the goal. The higher the contribution, the more important the business is; the urgency is reflected in the time limit for completing the business. The more urgent the business is to be processed, the more important the business is; the dependency is reflected in the business's demand for computing resources, storage resources and network resources. Among them, intensive businesses may require higher priority when resources are limited, and the higher the importance; if the business is associated with the SLA, it must comply with the performance standards specified in the SLA, such as response time, reliability, availability, etc., and the importance of the business associated with the SLA is usually relatively high.
[0085] The following describes how to obtain cluster priorities for computing devices.
[0086] For example, after configuring the services of each cluster, the computing device can obtain the importance of each service in the cluster according to the importance evaluation model, and then obtain the importance score of the cluster, and then obtain the priority of each cluster according to the importance score of each cluster.
[0087] For example, if server cluster 1 is configured with services 1 and 2, with importance scores of 5 and 4, respectively, then server cluster 1's corresponding importance score is 9. Server cluster 2 is configured with services 3 and 4, with importance scores of 3 and 2, respectively, then server cluster 2's corresponding importance score is 5. Server cluster 3 is configured with service 5, with importance score of 1, then server cluster 3's corresponding importance score is 1. Therefore, among server cluster 1, server cluster 2, and server cluster 3, server cluster 1 has the highest priority, server cluster 2 has the second highest, and server cluster 3 has the lowest.
[0088] Exemplarily, the computing device may also determine the priority of each cluster according to the number of important services configured in the cluster.
[0089] For example, important services include Service 1, Service 2, and Service 3. If server cluster 1 is configured with Service 1 and Service 2, server cluster 1 includes two important services, server cluster 2 is configured with Service 3, server cluster 2 includes one important service, and server cluster 3 does not have any important services, then server cluster 1 has the highest priority, server cluster 2 has the second highest priority, and server cluster 3 has the lowest priority.
[0090] For example, the services to be performed by each cluster in each time period are planned in advance, and the computing device can obtain the priority of each cluster in the corresponding time period by reading the cluster tag.
[0091] For example, the label corresponding to server cluster 1 indicates that the priority of server cluster 1 is 3, the label corresponding to server cluster 2 indicates that the priority of server cluster 2 is 2, and the label corresponding to server cluster 3 indicates that the priority of server cluster 3 is 1. A higher number corresponds to a higher priority.
[0092] It should be noted that in the embodiment of the present application, the priority of each cluster can be updated in real time, thereby more reasonably configuring power capping values for the cluster and the computing nodes within the cluster, thereby improving the operating efficiency of the cluster and the computing nodes.
[0093] In the aforementioned embodiment, the manner in which the computing device obtains the priority of the cluster is merely exemplary, and the embodiment of the present application does not limit the manner in which the priority of the cluster is obtained.
[0094] The following describes how computing devices can reduce the power cap of a cluster.
[0095] For example, among server cluster 1, server cluster 2, and server cluster 3, server cluster 1 has the highest priority, server cluster 2 has the second highest priority, and server cluster 3 has the lowest priority. The computing device first lowers the power cap of server cluster 3. If, after lowering the power cap of server cluster 3, the sum of the power caps of all clusters is still greater than the maximum output power of the power supply, the computing device lowers the power cap of server cluster 2, and so on, until the sum of the power caps of all clusters is less than or equal to the maximum output power of the power supply.
[0096] For easier understanding, the following Table 2 is provided for further introduction:
[0097] Table 2
[0098]
[0099] If the maximum output power of the power supply is 2000W, and the corresponding power caps for server cluster 1, server cluster 2, and server cluster 3 are 750W, 700W, and 700W, respectively, the total power cap reduction for each cluster is 150W. In descending order of priority, reduce the power cap for server cluster 3 from 700W to 600W; then reduce the power cap for server cluster 2 from 700W to 650W. At this point, the total power cap for server cluster 1, server cluster 2, and server cluster 3 equals the maximum output power of the power supply. There is no need to adjust the power cap for server cluster 3. The latest power caps for server cluster 1, server cluster 2, and server cluster 3 are 600W, 650W, and 750W, respectively.
[0100] In this embodiment, the power capping values for each cluster are lowered in ascending order of priority until the sum of all cluster power capping values is less than the maximum output power of the power supply. As much power capping value as possible is reserved for higher-priority clusters, allowing important services in these higher-priority clusters to be executed smoothly.
[0101] S330: For each of the multiple clusters, the computing device configures a power capping value of a corresponding computing node according to the latest power capping value of the cluster and the number of CPU cores running on each computing node in the cluster.
[0102] The power capping value of each computing node is greater than or equal to its corresponding minimum operating power.
[0103] In one possible implementation, an initial power cap is set for each computing node in the cluster according to the corresponding minimum operating power of each computing node. When the power cap of the cluster is greater than the sum of the minimum operating powers of each computing node, the initial power cap of each computing node is adjusted according to the number of CPU cores running on each computing node and the remaining power index. The remaining power index is the portion of the power cap of the cluster that exceeds the sum of the minimum operating powers of each computing node.
[0104] To facilitate understanding, the following describes how to configure the power capping value for a compute node, in conjunction with Table 3:
[0105] Table 3
[0106]
[0107] For server cluster 3, its power cap is equal to the sum of the minimum operating power of server 1, server 2, and server 3. The computing device directly configures the power cap of server 1, server 2, and server 3 to 200 W respectively.
[0108] For server cluster 2, whose power cap is greater than the sum of the minimum operating power of server 1, server 2, and server 3, the computing device first sets the initial power caps of server 1, server 2, and server 3 to 200 W respectively. It then allocates the remaining 50 W to server 1, server 2, and server 3 based on the number of CPU cores running on each computing node, and adjusts the corresponding initial power caps.
[0109] For example, the initial power cap value for server 1 is adjusted by [50 / (15+15+20)]*15, the initial power cap value for server 2 is adjusted by [50 / (15+15+20)]*15, and the initial power cap value for server 3 is adjusted by [50 / (15+15+20)]*20. After adjustment, the power cap values for server 1, server 2, and server 3 are 215W, 215W, and 220W, respectively.
[0110] For server cluster 2, its power cap is greater than the sum of the minimum operating power of server 1, server 2, and server 3. The computing device first sets the initial power caps of server 1, server 2, and server 3 to 200 W respectively. It then allocates the remaining 150 W to server 1, server 2, and server 3 according to the number of CPU cores running on each computing node, and adjusts the corresponding initial power caps. After the adjustment, the power caps for server 1, server 2, and server 3 are 245 W, 245 W, and 260 W respectively.
[0111] In an embodiment of the present application, while ensuring the operation of each computing node in the cluster, the power capping value of the computing node is reasonably configured based on the number of CPU cores running in the computing node in the cluster, that is, the operating status of the computing node, to improve the operating efficiency of each computing node.
[0112] In addition, the cluster power cap configuration method provided in the embodiments of the present application also includes the inter-cluster power quota adjustment and the configuration of the power cap of the computing nodes within the cluster after the adjustment is completed. The following describes the implementation of the inter-cluster power quota adjustment and the implementation of the power cap adjustment of the computing nodes within the cluster.
[0113] See also Figure 4 , which is a flowchart of another method for configuring cluster power capping values provided in an embodiment of the present application.
[0114] like Figure 4 As shown, the method includes:
[0115] S410: The computing device obtains power capping values corresponding to the plurality of clusters.
[0116] The contents involved in this step have been described in detail in the above embodiments and will not be repeated here.
[0117] S420: The computing device increases the power capping value of a cluster whose power capping value is less than the minimum operating power of a corresponding computing node.
[0118] exist Figure 4 In, P SC Indicates the power capping value of the cluster, P server-Min Indicates the minimum operating power of the compute nodes in the cluster.
[0119] For example, server cluster 1 includes server 1, server 2, and server 3. The power capping value of server cluster 1 is P.SC1 , the minimum operating power corresponding to server1 is P server1-Min , the minimum operating power corresponding to server2 is P server2-Min , the minimum operating power corresponding to server3 is P server3-Min If P SC1 Less than P server1-Min 、P server2-Min and P server3-Min The sum of the power consumption of server cluster 1 is increased.
[0120] S430: When the sum of the power capping values corresponding to the plurality of clusters is greater than the maximum output power of the power supply, the computing device reduces the power capping value corresponding to one or more clusters.
[0121] The power capping value after the cluster is reduced is greater than or equal to the sum of the minimum operating power of each computing node in the corresponding cluster, P SC1 .
[0122] exist Figure 4 The sum of the power capping values of each cluster is expressed as P SC-TOTAL .
[0123] The contents involved in this step have been described in detail in the above embodiments and will not be repeated here.
[0124] S440: For each of the multiple clusters, the computing device configures a power capping value of a corresponding computing node according to the latest power capping value of the cluster and the number of CPU cores running on each computing node in the cluster.
[0125] The power capping value of each computing node is greater than or equal to its corresponding minimum operating power.
[0126] The contents involved in this step have been described in detail in the above embodiments and will not be repeated here.
[0127] Because clusters' power capping requirements are dynamic, some clusters have sufficient power to lend, while others have insufficient power and need to borrow. To better utilize power caps and improve the performance of clusters with insufficient power, the following embodiments of this application will describe how to implement inter-cluster power allocation.
[0128] S450: The computing device obtains the power index that can be lent or the power index that needs to be borrowed by each cluster.
[0129] Among them, the clusters that can borrow output power indicators are classified as the first type of clusters, and the clusters that need to borrow power indicators are classified as the second type of clusters.
[0130] In one possible implementation, a computing device obtains the power index that can be lent or the power index that needs to be borrowed of the running computing nodes in each cluster; wherein, the power index that can be lent by the computing node is the portion of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can borrow the power index is classified as a first-category computing node; the power index that needs to be borrowed by the computing node is the portion of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is classified as a second-category computing node; for any cluster, the power index that can be lent by the cluster is obtained based on the sum of the power indexes that can be lent by all first-category computing nodes in the cluster and the sum of the power indexes that need to be borrowed by all second-category computing nodes; wherein, the power index that can be lent by the cluster is the portion of the sum of the power indexes that can be lent by all first-category computing nodes that exceeds the sum of the power indexes that need to be borrowed by all second-category computing nodes.
[0131] It should be understood that the number of CPU cores running on the computing node in the embodiment of the present application is not 0.
[0132] The following describes how a computing device obtains the power indicators that can be lent by the cluster and the power indicators that the cluster needs to borrow, based on Table 4:
[0133] Table 4
[0134]
[0135] If the power capping value corresponding to the computing node is greater than the maximum operating power, the computing node can borrow the output power indicator. If the power capping value corresponding to the computing node is less than the average operating power, the computing node needs to borrow the power indicator.
[0136] In server cluster 1, the power cap value of server 1 is greater than the maximum operating power, so server 1 can borrow 20W of power. The power cap value of server 2 is greater than the maximum operating power, so server 2 can borrow 30W of power. The power cap value of server 3 is less than the average operating power, so server 3 needs to borrow 40W of power. Among them, server 1 and server 2 are first-class computing nodes, and the total power index that can be borrowed is 50W. Server 3 is a second-class computing node, and the power index that needs to be borrowed is 40W. Therefore, server cluster 1, as a first-class cluster, can borrow 10W of power, that is, P OUT1 10W.
[0137] In server cluster 2, the power cap value of server 1 is greater than the maximum operating power, so server 1 can borrow 15W of power. The power cap value of server 2 is greater than the maximum operating power, so server 1 can borrow 15W of power. The power cap value of server 3 is less than the average operating power, so server 3 needs to borrow 10W of power. Among them, server 1 and server 2 are first-class computing nodes, and the total power index that can be borrowed is 30W. Server 3 is a second-class computing node, and the power index that needs to be borrowed is 10W. Therefore, server cluster 1, as a first-class cluster, can borrow 20W of power, that is, P OUT2 It is 20W.
[0138] In server cluster 3, the power cap value of server1 is less than the average operating power, so server1 needs to borrow 20W of power. The power cap value of server2 is less than the average operating power, so server needs to borrow 10W of power. The power cap value of server3 is greater than the maximum operating power, so server3 can borrow 10W of power. As a first-class computing node, server3 can borrow 10W of power. As second-class computing nodes, server1 and server2 need to borrow a total of 30W of power. Therefore, server cluster 1, as a second-class cluster, needs to borrow 20W of power, that is, P IN3 It is 20W.
[0139] It should be noted that the average operating power of the computing node in the embodiment of the present application is the average value of the historical operating power of the computing node under the current number of running CPU cores; correspondingly, the maximum operating power of the computing node is the maximum value of the historical operating power of the computing node under the current number of running CPU cores.
[0140] For example, if server 1 currently has 15 CPU cores, the historical operating power of all server 1 running 15-core CPUs is obtained, and the average value of all historical operating powers and the maximum value of the historical operating powers are calculated, which are used as the average operating power and maximum operating power of server 1 respectively.
[0141] In one possible implementation, the computing device obtains the number of CPU cores running in each cluster, as well as the maximum operating power and average operating power of each cluster; obtains the power index that can be borrowed by the cluster based on the power cap value and maximum operating power of each cluster; obtains the power index that the cluster needs to borrow based on the average operating power of each cluster and the power cap value of the cluster; wherein, the power index that can be borrowed by the cluster is the portion of the cluster's power cap value that exceeds the maximum operating power, and the power index that the cluster needs to borrow is the portion of the cluster's average operating power that exceeds the power cap value.
[0142] For example, server cluster 1 has 50 CPU cores. The computing device obtains historical operating data from server cluster 1's historical operating power data when server cluster 1 was running with 50 CPU cores, and based on this historical operating data, obtains server cluster 1's average operating power and maximum operating power. If server cluster 1's power cap is greater than server cluster 1's maximum operating power, server cluster 1 is a first-class cluster, and the power indicator that can be borrowed is the portion of server cluster 1's power cap that exceeds server cluster 1's maximum operating power. If server cluster 1's power cap is less than server cluster 1's average operating power, server cluster 1 is a second-class cluster, and the power indicator that needs to be borrowed is the portion of server cluster 1's average operating power that exceeds server cluster 1's power cap.
[0143] In an embodiment of the present application, the average operating power and the maximum operating power of the cluster are obtained by the number of CPU cores running in the cluster, and the power index that the cluster can lend or the power index that needs to be borrowed are obtained. This can quickly respond to the adjustment of power capping values between clusters, so that the second type of cluster can quickly obtain the required power capping value, thereby improving the performance of the second type of cluster.
[0144] S460: The computing device distributes the sum of the power indicators that can be borrowed by each first-type cluster to each second-type cluster.
[0145] In a possible implementation, the sum of the power indicators that can be borrowed by the first-category clusters is distributed to the second-category clusters in descending order of priority of the second-category clusters.
[0146] It should be understood that the above embodiment introduces a method for obtaining the priority of a cluster, and the method for obtaining the priority of the second type of cluster is not described in detail here.
[0147] For example, the first type of clusters includes server cluster 1, server cluster 2, and server cluster 5, wherein the loanable power index corresponding to each of server cluster 1, server cluster 2, and server cluster 5 is P OUT1 、P OUT2 and P OUT5 The second type of clusters includes server cluster 3 and server cluster 5, where the power indicators required to be borrowed by server cluster 3 and server cluster 5 are P IN 3 and P IN 5 Among them, servercluster3 has a higher priority than server cluster5.
[0148] For ease of understanding, this embodiment of the application will introduce the inter-cluster power indicator secondment method in conjunction with the following Table 5:
[0149] Table 5
[0150]
[0151] The total power quota that can be borrowed by the first type of cluster is 40 W, and the total power quota that needs to be borrowed by the second type of cluster is 50 W. According to the priority of the second type of clusters from high to low, the computing device first allocates 30 W of power quota to server cluster 3, and then allocates the remaining 20 W of power quota to server cluster 5.
[0152] For ease of understanding, Figure 5 As shown, the computing device will P OUT1 、P OUT2 and P OUT4 The sum of the power required by server cluster 3 is P IN3 , the power index required by server cluster 5 is P IN5 .
[0153] In the embodiment of the present application, for the two second-category clusters (server cluster 3 and server cluster 5), the power indicators obtained are as follows: in the first case, server cluster 3 obtains the power indicator it needs to borrow, and server cluster 5 obtains part of the power indicator it needs to borrow; in the second case, server cluster 3 obtains part of the power indicator it needs to borrow, and server cluster 5 does not obtain the power indicator it needs to borrow; in the third case, server cluster 3 and server cluster 5 both obtain the power indicator they need to borrow.
[0154] The following describes how to adjust the power capping values of computing nodes in a cluster for the above three situations.
[0155] S470: The computing device adjusts the power capping value of each computing node according to the power index obtained by each second type cluster and the average operating power of each computing node in each second type cluster.
[0156] In a possible implementation, server cluster 3 obtains the power index that it needs to borrow, and server cluster 5 obtains part of the power index that it needs to borrow.
[0157] For server cluster 3, adjust the power caps for all compute nodes in the cluster (both Category 1 and Category 2 compute nodes) to the corresponding average operating power. Specifically, lower the power caps for Category 1 compute nodes to the corresponding average operating power, while raise the power caps for Category 2 compute nodes to the corresponding average operating power. For server cluster 5, lower the power caps for Category 1 compute nodes to the corresponding average operating power. Raise the power caps for Category 2 compute nodes to the corresponding average operating power, in ascending order of the power requirements required by the Category 2 compute nodes.
[0158] In the embodiment of the present application, the power capping requirements of the second-category computing nodes that require less borrowed power indicators are preferentially met, and as many second-category computing nodes as possible are met, thereby improving the performance of the cluster.
[0159] In a possible implementation, server cluster 3 obtains part of the power indicators that need to be borrowed, and server cluster 5 obtains the power indicators that do not need to be borrowed.
[0160] For server cluster 3, the power caps for the first-category compute nodes in the cluster were lowered to the corresponding average operating power. The power caps for the second-category compute nodes were then raised to the corresponding average operating power, in ascending order of the required power metrics for the second-category compute nodes. For server cluster 5, the power caps for the first-category compute nodes in the cluster were lowered to the corresponding average operating power. The power caps for the second-category compute nodes were then raised to the corresponding average operating power, in ascending order of the required power metrics for the second-category compute nodes.
[0161] In a possible implementation, server cluster 3 obtains part of the power indicators that need to be borrowed, and server cluster 5 obtains the power indicators that do not need to be borrowed.
[0162] For server cluster 3, adjust the power caps for each compute node in the cluster (both Category 1 and Category 2 compute nodes) to the corresponding average operating power. Specifically, the power caps for Category 1 compute nodes are lowered to the corresponding average operating power, while the power caps for Category 2 compute nodes are increased to the corresponding average operating power. For server cluster 5, adjust the power caps for each compute node in the cluster (both Category 1 and Category 2 compute nodes) to the corresponding average operating power. Specifically, the power caps for Category 1 compute nodes are lowered to the corresponding average operating power, while the power caps for Category 2 compute nodes are increased to the corresponding average operating power.
[0163] In the embodiment of the present application, by redistributing the power index among clusters, power is more fully utilized and power waste is reduced. In addition, for the second type of clusters (clusters that need to borrow power indexes), borrowing power indexes is beneficial to improving the performance of the second type of clusters.
[0164] In addition, the embodiment of the present application also provides a device for configuring the cluster power capping value, the structural diagram of which is shown in FIG. Figure 6a shown.
[0165] See also Figure 6a , the apparatus includes: an acquisition module 610, a reduction module 620 and a configuration module 630;
[0166] An acquisition module 610 is configured to acquire power capping values corresponding to a plurality of clusters;
[0167] A reduction module 620 is configured to reduce the power capping values corresponding to one or more clusters when the sum of the power capping values corresponding to the multiple clusters is greater than the maximum output power of the power supply; wherein the power capping value of the cluster after the reduction is greater than or equal to the sum of the minimum operating powers of the computing nodes in the corresponding cluster;
[0168] Configuration module 630 is used to configure the power cap value of the corresponding computing node for each cluster in the multiple clusters based on the cluster's latest power cap value and the number of CPU cores running on each computing node in the cluster; wherein the power cap value of each computing node is greater than or equal to its corresponding minimum operating power.
[0169] In an embodiment of the present application, when the sum of the power capping values of the clusters is greater than the maximum output power of the power supply, the power capping values of one or more clusters are reduced so that the sum of the power capping values of the clusters is less than or equal to the maximum output power of the power supply, thereby enabling each cluster to operate normally. In addition, when reducing the power capping value of the cluster, the embodiment of the present application takes into account the minimum operating power of each computing node in the cluster to ensure the operation of each computing node in the cluster, and further reasonably configures the power capping value of the computing node based on the number of CPU cores running on each computing node in the cluster, that is, the operating status of the computing node, thereby improving the operating efficiency of each computing node.
[0170] Optionally, the reduction module 620 is specifically configured to reduce the power capping value corresponding to one or more clusters in descending order of priority of each cluster; wherein the priority of the cluster is determined based on the importance of the service undertaken by the cluster.
[0171] Optionally, the acquisition module 610 is specifically configured to query power capping value data tables corresponding to multiple clusters to obtain the power capping value corresponding to each cluster at the current time; wherein the power capping value data table includes the power capping value of the corresponding cluster in each time period.
[0172] Optionally, the acquisition module 610 is specifically configured to acquire the average operating power corresponding to each cluster, and use each average operating power as the power capping value of the corresponding cluster.
[0173] Optionally, configuration module 630 is specifically used to set an initial value of the power capping value for the corresponding computing node according to the minimum operating power corresponding to each computing node in the cluster; when the power capping value of the cluster is greater than the sum of the minimum operating powers of each computing node, the initial value of the power capping value of each computing node is adjusted according to the number of CPU cores running on each computing node and the remaining power index; wherein the remaining power index is the part of the power capping value of the cluster that exceeds the sum of the minimum operating powers of each computing node.
[0174] In addition, the embodiment of the present application also provides a structural diagram of another device for configuring cluster power capping value, such as Figure 6b shown.
[0175] like Figure 6b As shown, the apparatus includes: an acquisition module 610, a reduction module 620, a configuration module 630 and a secondment module 640;
[0176] The above embodiment has introduced the acquisition module 610, the reduction module 620 and the configuration module 630, which will not be repeated here.
[0177] The loan module 640 is used to obtain the power indicators that can be borrowed or the power indicators that need to be borrowed by each cluster; among them, the clusters that can borrow power indicators are classified as first-category clusters, and the clusters that need to borrow power indicators are classified as second-category clusters; the sum of the power indicators that can be borrowed by all first-category clusters is distributed to each second-category cluster.
[0178] Optionally, the loan module 640 is specifically used to obtain the power index that can be borrowed or the power index that needs to be borrowed of the running computing nodes in each cluster; wherein, the power index that can be borrowed by the computing node is the part of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can borrow the power index is regarded as the first category computing node; the power index that needs to be borrowed by the computing node is the part of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as the second category computing node; for any cluster, the power index that can be borrowed by the cluster is obtained based on the sum of the power indexes that can be borrowed by all first category computing nodes in the cluster and the sum of the power indexes that need to be borrowed by all second category computing nodes; wherein, the power index that can be borrowed by the cluster is the part of the sum of the power indexes that can be borrowed by all first category computing nodes that exceeds the sum of the power indexes that need to be borrowed by all second category computing nodes.
[0179] Optionally, the loan module 640 is specifically used to obtain the power index that can be borrowed or the power index that needs to be borrowed of the running computing nodes corresponding to each cluster; wherein, the power index that can be borrowed by the computing node is the part of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can borrow the power index is regarded as the first type of computing node; the power index that the computing node needs to borrow is the part of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as the second type of computing node; for any cluster, the power index that the cluster needs to borrow is obtained based on the sum of the power indexes that can be borrowed by all first type computing nodes in the cluster and the sum of the power indexes that need to be borrowed by all second type computing nodes; wherein, the power index that the cluster needs to borrow is the part of the sum of the power indexes that need to be borrowed by all second type computing nodes that exceeds the sum of the power indexes that can be borrowed by all first type computing nodes.
[0180] Optionally, the loaning module 640 is specifically configured to allocate the sum of the loanable power indicators of all first-category clusters to one or more second-category clusters in descending order of priority of each second-category cluster.
[0181] Optionally, the adjustment module 640 is specifically configured to adjust the power capping value of each computing node according to the power index allocated by the second type of cluster and the average operating power of each computing node in operation.
[0182] Optionally, the adjustment module 640 is specifically configured to adjust the power capping value of each computing node according to the average operating power of each computing node in operation.
[0183] In addition, the embodiment of the present application also provides a data center, the structural diagram of the data center is as follows Figure 7 shown.
[0184] like Figure 7 As shown, the data center includes multiple clusters 120, such as server cluster 1, server cluster 2...server cluster n, each cluster 120 includes multiple computing nodes, such as server 1, server 2...server m, one of which serves as a management node 110; the management node 110 is used to execute the configuration method of the cluster power capping value described in any of the aforementioned embodiments.
[0185] It should be noted that the management node 110 that executes the cluster power capping value configuration method described in any of the aforementioned embodiments in the embodiments of the present application may be a computing device independent of each cluster.
[0186] In addition, an embodiment of the present application further provides a computer program product, which, when executed on a computing device, executes the method for configuring the cluster power capping value as described in the aforementioned embodiment.
[0187] Through the description of the above embodiments, it can be seen that those skilled in the art can clearly understand that all or part of the steps in the above-mentioned embodiment methods can be implemented by means of software plus a general hardware platform. Based on this understanding, the technical solution of the present application can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as a read-only memory (ROM) / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in various embodiments or certain parts of the embodiments of the present application.
[0188] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment. Those of ordinary skill in the art can understand and implement it without paying any creative work.
[0189] The above description is merely an exemplary embodiment of the present application and is not intended to limit the scope of protection of the present application.
Claims
1. A method for configuring a cluster power capping value, characterized in that: The method comprises: Obtain the power capping values corresponding to multiple clusters; When the sum of the power capping values of the multiple clusters is greater than the maximum output power of the power supply, reduce the power capping value of one or more of the clusters; wherein the power capping value of the cluster after the reduction is greater than or equal to the sum of the minimum operating power of each computing node in the cluster; For each of the multiple clusters, configure a power capping value for the corresponding computing node based on the latest power capping value of the cluster and the number of CPU cores running on each computing node in the cluster; wherein the power capping value of each computing node is greater than or equal to its corresponding minimum operating power; The obtaining of power capping values corresponding to the plurality of clusters includes: The power capping value data tables corresponding to the plurality of clusters are queried to obtain the power capping value of each cluster at the current time; wherein the power capping value data table includes the power capping value of the cluster during peak hours, normal hours, and off-peak hours.
2. The method according to claim 1, characterized in that The lowering of the power capping value of one or more clusters includes: The power capping value of one or more clusters is reduced in order of priority of the clusters from low to high; wherein the priority of the cluster is determined based on the importance of the service undertaken by the cluster.
3. The method according to claim 1, characterized in that The obtaining of power capping values corresponding to the plurality of clusters includes: An average operating power corresponding to each of the clusters is obtained, and each of the average operating powers is used as a power capping value of each of the clusters.
4. The method according to claim 1, wherein The configuring the power capping value of the corresponding computing node according to the latest power capping value of the cluster and the number of CPU cores running on each computing node in the cluster includes: According to the minimum operating power of each computing node in the cluster, setting an initial power capping value for the corresponding computing node; When the power capping value of the cluster is greater than the sum of the minimum operating powers of each of the computing nodes, the initial value of the power capping value of each of the computing nodes is adjusted according to the number of CPU cores running on each of the computing nodes and the remaining power index; wherein the remaining power index is the portion of the power capping value of the cluster that exceeds the sum of the minimum operating powers of each of the computing nodes.
5. The method according to any one of claims 1 to 4, characterized in that After configuring the power capping value of the corresponding computing node, the method further includes: Obtaining the power index that can be lent or the power index that needs to be borrowed of each cluster; wherein the clusters that can lend the power index are classified as first-category clusters, and the clusters that need to borrow the power index are classified as second-category clusters; The sum of the power indicators that can be borrowed by all the first-category clusters is distributed to each of the second-category clusters.
6. The method according to claim 5, characterized in that Obtain the power indicators that can be borrowed by each cluster in the following way: Obtain the power index that can be lent or the power index that needs to be borrowed of each computing node running in each of the first-category clusters; wherein, the power index that can be lent by the computing node is the portion of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can lend the power index is regarded as a first-category computing node; the power index that needs to be borrowed by the computing node is the portion of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as a second-category computing node; For any of the clusters, the power index that can be lent by the cluster is obtained based on the total power index that can be lent by all the first-category computing nodes in the cluster and the total power index that all the second-category computing nodes need to borrow; wherein, the power index that can be lent by the cluster is the part of the total power index that can be lent by all the first-category computing nodes that exceeds the total power index that all the second-category computing nodes need to borrow.
7. The method according to claim 5, characterized in that Obtain the power indicators required for each cluster using the following methods: Obtain the power index that can be lent or the power index that needs to be borrowed of each computing node running in each second-category cluster; wherein, the power index that can be lent by the computing node is the portion of the power cap value of the computing node that exceeds the maximum operating power of the computing node, and the computing node that can lend the power index is regarded as a first-category computing node; the power index that needs to be borrowed by the computing node is the portion of the average operating power of the computing node that exceeds the power cap value of the computing node, and the computing node that needs to borrow the power index is regarded as a second-category computing node; For any of the clusters, the power index that the cluster needs to borrow is obtained based on the total power index that can be borrowed by all the first-category computing nodes in the cluster and the total power index that all the second-category computing nodes need to borrow; wherein, the power index that the cluster needs to borrow is the part of the total power index that needs to be borrowed by all the second-category computing nodes that exceeds the total power index that can be borrowed by all the first-category computing nodes.
8. The method according to claim 5, characterized in that The allocating the sum of the power indicators that can be borrowed by all the first-category clusters to each of the second-category clusters includes: The sum of the lentable power indicators of all the first-category clusters is distributed to one or more second-category clusters in descending order of priority of the second-category clusters.
9. A computing device, characterized in that include: at least one memory for storing a program; At least one processor is configured to execute the program stored in the memory; wherein, when the program stored in the memory is executed, the processor is configured to execute the cluster power capping value configuration method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Apparatus, system and method for power management
US20090077407A1