Power supply system, power management method, computer program product and storage medium
By monitoring the server load status and dynamic power distribution, combined with machine learning model and backup power module, the problems of energy waste and unstable operation in the server power supply system are solved, and efficient power management and stable power supply are achieved.
Patent Information
- Application Number
- CN202510104213.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The existing server power system adopts a fixed power distribution method, resulting in energy waste and operational instability problems when the server load changes.
The server load status is monitored through the sensing device, the control device is used to perform dynamic power distribution adjustments, and the load changes are predicted in combination with the machine learning model, and the power supply stability is ensured through the backup power module and the fault monitoring device.
The rationality of power distribution and the improvement of energy utilization efficiency have been achieved, and the stability and reliability of server operation have been ensured.
Smart Images

Figure CN119536486B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of servers, and in particular, to a power supply system, a power management method, a computer program product, and a storage medium. Background Art
[0002] With the continuous development of information technology, the status of servers has become increasingly important. In scenarios such as data centers, servers often need to run continuously for a long time to support the processing and storage of massive amounts of data, and the performance of the power supply system becomes crucial.
[0003] Currently, common server power supply systems mainly adopt a fixed power distribution method. Specifically, a power distribution plan for the server cluster is designed in advance according to the maximum power consumption of each server in the server cluster. This static power distribution mode is simple and easy to implement. However, in actual applications, the load of the server will change dynamically with different running tasks, resulting in energy waste in many cases for this fixed power distribution method. Especially when the entire server cluster is in a low-load state, the preset power supply is often much higher, causing the excess electric energy to be not effectively utilized. In addition, in some load peak periods in the current solution, the demand of the server may not be met, resulting in a decline in system performance and even affecting the stability of server operation.
[0004] In summary, how to effectively implement power management, improve energy utilization efficiency, and ensure the stability of server operation is a technical problem that those skilled in the art urgently need to solve at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a power supply system, a power management method, a computer program product, and a storage medium to effectively implement power management, improve energy utilization efficiency, and ensure the stability of server operation.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a power supply system, including: a plurality of power modules for supplying power to a server cluster, and a power management device connected to each of the power modules; the power management device includes:
[0008] a sensing device for monitoring the load status of each server in the server cluster;
[0009] A control device connected to the sensing device, configured to determine a load prediction value of each server after a first time period based on the load status of each server; adjust the power distribution of the power system so that the total power distribution error is lower than a preset first threshold and each server in the server cluster receives power supply from at least K power modules.
[0010] Wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed; K is a positive integer not less than 2, and the total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power system and the load prediction value of the server.
[0011] On the other hand, it further includes:
[0012] A fault monitoring device, configured to monitor the fault status of each power module in the power system and send it to the control device;
[0013] The control device is connected to the fault monitoring device and is further configured to perform a preset fault response operation when the fault monitoring device monitors a fault of the power module.
[0014] On the other hand, the power system further includes standby power modules;
[0015] Correspondingly, the control device performing the preset fault response operation includes:
[0016] Determining the number of the faulty power modules;
[0017] Selecting the same number of standby power modules from the standby power modules in the standby state to replace the faulty power modules and adjusting the power distribution of the power system.
[0018] On the other hand, the control device adjusting the power distribution of the power system includes:
[0019] The control device adjusts the power distribution of the power system according to a preset constraint rule;
[0020] Wherein, the constraint rule includes: for any power module in the power system, the total power provided by the power module does not exceed the maximum power output limit value of the power module.
[0021] On the other hand, the constraint rule further includes:
[0022] The total power obtained by a single server from the power system is greater than A times the load prediction value of the server; where A is a preset magnification value and A is greater than 1.
[0023] On the other hand, it further includes a communication device, and the communication device is connected to the control device;
[0024] The control device is further configured to: receive a power distribution instruction through the communication device, and adjust the power distribution of the power system according to the power distribution instruction.
[0025] On the other hand, the sensing device is specifically configured to:
[0026] Monitor the load status of each server in the server cluster in the most recent second time period;
[0027] Correspondingly, the control device determines the load prediction value of each server after the first time period based on the load status of each server, including:
[0028] The control device determines the load prediction value of each server after the first time period based on the load status of each server in the most recent second time period.
[0029] On the other hand, the sensing device is specifically configured to:
[0030] Monitor each server in the server cluster, and for any one of the servers, determine a central processing unit utilization rate sequence for reflecting the change in the central processing unit utilization rate of the server in the most recent second time period, a memory occupancy rate sequence for reflecting the change in the memory occupancy rate of the server in the most recent second time period, a request quantity sequence for reflecting the change in the request quantity of the server in the most recent second time period, and a network bandwidth sequence for reflecting the change in the network bandwidth of the server in the most recent second time period, and use the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request quantity sequence, and the network bandwidth sequence as the load status of the monitored server in the most recent second time period.
[0031] On the other hand, the control device determines the load prediction value of each server after the first time period based on the load status of each server in the most recent second time period, including:
[0032] For any one of the servers in the server cluster, input the load status of the server into a pre-trained first model to obtain a load prediction result output by the first model, which is used to reflect the load status of the server after a first time period; wherein, the load prediction result includes the central processing unit utilization rate, memory occupancy rate, number of requests, and network bandwidth of the server after the first time period;
[0033] Input the load prediction result into a pre-trained second model to obtain a load prediction value output by the second model, and use it as the determined load prediction value of the server after the first time period.
[0034] On the other hand, the control device determines the load prediction value of each server after the first time period based on the load status of each server in the most recent second time period, including:
[0035] For any one of the servers in the server cluster, based on the load status of the server in the most recent second time period, respectively determine N pending prediction values of the server after the first time period through N different calculation methods; wherein, N is a positive integer not less than 2;
[0036] Based on the N pending prediction values, determine the load prediction value of the server after the first time period.
[0037] On the other hand, based on the N pending prediction values, determining the load prediction value of the server after the first time period includes:
[0038] Based on the N pending prediction values and the respective weight parameters of the N different calculation methods, by weighted superposition of the N pending prediction values, obtain the load prediction value of the server after the first time period.
[0039] On the other hand, the control device is further configured to:
[0040] For any one server, determine the input power value currently required by the server;
[0041] Judge whether the error between the input power value and the load prediction value of the server exceeds a preset error range;
[0042] If so, use the input power value currently required by the server to replace the load prediction value of the server, and perform an operation to adjust the power distribution of the power system that powers the server cluster.
[0043] In a second aspect, the present invention provides a power management method, including:
[0044] Monitor the load status of each server in the server cluster;
[0045] Based on the load status of each server, determine the load prediction value of each server after a first time period; wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed;
[0046] Adjust the power distribution of the power system that powers the server cluster so that the total power distribution error is lower than a preset first threshold, and each server in the server cluster receives power supply from at least K power modules;
[0047] Wherein, the power system includes multiple power modules, K is a positive integer not less than 2, the total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power system and the load prediction value of the server.
[0048] In a third aspect, the present invention provides a computer program product, including computer programs / instructions, which when executed by a processor implement the steps of the power management method as described above.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the power management method as described above.
[0050] Applying the technical solution provided by the embodiments of the present invention, considering that the traditional fixed power distribution method is prone to energy waste and affects the stability of server operation. In this regard, in the solution of this application, dynamic adjustment of power distribution will be carried out. Specifically, in the solution of this application, it is necessary to monitor the load status of each server in the server cluster, and then based on the load status of each server, determine the load prediction value of each server after a first time period at the current moment. This is because the load will change, and the load prediction value of each server after a first time period is determined, so as to achieve dynamic power distribution, which is beneficial to ensuring the rationality of power distribution and thus beneficial to ensuring the stability of server operation.
[0051] When adjusting the power distribution of the power supply system that powers the server cluster, it is necessary to make the total power distribution error lower than a preset first threshold. The total power distribution error represents the sum of the power distribution errors of each server in the server cluster. That is to say, making the total power distribution error lower than the preset first threshold means making the power distribution situation of each server conform to the load prediction value of that server. This also means that when it is predicted that the future load of the server increases, the power supply can be increased in advance, thus avoiding the situation of unstable voltage or degraded system performance caused by the increased load. Conversely, when it is predicted that the future load of the server decreases, the power supply can be reduced in advance to prevent power waste and improve energy utilization efficiency.
[0052] Furthermore, in order to ensure power supply stability, it is also necessary to make each server in the server cluster receive power supply from at least K power modules, where K is a positive integer not less than 2. That is to say, each server will receive power supply from 2 or more power modules. In this way, when a single power module that powers the server fails, the server will not immediately lose all power supply, which is also beneficial to ensuring the stability of server operation to a certain extent. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 The flowchart of the power management method provided by a specific embodiment of the present invention;
[0055] Figure 2 The structural schematic diagram of the power supply system provided by a specific embodiment of the present invention;
[0056] Figure 3 The structural schematic diagram of the device provided by a specific embodiment of the present invention;
[0057] Figure 4 The structural schematic diagram of a computer-readable storage medium of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The core of the present invention is to provide a power supply system, a power management method, a computer program product and a storage medium, which can effectively implement power management, improve energy utilization efficiency, and ensure the stability of server operation.
[0059] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0060] Please refer to Figure 1 and Figure 2 , Figure 1 which is the implementation flowchart of the power management method provided by a specific embodiment of the present invention. Figure 2 which is the structural schematic diagram of the power system provided by a specific embodiment of the present invention. The power system may include: a plurality of power modules for powering the server cluster, and a power management device connected to each power module; the power management device includes:
[0061] a sensing device for monitoring the load status of each server in the server cluster;
[0062] a control device connected to the sensing device for determining the load prediction value of each server after the first time period based on the load status of each server; adjusting the power distribution of the power system so that the total power distribution error is lower than a preset first threshold and each server in the server cluster receives power supply from at least K power modules;
[0063] wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed; K is a positive integer not less than 2, and the total power distribution error represents the sum of the power distribution errors of each server in the server cluster. The power distribution error of a single server represents the error between the total power obtained by the server from the power system and the load prediction value of the server.
[0064] Specifically, the sensing device can monitor the load status of each server in the server cluster. The load status of the server can be measured by various indicators, and the specific measurement method can be set and adjusted according to actual needs. Moreover, for the load status of a certain server, it can be measured by various indicators at the current moment or based on various indicators within a certain recent period of time, which does not affect the implementation of the present invention and can be set and selected according to actual needs.
[0065] In a specific embodiment of the present invention, the sensing device is specifically used for: monitoring the load status of each server in the server cluster within the most recent second time period.
[0066] This implementation mode takes into account that monitoring the load status of each server in the monitoring server cluster is to determine the load prediction value of each server after the first duration at the current moment. Therefore, if the load status of each server in the recent period is monitored, it is beneficial to obtain an accurate load prediction value subsequently. Thus, in this implementation mode, specifically monitored is the load status of each server in the server cluster in the recent second duration.
[0067] Furthermore, in a specific implementation mode of the present invention, monitoring the load status of each server in the server cluster in the recent second duration may specifically include:
[0068] Monitoring each server in the server cluster, and for any one server, determining a central processing unit utilization rate sequence for reflecting the change of the central processing unit utilization rate of the server in the recent second duration, a memory occupancy rate sequence for reflecting the change of the memory occupancy rate of the server in the recent second duration, a request quantity sequence for reflecting the change of the request quantity of the server in the recent second duration, and a network bandwidth sequence for reflecting the change of the network bandwidth of the server in the recent second duration, and taking the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request quantity sequence, and the network bandwidth sequence as the load status of the monitored server in the recent second duration.
[0069] This implementation mode takes into account that for a server, the CPU (Central Processing Unit), memory, I / O (Input / Output) request quantity, and network bandwidth will all affect the load of the server. In other words, the load situation of the server will be reflected in the CPU, memory, I / O quantity, and network bandwidth. Therefore, in this implementation mode, the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request quantity sequence, and the network bandwidth sequence will be taken as the load status of the monitored server in the recent second duration.
[0070] The central processing unit utilization rate sequence reflects the change in the central processing unit utilization rate of the server within the second most recent time period. For example, the central processing unit utilization rate of the server can be monitored periodically and arranged in chronological order to obtain a central processing unit utilization rate sequence corresponding to the second most recent time period. Similarly, the memory occupancy rate of the server can be monitored periodically and arranged in chronological order to obtain a memory occupancy rate sequence corresponding to the second most recent time period. The number of I / O operations of the server, that is, the number of requests, can be monitored periodically and arranged in chronological order to obtain a request number sequence corresponding to the second most recent time period. The network bandwidth of the server can be monitored periodically and arranged in chronological order to obtain a network bandwidth sequence corresponding to the second most recent time period.
[0071] In addition, it can be seen that in this implementation manner, in the form of sequences, that is, through the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request number sequence, and the network bandwidth sequence, the load change situation of the server within the second most recent time period can be effectively reflected. When these sequences are used as the load status of the server within the second most recent time period, a more accurate load prediction value can be obtained subsequently.
[0072] The control device connected to the sensing device can determine the load prediction value of each server after the first time period based on the load status of each server. The load prediction value reflects the power required by the server after the first time period at the current moment. The current moment described here is also Figure 1 the triggering moment of step S102 in. At this time, the monitoring of the load status of each server in the server cluster has been completed. Therefore, based on the monitoring results, that is, based on the load status of each server, the load prediction value of each server at a certain future moment (after the first time period of the current moment) can be predicted.
[0073] In one implementation manner above, when monitoring the load status of each server in the server cluster, the load status of each server in the server cluster within the second most recent time period from the current moment is monitored. Therefore, in this implementation manner, the control device can determine the load prediction value of each server after the first time period of the current moment based on the load status of each server within the second most recent time period.
[0074] In a specific implementation manner of the present invention, the control device determines the load prediction value of each server after the first time period based on the load status of each server within the second most recent time period, which may specifically include:
[0075] For any server in the server cluster, input the load status of the server into the pre-trained first model to obtain a load prediction result output by the first model, which is used to reflect the load status of the server after the first time period; wherein, the load prediction result includes the central processing unit utilization rate, memory occupancy rate, number of requests, and network bandwidth of the server after the first time period.
[0076] Input the load prediction result into the pre-trained second model to obtain a load prediction value output by the second model, which is used as the determined load prediction value of the server after the first time period.
[0077] This implementation method considers that, first, based on Figure 1 the load status of each server obtained in step S101, predict the load status of each server after the first time period, and then, according to the predicted load status of each server after the first time period, obtain the load prediction value of each server, that is, obtain the power size required by each server.
[0078] In this regard, in this implementation method, the first model is a model that can predict the load status of a server at a future moment based on the load status of the server in the past period of time. Therefore, for any server in the server cluster, the load status of the server can be first input into the pre-trained first model. Considering that the load status of the server input into the first model usually includes the CPU situation, memory situation, number of IOs, and network bandwidth of the server in the most recent second time period, in this implementation method, the load prediction result output by the first model can include the central processing unit utilization rate, memory occupancy rate, number of requests, and network bandwidth of the server after the first time period.
[0079] The second model is a model that can predict the power size required by the server based on the load prediction result. Therefore, after obtaining the load prediction result output by the first model and inputting it into the pre-trained second model, the load prediction value output by the second model can be obtained, which is used as the determined load prediction value of the server after the first time period.
[0080] In practical applications, considering that machine learning models can accurately predict load changes, the first model and the second model used in this implementation method can usually be prediction models based on machine learning. By learning historical load data, the first model and the second model can meet the usage requirements of this application. Commonly used machine learning models include long short-term memory network models, random forest regression models, etc. The specific selection of the model type can be set and adjusted according to actual needs.
[0081] The first model and the second model need to be pre-trained. Taking the first model as an example, the training samples are the historical load states of the server, including the central processing unit utilization rate, memory occupancy rate, number of IO requests, and network bandwidth, which are arranged according to time respectively, and corresponding sequences can be obtained. The labels are the load conditions at a future time point corresponding to the respective samples, specifically including the central processing unit utilization rate, memory occupancy rate, number of IO requests, and network bandwidth at a future time point. The loss function can specifically use the mean square error, for example, which can minimize the gap between the predicted value and the true load value, and gradually optimize the performance of the model. Through multiple rounds of training, a first model that can effectively predict the load state of the server at a future time point can be obtained. The training of the second model is the same in principle. Through effective training, a second model that can effectively predict the power required by the server based on the load state of the server can be obtained.
[0082] Further, in a specific implementation manner of the present invention, the control device determines the load prediction values of each server after the first time period based on the load states of each server in the most recent second time period, which may include:
[0083] For any server in the server cluster, based on the load state of the server in the most recent second time period, N pending prediction values of the server after the first time period are determined respectively through N different calculation methods; where N is a positive integer not less than 2;
[0084] Based on the N pending prediction values, the load prediction value of the server after the first time period is determined.
[0085] This implementation manner takes into account that in addition to obtaining the load prediction value by the machine learning model, multiple analysis methods can also be combined, such as a model like the autoregressive moving average model that can effectively capture the periodic fluctuations and trend changes of the server load, and comprehensively obtain the load prediction value of the server after the first time period.
[0086] In this regard, in this implementation manner, for any server in the server cluster, based on the load state of the server in the most recent second time period, N pending prediction values of the server after the first time period are determined respectively through N different calculation methods.
[0087] For example, one of the calculation methods is the machine learning-based method described above, that is, through the first model and the second model described above, one undetermined predicted value after the first time period of the server is determined. Since each method can obtain one undetermined predicted value, a total of N undetermined predicted values can be obtained in this implementation. For example, another calculation method uses an autoregressive moving average model. The autoregressive moving average model also needs to be trained, and the samples used can also be the historical load data of the server, but it can pay more attention to the time series characteristics in the data, such as the periodic changes of the load status, long-term trends, etc. In addition, according to actual needs, the autoregressive and moving average parameters of the autoregressive moving average model can be adjusted to continuously optimize its prediction ability, which is particularly suitable for predicting long-term trends.
[0088] In other embodiments, the specific implementation of the N different calculation methods can be set and adjusted according to actual needs. In some cases, to simplify the calculation, N can be set to the minimum value of 2.
[0089] When determining the load prediction value of the server after the first time period based on the N undetermined predicted values, there can be various specific implementation methods. For example, a simple implementation method is to directly take the average of the N undetermined predicted values as the load prediction value of the server determined after the first time period.
[0090] In a specific embodiment of the present invention, determining the load prediction value of the server after the first time period based on the N undetermined predicted values may specifically include:
[0091] Based on the N undetermined predicted values and the respective weight parameters of the N different calculation methods, by weighted superposition of the N undetermined predicted values, the load prediction value of the server after the first time period is obtained.
[0092] This implementation takes into account that respective weight parameters can be set for the N different calculation methods, so that the influence degrees of the N different calculation methods on the load prediction value are different, improving the flexibility of the solution. And it can be understood that if it is found that the undetermined predicted value obtained by a certain calculation method is relatively accurate, the weight parameter of this calculation method can be appropriately increased to further improve the accuracy of the obtained load prediction value.
[0093] Taking N = 2 as an example, the load prediction value of the server after the first time period obtained can be expressed as P final=αP1 + βP2. Here, P1 and P2 respectively represent the undetermined predicted values determined by the first calculation method and the undetermined predicted values determined by the second calculation method. α and β respectively represent the weight coefficients of these two calculation methods, and α + β = 1. The specific values of α and β can be dynamically adjusted based on the specific performance of these two calculation methods, which ensures the flexibility of the solution of this application and is also conducive to improving the accuracy of the obtained load prediction value.
[0094] In addition, in addition to weighted superposition, other methods can also be used to determine the load prediction value of the server after the first time period using N undetermined predicted values. For example, methods such as voting method, weighted majority method, and validation dataset method can be adopted.
[0095] After the control device determines the load prediction values of each server after the first time period, it can adjust the power distribution of the power system that powers the server cluster, so that the total power distribution error is lower than a preset first threshold, and each server in the server cluster receives power supply from at least K power modules.
[0096] In the solution of this application, it is necessary to perform dynamic power distribution on the power system to ensure the rationality of power distribution, which is also conducive to ensuring the stability of server operation. When adjusting the power distribution of the power system, it can be adjusted in real time or periodically. In practical applications, considering that although the server load will fluctuate, it usually does not fluctuate frequently and violently in a short period of time. Therefore, usually, the power distribution of the power system can be adjusted every once in a while to avoid over-adjustment.
[0097] Moreover, when adjusting the power distribution of the power system, it is necessary to make the total power distribution error lower than a preset first threshold. The total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power system and the load prediction value of the server. That is to say, performing dynamic power distribution on the power system so that the total power distribution error is lower than a preset first threshold means that the power distribution situation of each server can conform to the load prediction value of the server. This also means that for any server, if it is predicted that the future load of the server will increase, the power supply can be increased in advance to avoid the situation of unstable voltage or degraded system performance caused by the increased load. On the contrary, if it is predicted that the future load of the server will decrease, the power supply can be reduced in advance to prevent power waste and improve energy utilization efficiency.
[0098] In practical applications, the power distribution error of a single server, for example, can be expressed by to represent, or can also be expressed by It is represented by, and the specific form can be set and adjusted according to actual needs, as long as it can effectively reflect the error between the total power obtained by the server from the power system and the load prediction value of the server. In this application, is used to represent the power distribution error of a single server. In this formula, x ij represents the power output from power module j to server i, and M is the total number of power modules currently working in the power system. is the load prediction value of server i.
[0099] If is used to represent the power distribution error of a single server, then the total power distribution error can be specifically expressed as , where N here represents the total number of servers. That is to say, in this specific implementation, after power distribution, it is required that is lower than a preset first threshold. It can be understood that, ideally, should be equal to 0.
[0100] And in the solution of this application, after power distribution, each server in the server cluster should receive power supply from at least K power modules. K is a positive integer not less than 2. This is because if each server receives power supply from 2 or more power modules, even if a single power module supplying power to the server fails, the server will not immediately lose all power supply, which is also beneficial to ensuring the stability of server operation to a certain extent and achieving the purpose of redundancy.
[0101] In a specific implementation of the present invention, the control device can specifically adjust the power distribution of the power system as follows: The control device adjusts the power distribution of the power system according to a preset constraint rule.
[0102] In this implementation, when adjusting the power distribution of the power system, it is also required to meet the constraint rule. And the constraint rule specifically includes: For any power module in the power system, the total power provided by the power module does not exceed the maximum power output limit value of the power module.
[0103] This is because in some cases, after power distribution, some power modules will provide a lot of power, while some power modules will provide less power, and even in some cases, the power provided by some power modules exceeds the limit of the power module. Therefore, in this implementation, it is required to meet the constraint rule when adjusting the power distribution, so that after the power distribution is completed, for any power module in the power system, the total power provided by the power module does not exceed the maximum power output limit value of the power module.
[0104] For each power supply module, its maximum power output limit value can be determined in advance. For example, the maximum output power set when the power supply module leaves the factory can be directly used as the maximum power output limit value of the power supply module in this implementation manner. Of course, in some cases, the power supply module may be adjusted. For example, a power amplifier is added to the output end of the power supply module to allow the power supply module to output greater power. Then, according to the actual situation, the maximum power output limit value of each power supply module can be adjusted.
[0105] In this implementation manner, since the total power provided by the power supply module does not exceed the maximum power output limit value of the power supply module, the power supply safety of each power supply module can be effectively guaranteed.
[0106] Furthermore, in a specific implementation manner of the present invention, the constraint rule may further include:
[0107] The total power obtained by a single server from the power supply system is greater than A times the load prediction value of the server; where A is a preset magnification value and A>1.
[0108] This implementation manner takes into account that in the solution of the present application, the current power distribution is based on the load prediction values of each server, that is, the current power distribution of each server is based on the future load conditions of each server. Therefore, in order to further ensure the power supply stability, in this implementation manner, the total power obtained by a single server from the power supply system is greater than A times the load prediction value of the server, so that when the load suddenly increases to a certain extent, the power supply stability of the server will not be damaged. For example, when A is set to 120%, the constraint rule of this implementation manner can be expressed as . is the load prediction value of server i, and x ij represents the power output from power supply module j to server i.
[0109] For example, in a specific scenario, there are 3 power supply modules in the power supply system. The maximum power output limit value of power supply module 1 is 1500W, the maximum power output limit value of power supply module 2 is 800W, and the maximum power output limit value of power supply module 3 is 1000W. For example, the load prediction value of server 1 is 800W, 800×1.2 = 960W. The load prediction value of server 2 is 500W, 500×1.2 = 600W. The load prediction value of server 3 is 700W, 700×1.2 = 840W. The load prediction value of server 4 is 300W, 300×1.2 = 360W.
[0110] After adjusting the power distribution of the power supply system that powers the server cluster, for example, in this implementation, Server 1 obtains 600W from Power Module 1 and 360W from Power Module 2, for a total of 960W. Server 2 obtains 500W from Power Module 1 and 100W from Power Module 2, for a total of 600W. Server 3 obtains 300W from Power Module 1 and 540W from Power Module 3, for a total of 840W. Server 4 obtains 200W from Power Module 2 and 160W from Power Module 3, for a total of 360W.
[0111] In this example, for Power Module 1, the total power provided is 600W + 500W + 300W = 1400W, which does not exceed its maximum power output limit value of 1500W. For Power Module 2, the total power provided is 360W + 100W + 200W = 660W, which does not exceed its maximum power output limit value of 800W. For Power Module 3, the total power provided is 540W + 160W = 700W, which does not exceed its maximum power output limit value of 800W.
[0112] And it can be seen that in this example, after adjusting the power distribution of the power supply system, the total power distribution error reaches the minimum value, and each server receives power supply from at least 2 power modules, and also meets the above-mentioned constraint rules.
[0113] In the solution of this application, it is necessary to be able to adjust the power distribution of the power supply system. There can be various specific implementation methods. For example, in one specific implementation, for any power module in the power supply system, the power module can have multiple output interfaces, and each output interface can be used to connect to a server to supply power to the corresponding server. For different output interfaces of the power module, different power outputs can be provided. For example, multiple sets of adjustable windings can be set on the secondary side of the transformer in the power module, and by adjusting each set of adjustable windings, the power adjustment of the corresponding output interface can be achieved. Another example is that for different output interfaces of the power module, the power control of the output interface can also be achieved through methods such as PWM pulse width modulation. There are various specific implementation methods, which do not affect the implementation of the present invention. And because the power module can have multiple output interfaces for connecting different servers and the output power of each output interface is adjustable, therefore, by adjusting each power module in the power supply system, the adjustment of the power distribution of the power supply system is realized, and after the adjustment, it can meet the various requirements described above.
[0114] In one specific implementation of the present invention, a fault monitoring device can also be included, which is used to monitor the fault status of each power module in the power supply system, and when a power module fault is detected, a preset fault response operation is executed.
[0115] This implementation mode takes into account that the fault status of each power module in the power supply system can also be monitored by a fault monitoring device. If a power module fails, the control device can perform a preset fault response operation, thereby ensuring the reliability of the power supply system and also facilitating the operation stability of the server.
[0116] The specific content of the fault response operation can be set and adjusted according to actual needs. For example, in a specific implementation mode of the present invention, the power supply system further includes a standby power module. Correspondingly, the control device performing the preset fault response operation can specifically include:
[0117] Determine the number of power modules that have failed;
[0118] Select the same number of standby power modules from the standby power modules in the standby state to take over the work of the failed power modules, and return to perform the operation of adjusting the power distribution of the power supply system for the server cluster.
[0119] This implementation mode takes into account that a certain number of standby power modules can be preset in advance. If a power module in the power supply system fails, the same number of standby power modules can be selected based on the number of faults to take over the work of the failed power modules. In practical applications, when a power module fails, the standby power module can be immediately started to take over the work of this power module, and the power switch can be completed within milliseconds to ensure that the power supply of the server is not interrupted.
[0120] And this implementation mode takes into account that after selecting the same number of standby power modules to take over the work of the failed power modules, since the power that the standby power module can provide may not be exactly the same as that of the original power module, therefore, the power distribution of the power supply system can be immediately re-performed to ensure the operation stability of the server.
[0121] There can be various specific implementation manners for determining the power module failure. For example, when abnormal current / voltage, overheating and other situations are detected, the power module failure can be determined. In addition, it can be understood that in practical applications, after the power module has been repaired, it can be re-connected to the power supply system, and through re-performing the power distribution of the power supply system, this power module can re-power the server.
[0122] In a specific implementation mode of the present invention, a communication device can also be included, and the communication device is connected to the control device:
[0123] The control device is further configured to: receive a power distribution instruction through the communication device, and adjust the power distribution of the power supply system for the server cluster according to the power distribution instruction.
[0124] In this implementation, it is considered that a communication device can be provided in the power supply system, and then the power distribution of the power supply system can be adjusted based on the power distribution instruction sent by the host computer received by the communication device. The communication device can support various industrial standard communication protocols, including Ethernet, CAN bus, RS-485, etc., to ensure compatibility.
[0125] In this implementation, the power distribution of the power supply system can be directly adjusted according to the power distribution instruction, which ensures the flexibility of the solution. It should be noted that since the power distribution of the power supply system is adjusted according to the power distribution instruction, relevant constraint rules can be ignored. Of course, generally speaking, the power distribution instructions are relatively reasonable power distribution instructions sent by the staff according to actual needs. After adjusting the power distribution of the power supply system according to the power distribution instruction, the power supply conditions of each server usually also meet the requirements of relevant constraint rules.
[0126] In a specific implementation of the present invention, the control device can also be used for:
[0127] For any server, determine the input power value currently required by the server;
[0128] Judge whether the error between the input power value and the load prediction value of the server exceeds a preset error range;
[0129] If so, use the input power value currently required by the server to replace the load prediction value of the server, and perform the operation of adjusting the power distribution of the power supply system that powers the server cluster.
[0130] As can be seen from the above description, in the solution of this application, the power supply system is dynamically power-distributed, so that the total power distribution error is lower than a preset first threshold, which means that the power distribution situation of each server can conform to the load prediction value of the server. That is to say, for any server, if it is predicted that the future load of the server will increase, the power supply can be increased in advance to avoid the situation of unstable voltage or degraded system performance caused by the increased load. On the contrary, if it is predicted that the future load of the server will decrease, the power supply can be reduced in advance to prevent power waste and improve energy utilization efficiency. That is to say, in the solution of this application, in fact, the current power distribution of the power supply system is adjusted based on the future load of each server. Such a method usually has good effects, but in some special scenarios, it may affect the stability of the server operation. Such a special scenario is specifically that the server load has an unexpected mutation. At this time, the server should not be powered according to the predicted load prediction value of the server, but should be powered immediately according to the current load situation of the server.
[0131] In this implementation, for any server, the input power value currently required by the server can be determined. For example, due to a sudden computing task, the input power value required by a certain server suddenly increases to 1000W, far exceeding its load prediction value of 600W. Therefore, it can be determined that the error between the input power value of the server and the load prediction value of the server exceeds the preset error range. At this time, the load prediction value of 600W can be ignored, and instead, the input power value of 1000W currently required by the server is directly used as the load prediction value of the server, that is, the input power value of 1000W required by the server is used to replace the original load prediction value of 600W of the server, and a power distribution adjustment is immediately performed so that the server can be allocated sufficient power. Another example is that due to a sudden stop of a certain task on a server, the input power value required suddenly drops to 500W, far lower than its load prediction value of 800W. Therefore, it can be determined that the error between the input power value of the server and the load prediction value of the server exceeds the preset error range. At this time, the load prediction value of 800W can be ignored, and instead, the input power value of 500W currently required by the server is directly used as the load prediction value of the server, and a power distribution adjustment is immediately performed to avoid power waste.
[0132] The specific value of the error range can be set and adjusted as needed. There can also be various specific implementation methods for determining the input power value currently required by the server. For example, it can be determined based on the current business of the server. Another example is that it can be determined based on the CPU status, memory status, number of IOs, etc. of the server.
[0133] Applying the technical solution provided by the embodiments of the present invention, considering that the traditional fixed power distribution method is prone to energy waste and will affect the stability of server operation. In this regard, in the solution of this application, dynamic adjustment of power distribution will be performed. Specifically, in the solution of this application, it is necessary to monitor the load status of each server in the server cluster, and then based on the load status of each server, determine the load prediction value of each server after a first period of time at the current moment. This is because the load will change, and the load prediction value of each server after the first period of time is determined, so as to achieve dynamic power distribution, which is beneficial to ensuring the rationality of power distribution and thus beneficial to ensuring the stability of server operation.
[0134] When adjusting the power distribution of the power supply system that powers the server cluster, it is necessary to make the total power distribution error lower than a preset first threshold. The total power distribution error represents the sum of the power distribution errors of each server in the server cluster. That is to say, making the total power distribution error lower than the preset first threshold means making the power distribution situation of each server conform to the load prediction value of that server. This also means that when it is predicted that the future load of the server will increase, the power supply can be increased in advance, thereby avoiding the situation where voltage instability or system performance degradation is caused by the increased load. Conversely, when it is predicted that the future load of the server will decrease, the power supply can be reduced in advance to prevent power waste and improve energy utilization efficiency.
[0135] Furthermore, in order to ensure power supply stability, it is also necessary to make each server in the server cluster receive power supply from at least K power modules. K is a positive integer not less than 2, which means that each server will receive power supply from 2 or more power modules. In this way, when a single power module that powers the server fails, the server will not immediately lose all power supply, which is also beneficial to ensuring the stability of server operation to a certain extent.
[0136] In summary, the solution of this application can effectively achieve power management, improve energy utilization efficiency, and ensure the stability of server operation.
[0137] Corresponding to the above power supply system embodiment, the embodiment of the present invention also provides a power management method, which can be mutually corresponding and referred to with the above text.
[0138] See Figure 1 As shown, this power management method can be applied to a power management device and includes the following steps:
[0139] Step S101: Monitor the load status of each server in the server cluster;
[0140] Step S102: Based on the load status of each server, determine the load prediction value of each server after the first time period; wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed;
[0141] Step S103: Adjust the power distribution of the power supply system that powers the server cluster to make the total power distribution error lower than a preset first threshold and make each server in the server cluster receive power supply from at least K power modules;
[0142] Among them, the power supply system includes multiple power modules, K is a positive integer not less than 2, the total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power supply system and the load prediction value of the server.
[0143] Corresponding to the above embodiments, the embodiments of the present invention further provide a power management device, a computer-readable storage medium, and a computer program product, which can be mutually corresponded and referred to with the above text.
[0144] See Figure 3 As shown, the device may include:
[0145] A memory 301 for storing a computer program;
[0146] A processor 302 for executing the computer program to implement the steps of the power management method in any of the above embodiments.
[0147] The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by the processor, the steps of the power management method in any of the above embodiments are implemented.
[0148] Refer to Figure 4 There is a computer program 41 stored on the computer-readable storage medium 40. When the computer program 41 is executed by the processor, the steps of the power management method in any of the above embodiments are implemented. The computer-readable storage medium 40 mentioned here includes RAM (Random Access Memory), memory, ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), registers, hard disks, removable disks, or any other form of storage medium well-known in the technical field.
[0149] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0150] Professionals may further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention, and the description of the above embodiments is only used to help understand the technical solution and its core idea of the present invention. It should be pointed out that for ordinary technicians in the technical field, without departing from the principle of the present invention, the present invention can also be improved and modified in a number of ways, and these improvements and modifications also fall within the scope of protection of the present invention.
Claims
1. A power supply system, characterized in that, Including: A plurality of power modules for powering a server cluster, and a power management device connected to each of the power modules; The power management device includes: A sensing device for monitoring the load status of each server in the server cluster; A control device connected to the sensing device, configured to determine a load prediction value of each server after a first time period based on the load status of each server; adjust the power distribution of the power system so that the total power distribution error is lower than a preset first threshold, and ensure that each server in the server cluster receives power supply from at least K power modules; Wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed; K is a positive integer not less than 2, and the total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power system and the load prediction value of the server; Specifically, the sensing device is configured to: Monitor the load status of each server in the server cluster in the most recent second time period; Correspondingly, the control device determines the load prediction value of each server after a first time period based on the load status of each server, including: The control device determines the load prediction value of each server after a first time period based on the load status of each server in the most recent second time period; Specifically, the sensing device is configured to: Monitor each server in the server cluster, and for any one of the servers, determine a central processing unit utilization rate sequence for reflecting the change of the central processing unit utilization rate of the server in the most recent second time period, a memory occupancy rate sequence for reflecting the change of the memory occupancy rate of the server in the most recent second time period, a request quantity sequence for reflecting the change of the request quantity of the server in the most recent second time period, and a network bandwidth sequence for reflecting the change of the network bandwidth of the server in the most recent second time period, and use the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request quantity sequence, and the network bandwidth sequence as the load status of the server monitored in the most recent second time period; The control device determines the load prediction value of each server after a first time period based on the load status of each server in the most recent second time period, including: For any one of the servers in the server cluster, input the load status of the server into a pre-trained first model to obtain a load prediction result output by the first model for reflecting the load status of the server after a first time period; wherein, the load prediction result includes the central processing unit utilization rate, memory occupancy rate, request quantity, and network bandwidth of the server after a first time period; Input the load prediction result into a pre-trained second model to obtain the load prediction value output by the second model, which is used as the determined load prediction value of the server after the first time period. The control device is further configured to: For any server, determine the input power value currently required by the server. Determine whether the error between the input power value and the load prediction value of the server exceeds a preset error range. If so, use the input power value currently required by the server to replace the load prediction value of the server, and perform an operation to adjust the power distribution of the power system that powers the server cluster. Among them, for any power module in the power system, the power module has multiple output interfaces, and each output interface is used to connect to a server to supply power to the corresponding server; for different output interfaces of the power module, the power adjustment of the output interface is achieved by setting multiple sets of adjustable windings on the secondary side of the transformer or by PWM pulse width modulation.
2. The power supply system according to claim 1, characterized in that, It further includes: A fault monitoring device for monitoring the fault status of each power module in the power system and sending it to the control device. The control device is connected to the fault monitoring device and is further configured to perform a preset fault response operation when the fault monitoring device monitors a fault of the power module.
3. The power supply system according to claim 2, wherein The power system further includes a standby power module. Correspondingly, the control device performing the preset fault response operation includes: Determine the number of the faulty power modules. Select the same number of standby power modules from the standby power modules in the standby state to take over the work of the faulty power modules, and adjust the power distribution of the power system.
4. The power supply system according to claim 1, characterized in that, The control device adjusting the power distribution of the power system includes: The control device adjusts the power distribution of the power system according to a preset constraint rule. Among them, the constraint rule includes: for any power module in the power system, the total power provided by the power module does not exceed the maximum power output limit value of the power module.
5. The power supply system according to claim 4, wherein The constraint rule further includes: The total power obtained by a single server from the power system is greater than A times the load prediction value of the server; where A is a preset magnification value and A is greater than 1.
6. The power supply system according to claim 1, wherein It further includes a communication device, and the communication device is connected to the control device. The control device is further configured to: receive a power distribution instruction through the communication device and adjust the power distribution of the power system according to the power distribution instruction.
7. The power supply system according to claim 1, characterized in that The control device determining the load prediction value of each server after the first time period based on the load status of each server in the most recent second time period includes: For any server in the server cluster, based on the load status of the server in the most recent second time period, respectively determine N pending prediction values of the server after the first time period through N different calculation methods; where N is a positive integer not less than 2. Based on N pending prediction values, determine the load prediction value of the server after the first time period.
8. The power supply system according to claim 7, wherein Based on N pending prediction values, determining the load prediction value of the server after the first time period includes: Based on the N pending prediction values and the weight parameters of N different calculation methods respectively, by weighted superposition of the N pending prediction values, obtain the load prediction value of the server after the first time period.
9. A power management method, characterized in that, Includes: Monitor the load status of each server in the server cluster; Based on the load status of each server, determine the load prediction value of each server after the first time period; wherein, the starting moment of the first time period is the moment when the load status monitoring of each server is completed; Adjust the power distribution of the power supply system that powers the server cluster, so that the total power distribution error is lower than a preset first threshold, and each server in the server cluster receives power supply from at least K power modules; Wherein, the power supply system includes a plurality of power modules, K is a positive integer not less than 2, the total power distribution error represents the sum of the power distribution errors of each server in the server cluster, and the power distribution error of a single server represents the error between the total power obtained by the server from the power supply system and the load prediction value of the server; Monitoring the load status of each server in the server cluster includes: Monitor the load status of each server in the server cluster within the most recent second time period; Correspondingly, based on the load status of each server, determining the load prediction value of each server after the first time period includes: Based on the load status of each server within the most recent second time period, determine the load prediction value of each server after the first time period; Monitoring the load status of each server in the server cluster includes: Monitor each server in the server cluster, and for any one of the servers, determine a central processing unit utilization rate sequence for reflecting the change of the central processing unit utilization rate of the server within the most recent second time period, a memory occupancy rate sequence for reflecting the change of the memory occupancy rate of the server within the most recent second time period, a request quantity sequence for reflecting the change of the request quantity of the server within the most recent second time period, and a network bandwidth sequence for reflecting the change of the network bandwidth of the server within the most recent second time period, and use the central processing unit utilization rate sequence, the memory occupancy rate sequence, the request quantity sequence, and the network bandwidth sequence as the load status of the server monitored within the most recent second time period; Correspondingly, based on the load status of each server, determining the load prediction value of each server after the first time period includes: For any one of the servers in the server cluster, input the load status of the server into a pre-trained first model to obtain a load prediction result output by the first model, which is used to reflect the load status of the server after a first time period; wherein, the load prediction result includes the central processing unit utilization rate, memory occupancy rate, number of requests, and network bandwidth of the server after the first time period; Input the load prediction result into a pre-trained second model to obtain a load prediction value output by the second model, which is used as the determined load prediction value of the server after the first time period; It further includes: For any one of the servers, determine the input power value currently required by the server; Judge whether the error between the input power value and the load prediction value of the server exceeds a preset error range; If so, use the input power value currently required by the server to replace the load prediction value of the server, and perform an operation of adjusting the power distribution of the power supply system that powers the server cluster; Among them, for any one of the power modules in the power supply system, the power module has multiple output interfaces, and each output interface is used to connect to a server to supply power to the corresponding server; for different output interfaces of the power module, the power adjustment of the output interface is realized by setting multiple sets of adjustable windings on the secondary side of the transformer or by PWM pulse width modulation.
10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the power management method as claimed in claim 9 are implemented.
11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the power management method as claimed in claim 9 are implemented.
Citation Information
Patent Citations
Power distribution method and system based on multiple cameras
CN119172639A