Device power supply intelligent regulation and control method and system based on large model application
By constructing a server load index and dynamically adjusting the power supply strategy, the problem of insufficient or excessive power supply in traditional equipment power supply control methods is solved, achieving fine-grained adaptive control and improving the system's energy efficiency and stability.
Patent Information
- Application Number
- CN202511441095.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-26
AI Technical Summary
Traditional power supply control methods cannot effectively cope with complex and ever-changing workloads, resulting in insufficient or excessive power supply. They cannot achieve fine-grained adaptive control, affecting business performance and wasting energy. They also lack collaborative analysis and closed-loop optimization of CPU/GPU utilization, temperature, and SLA requirements.
The intelligent power supply control method for devices based on large model applications constructs a server load index by acquiring power supply prediction models and server cluster operation data, divides high, medium and low load zones, and adjusts the power supply strategy according to prediction errors, temperature rise deviations and SLA default rates to achieve fine-grained adaptive control.
It enables fine-grained partitioning and power supply control of server clusters, improves energy efficiency, enhances system stability and reliability, prevents equipment overheating, and achieves closed-loop optimization based on SLA default feedback, realizing fine-grained management and elastic resource allocation at the server level.
Smart Images

Figure CN121209679A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of intelligent power supply regulation, and in particular to a device power supply intelligent regulation method and system based on a large model application. BACKGROUND
[0002] The device power supply intelligent regulation method is based on artificial intelligence, the Internet of Things and big data analysis technology, realizes precise control and efficient operation of a multi-source heterogeneous power supply system by monitoring the power supply and demand state in real time and dynamically optimizing the energy distribution strategy. The core lies in the fusion of the "source network load storage" collaborative mechanism, the use of prediction algorithms (such as load prediction and new energy output prediction) and optimization models (such as multi-objective optimization and reinforcement learning), and the dynamic adjustment of the scheduling priority of photovoltaic, energy storage and power grid energy, so as to solve the imbalance between supply and demand caused by load fluctuation, intermittent new energy and power grid stability demand in the traditional power supply system.
[0003] Traditional device power supply regulation methods mostly rely on static thresholds or simple predictions based on local historical data, which are difficult to cope with complex and variable workloads, and are prone to cause insufficient power supply to affect business performance and SLA (Service Level Agreement), or excessive power supply to cause a large amount of energy waste. In addition, the existing methods lack collaborative analysis and closed-loop optimization of CPU / GPU utilization, temperature, power and SLA requirements, and cannot realize fine-grained adaptive regulation, which has significant shortcomings in dealing with sudden loads, preventing equipment overheating and ensuring service stability. SUMMARY
[0004] To solve the above technical problems, a device power supply intelligent regulation method and system based on a large model application are provided, which solves the problem that the existing methods lack collaborative analysis and closed-loop optimization of CPU / GPU utilization, temperature, power and SLA requirements, and cannot realize fine-grained adaptive regulation, which has significant shortcomings in dealing with sudden loads, preventing equipment overheating and ensuring service stability.
[0005] To achieve the above purposes, the technical scheme adopted by the application is as follows: A device power supply intelligent regulation method based on a large model application, comprising: obtaining a power supply prediction model, running data of a server cluster and available power on the power supply side, wherein the running data includes CPU utilization, GPU utilization, power, server SLA and temperature; based on the running data of the server cluster, obtaining a server load index of all servers in the server cluster; according to the server load index, dividing the server cluster into server cluster partitions, wherein the server cluster partitions include a high-load area, a medium-load area and a low-load area; According to the power supply prediction model, the server cluster partition and the server load index, an initial power supply strategy is obtained; According to the initial power supply strategy, an execution effect is obtained; According to the execution effect, a prediction error, a temperature rise deviation and an SLA violation rate are obtained; According to the prediction error, the temperature rise deviation and the SLA violation rate, the initial power supply strategy is adjusted.
[0006] Preferably, the server cluster running data is used to obtain a server load index of all servers in the server cluster, specifically including: According to the power, a power stress is obtained; According to the CPU utilization rate and the GPU utilization rate, a resource utilization rate is obtained; According to the temperature, a temperature normalization factor is obtained; According to the server SLA, a service level agreement weight factor is obtained; According to the power stress, the resource utilization rate, the temperature normalization factor and the service level agreement weight factor, a server load index is obtained; The service level agreement weight factor is specifically: In the formula, The service level agreement weight factor is, The maximum allowed interruption time of the server; The server load index is specifically: In the formula, The server load index is, The power stress is, The resource utilization rate is, The temperature normalization factor is.
[0007] Preferably, according to the server load index, the server cluster is divided into a server cluster partition, specifically including: The maximum power and the average power of the server within a week are obtained; According to the maximum power and the average power, an instantaneous peak adjustment factor is obtained; The standard deviation of the resource utilization rate and the mean value of the resource utilization rate of the server within a week are obtained; According to the standard deviation and the mean value of the resource utilization rate, a load fluctuation amplitude adjustment factor is obtained; According to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, the server load index is adjusted; The first threshold and the second threshold are obtained by analyzing all the adjusted server load indexes, if the server load index of a server is greater than the first threshold, the server is divided into a high load area, if the server load index of a server is between the second threshold and the first threshold, the server is divided into a medium load area, and if the server load index of a server is less than the second threshold, the server is divided into a low load area; The instantaneous peak adjustment factor is specifically: In the formula, The instantaneous peak adjustment factor is specifically: The maximum power of the server within a week, The average power of the server within a week; The load fluctuation amplitude adjustment factor is specifically: In the formula, The load fluctuation amplitude adjustment factor is specifically: The standard deviation of the resource utilization rate of the server within a week, The average value of the resource utilization rate of the server within a week; The adjustment of the server load index is specifically: In the formula, The adjusted server load index, The server load index before adjustment.
[0008] Preferably, the initial power supply strategy is obtained according to the power supply prediction model, the server cluster partition and the server load index, and specifically includes: The server load index of the server is normalized; The server cluster partition power supply power is obtained according to the normalized server load index, and the server cluster partition power supply power includes the total power supply power of the corresponding load area; The power supply power of all servers is obtained according to the server cluster partition power supply power and the server load index; The power supply power demand of all servers is obtained based on the power supply prediction model; The power supply power of all servers is adjusted according to the power supply power demand of all servers; The initial power supply strategy is obtained according to the adjusted power supply power of all servers; The server cluster partition power supply power is specifically: In the formula, Choose 1, 2, 3. , , These represent the total power supply for high, medium, and low load zones, respectively. This corresponds to the number of servers in the load zone. For the corresponding load area Server load index of each server Available power on the power supply side; The specific power consumption of all servers is as follows: In the formula, Choose 1, 2, 3. , , These represent the total power supply for high, medium, and low load zones, respectively. For the corresponding load area Power consumption of each server This represents the number of servers in the corresponding load zone; Specifically, adjusting the power supply to all servers involves: In the formula, Choose 1, 2, 3. To adjust the power supply of the server, For the corresponding load area Power consumption of each server For the corresponding load area Power requirements of each server.
[0009] Preferably, adjusting the initial power supply strategy based on prediction error, temperature rise deviation, and SLA default rate specifically includes: Based on prediction error, temperature rise deviation, and SLA default rate, obtain the adjustment weights for all servers; Based on the aforementioned adjustment weights, obtain the sum of the adjustment weights; Based on the sum of the adjusted weights, obtain the adjusted power supply of all servers; Adjust the initial power supply strategy based on the adjusted power supply power of all servers; Based on the adjustment of the initial power supply strategy, obtain the execution effect after strategy optimization; Based on the execution results of the optimized strategy, conduct the next round of optimization and adjustments; Specifically, obtaining the adjustment weights of all servers involves: In the formula, Choose 1, 2, 3. the adjustment weight of the server corresponding to the load area, the prediction error of the server corresponding to the load area, the temperature rise deviation of the server corresponding to the load area, the SLA violation rate of the load area. wherein, the adjustment weight of the server corresponding to the load area, 1, 2, 3, wherein, 1, 2, 3, the adjustment weight of the server corresponding to the load area, the available power of the power supply side, the total adjustment weight of all servers.
[0010] Preferably, the execution effect after the strategy optimization is used for the next round of optimization adjustment, specifically including: the initial power supply strategy is issued to all servers; if a server needs to urgently reduce the power supply power due to abnormal temperature rise, the released power supply power is distributed in the load area according to the server load index; if the SLA violation rate exceeds the preset threshold, a non-critical task server is obtained; the power supply power of the server in the high load area is compensated by reducing the power supply power of the non-critical task server; the execution effect is continuously monitored, and the prediction error, temperature rise deviation and SLA violation rate are fed back to the adjustment process to perform the next round of optimization adjustment.
[0011] Further, an intelligent power supply control system based on a large model application is proposed, which is used to realize the intelligent power supply control method based on the large model application as described above, including: a main control module, the main control module is used to obtain a power supply prediction model, running data of a server cluster and available power of a power supply side, based on the running data of the server cluster, obtain server load indexes of all servers in the server cluster, divide the server cluster into server cluster partitions according to the server load indexes, obtain an initial power supply strategy according to the power supply prediction model, the server cluster partitions and the server load indexes, obtain an execution effect according to the initial power supply strategy, obtain a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjust the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate; a server partition module, configured to obtain a power stress according to power, obtain a resource utilization according to CPU utilization and GPU utilization, obtain a temperature normalization factor according to temperature, obtain a service level agreement weight factor according to a server SLA, obtain a server load index according to the power stress, the resource utilization, the temperature normalization factor and the service level agreement weight factor, obtain maximum power and average power of the server within a week, obtain an instantaneous peak adjustment factor according to the maximum power and the average power, obtain a standard deviation of resource utilization and a mean value of resource utilization of the server within the week, obtain a load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of resource utilization, adjust the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyze all adjusted server load indexes, and obtain a first threshold value and a second threshold value, and divide the server into high, medium and low load areas according to the first threshold value and the second threshold value; an adjusted initial power supply strategy module, configured to obtain an adjustment weight of all servers according to a prediction error, a temperature rise deviation and an SLA violation rate, obtain a total adjustment weight according to the adjustment weight, obtain an adjusted power supply power of all servers according to the total adjustment weight, adjust an initial power supply strategy according to the adjusted power supply power of all servers, obtain an execution effect after optimization of the strategy according to the adjusted initial power supply strategy, and perform optimization adjustment in the next round according to the execution effect after optimization of the strategy. a display module, which interacts with the main control module and is configured to output display.
[0012] Optionally, the main control module specifically comprises: an information acquisition unit, configured to obtain a power supply prediction model, operation data of a server cluster and available power on a power supply side, and obtain a server load index of all servers in the server cluster based on the operation data of the server cluster; an initial power supply strategy unit, configured to divide the server cluster into a server cluster partition according to the server load index, obtain an initial power supply strategy according to the power supply prediction model, the server cluster partition and the server load index, obtain an execution effect according to the initial power supply strategy, obtain a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjust the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate.
[0013] Optionally, the server partition module specifically comprises: a load index acquisition unit configured to acquire a power stress according to power, acquire a resource utilization rate according to CPU utilization rate and GPU utilization rate, acquire a temperature normalization factor according to temperature, acquire a service level agreement weight factor according to server SLA, and acquire a server load index according to the power stress, the resource utilization rate, the temperature normalization factor, and the service level agreement weight factor; a partition unit configured to acquire maximum power and average power of the server within a week, acquire an instantaneous peak adjustment factor according to the maximum power and the average power, acquire a standard deviation of resource utilization rate and a mean value of resource utilization rate of the server within the week, acquire a load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of resource utilization rate, adjust the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyze all adjusted server load indexes, acquire a first threshold value and a second threshold value, and divide the server into high, medium and low load zones according to the first threshold value and the second threshold value.
[0014] Optionally, the adjustment initial power supply strategy module specifically comprises: a power supply power acquisition unit configured to acquire an adjustment weight of all servers according to prediction error, temperature rise deviation and SLA violation rate, acquire a total adjustment weight according to the adjustment weight, and acquire adjusted power supply power of all servers according to the total adjustment weight; an optimization unit configured to adjust the initial power supply strategy according to the adjusted power supply power of all servers, acquire an execution effect after strategy optimization according to the adjusted initial power supply strategy, and perform next round of optimization adjustment according to the execution effect after strategy optimization.
[0015] Compared with the prior art, the present application has the following beneficial effects: The present application proposes a device power supply intelligent regulation and control method and system based on large model application, which realizes fine partitioning and power supply regulation of server clusters by introducing a large model for power supply prediction and constructing a server load index by comprehensively considering multi-dimensional data such as CPU utilization rate, GPU utilization rate, power, temperature and SLA, improves energy efficiency utilization, enhances system stability and reliability through the prediction results and weight adjustment mechanism of the large model, prevents equipment overheating through real-time temperature rise monitoring, and dynamically adjusts partitioning and power supply strategies according to real-time load, realizes fine-grained management and elastic resource allocation at the server level, and relies on SLA violation feedback closed-loop optimization. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 A device power supply intelligent regulation and control method flowchart based on large model application is proposed in the present application; Figure 2Flow chart for server cluster partition acquisition in the application; Figure 3 Structural block diagram for adjusting initial power supply strategy in the application; Figure 4 Structural block diagram of a device power supply intelligent regulation and control system based on large model application proposed in the application. DETAILED DESCRIPTION
[0017] The following description is used to disclose the application so that those skilled in the art can implement the application. The preferred embodiments in the following description are only used as examples, and other obvious modifications can be thought of by those skilled in the art.
[0018] Referring to Figure 1 - Figure 3 As shown in the figure, the device power supply intelligent regulation and control method based on large model application in the embodiment of the application comprises: Obtaining a power supply prediction model, running data of a server cluster and available power on the power supply side, wherein the running data includes CPU utilization, GPU utilization, power, server SLA and temperature; Based on the running data of the server cluster, obtaining a server load index of all servers in the server cluster; According to the server load index, the server cluster is divided into server cluster partitions, which include high load area, medium load area and low load area; According to the power supply prediction model, the server cluster partition and the server load index, an initial power supply strategy is obtained; According to the initial power supply strategy, an execution effect is obtained; According to the execution effect, a prediction error, a temperature rise deviation and an SLA violation rate are obtained; According to the prediction error, the temperature rise deviation and the SLA violation rate, the initial power supply strategy is adjusted.
[0019] Among them, based on the running data of the server cluster, the server load index of all servers in the server cluster is obtained, specifically including: According to the power, the power pressure is obtained; According to the CPU utilization and the GPU utilization, the resource utilization is obtained; According to the temperature, the temperature normalization factor is obtained; According to the server SLA, the service level agreement weight factor is obtained; According to the power pressure, the resource utilization, the temperature normalization factor and the service level agreement weight factor, the server load index is obtained; Among them, the service level agreement weight factor is specifically: In the formula, a service level agreement weight factor, a maximum outage time allowed for the server; wherein the server load index is specifically: wherein, is a power stress, is a power stress, is a resource utilization, is a temperature normalization factor; In the present embodiment, the power stress is specifically: wherein, is a power stress, is a peak power of the server, is an average power of the server, is a standard deviation of the server historical power; The resource utilization is specifically: wherein, is a resource utilization, is a historical average of CPU utilization when the server is working, is a historical average of GPU utilization when the server is working; The temperature normalization factor is specifically: wherein, is a temperature normalization factor, is a CPU temperature of the server, is a minimum working temperature of the CPU of the server, is a maximum working temperature of the CPU of the server.
[0020] The server cluster is divided into server cluster partitions according to the server load index, and specifically includes: obtaining the maximum power and the average power of the server within a week; obtaining an instantaneous peak adjustment factor according to the maximum power and the average power; obtaining the standard deviation of the resource utilization and the average of the resource utilization of the server within a week; obtaining a load fluctuation amplitude adjustment factor according to the standard deviation and the average of the resource utilization; adjusting the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor; The first threshold and the second threshold are obtained by analyzing all the adjusted server load indexes, if the server load index of a server is greater than the first threshold, the server is divided into a high-load area, if the server load index of a server is between the second threshold and the first threshold, the server is divided into a medium-load area, and if the server load index of a server is less than the second threshold, the server is divided into a low-load area; The instantaneous peak adjustment factor is specifically as follows: In the formula, The instantaneous peak adjustment factor is specifically as follows: The maximum power of the server within a week is specifically as follows: The average power of the server within a week is specifically as follows: The load fluctuation amplitude adjustment factor is specifically as follows: In the formula, The load fluctuation amplitude adjustment factor is specifically as follows: The standard deviation of the resource utilization rate of the server within a week is specifically as follows: The average value of the resource utilization rate of the server within a week is specifically as follows: The adjustment of the server load index is specifically as follows: In the formula, The adjusted server load index is specifically as follows: The server load index before adjustment is specifically as follows: In the embodiment, the first threshold is 0.75, and the second threshold is 0.25.
[0021] In the scheme, the server load index and the adjustment mechanism are constructed. In the traditional scheme, the server cluster is regarded as a whole for unified power supply regulation and control, which may cause resource mismatch and low energy efficiency in actual application process. For example, when a shopping peak is encountered, the corresponding server load surges and more power is needed to ensure performance, and some servers are in an idle state. In the non-partition mode, the power supply system cannot distinguish the demands of the two, so that the overall power supply is insufficient to cause service lag and SLA violation, or the overall power supply causes the energy of the servers in the idle state to be wasted. In the scheme, the comprehensive quantitative evaluation of the server load state is realized by fusing the multi-dimensional factors such as power, computing resources, temperature and SLA contract, and the one-sidedness of a single index is avoided. Further, the historical peak and fluctuation adjustment are introduced, so that the server load index can not only reflect the instantaneous state, but also has the predictability and adaptability to the sudden load and performance fluctuation, and can accurately identify and distinguish the two kinds of loads, so as to realize the preferential power supply for the former and the energy-saving regulation for the latter.
[0022] The initial power supply strategy is obtained based on the power supply prediction model, server cluster partitioning, and server load index, specifically including: The server load index is normalized. Based on the normalized server load index, the power supply power of the server cluster partition is obtained, and the power supply power of the server cluster partition includes the total power supply power of the corresponding load area. Obtain the power consumption of all servers based on the power consumption of each server cluster partition and the server load index. Based on the power supply prediction model, obtain the power supply requirements of all servers; Adjust the power supply of all servers according to their power requirements. Based on the adjusted power supply of all servers, obtain the initial power supply strategy; Specifically, the power supply for the server cluster partition is as follows: In the formula, Choose 1, 2, 3. , , These represent the total power supply for high, medium, and low load zones, respectively. This corresponds to the number of servers in the load zone. For the corresponding load area Server load index of each server Available power on the power supply side; The specific power consumption of all servers is as follows: In the formula, Choose 1, 2, 3. , , These represent the total power supply for high, medium, and low load zones, respectively. For the corresponding load area Power consumption of each server This represents the number of servers in the corresponding load zone; Specifically, adjusting the power supply to all servers involves: In the formula, Choose 1, 2, 3. To adjust the power supply of the server, For the corresponding load area Power consumption of each server For the corresponding load area Power requirements of each server.
[0023] In the present scheme, the initial allocation of the total power supply side power is realized in a complex cluster environment with fairness and efficiency. In the traditional scheme, the difference of the load of different servers in the cluster is ignored, and in the actual application process, the problem of insufficient power supply and excess power supply coexists, for example, two servers belonging to the high load area, their SPS values are 0.9 and 0.8 respectively, if the total power of the partition is simply allocated, it may lead to insufficient power supply of the former and excess power supply of the latter, and in the present scheme, according to the server load index, the allocation is carried out again in each partition, which ensures that each unit of power can be invested in the node with the highest output, realizes the maximization of resource utilization, and the load of the server is fluctuant, the server load index can reflect these fluctuations, and the dynamic power allocation according to the server load index makes the power supply strategy close to the real-time state of each server, this kind of partition overall planning, the strategy of subdividing in the area realizes the differentiated power supply at the server level, and ensures the scientificity and executability of the strategy.
[0024] The initial power supply strategy is adjusted according to the prediction error, temperature rise deviation and SLA violation rate, and specifically comprises: According to the prediction error, the temperature rise deviation and the SLA violation rate, the adjustment weight of all servers is obtained; According to the adjustment weight, the total adjustment weight is obtained; According to the total adjustment weight, the adjusted power supply power of all servers is obtained; According to the adjusted power supply power of all servers, the initial power supply strategy is adjusted; According to the adjustment of the initial power supply strategy, the execution effect after the optimization of the strategy is obtained; According to the execution effect after the optimization of the strategy, the next round of optimization adjustment is carried out; Wherein, the adjustment weight of all servers is obtained as follows: In the formula, Take 1, 2, 3, The adjustment weight of the i-th server in the corresponding load area is The prediction error of the i-th server in the corresponding load area is The temperature rise deviation of the i-th server in the corresponding load area is The SLA violation rate of the corresponding load area is The adjustment weight of the i-th server in the corresponding load area is The prediction error of the i-th server in the corresponding load area is The temperature rise deviation of the i-th server in the corresponding load area is Wherein, the adjusted power supply power of all servers is obtained as follows: In the formula, Take 1, 2, 3, adjusted power supply power of the server, available power for the power supply side, total of the adjustment weights of all servers.
[0025] the execution effect after the optimization according to the strategy, and the next round of optimization adjustment is performed, specifically including: the initial power supply strategy is issued to all servers; if a server needs to urgently reduce the power supply power due to abnormal temperature rise, the released power supply power is distributed in the load area where the server is located according to the server load index; if the SLA violation rate exceeds the preset threshold, a non-critical task server is obtained; In this embodiment, the preset threshold is 0.2; the power supply power of the server in the high load area is compensated by reducing the power supply power of the non-critical task server; the execution effect of the system is continuously monitored, and the prediction error, temperature rise deviation and SLA violation rate are fed back to the current adjustment process for the next round of optimization adjustment; In this embodiment, an example of distributing the released power supply power in the load area where the server is located according to the server load index if a server needs to urgently reduce the power supply power due to abnormal temperature rise is as follows: The power supply power of server A (SPS1=0.24) in the low load area is 500W, the CPU temperature is 90℃ (the normal working temperature range of CPU: 35℃-85℃), and the best power supply power of server A at this time is obtained according to the power supply prediction model, which is 300W. There are also server B (SPS2=0.18), server C (SPS3=0.02) and server D (SPS4=0.20) in the low load area; The server load indexes of server B, server C and server D are normalized: SPS2=0.18 / (0.18+0.02+0.2)=0.45; SPS3=0.02 / (0.18+0.02+0.2)=0.05; SPS4=0.20 / (0.18+0.02+0.2)=0.50; Server B, server C and server D distribute the power supply power released by server A: Server B gets the power supply power of: (500-300)*0.45=90W; Server C gets the power supply power of: (500-300)*0.05=10W; Server D gets the power supply power of: (500-300)*0.5=100W; The adjusted power supply of server B, server C and server D is the sum of the original power supply and the allocated power supply; If the maximum power supply of server B is 400W, the power supply of server B before adjustment is 350W, and the power supply of server B after adjustment is 400W, the extra 40W will be allocated again in server C and server D.
[0026] One example of compensating the power supply of servers in the high-load area by reducing the power supply of non-critical task servers is: Server A in the high-load area (current power: 500W) needs to reach 700W due to task requirements; Non-critical server B (current power: 300W, minimum power: 150W) and non-critical server C (current power: 200W, minimum power: 50W) in the high-load area; Non-critical server B and non-critical server run at minimum power, and the released power supply will compensate server A; If the power supply released by the power supply of non-critical task servers cannot meet the power supply gap of server A, all the power supply released by non-critical task servers will compensate server A; If the power supply released by the power supply of non-critical task servers is greater than the power supply gap of server A, it will be supplemented in non-critical task servers according to the server load index from small to large, for example: Server A in the high-load area (current power: 500W) needs to reach 700W due to task requirements; Non-critical server B (current power: 300W, minimum power: 160W, SPS=0.89), non-critical server C (current power: 200W, minimum power: 50W, SPS=0.90) and non-critical server D (current power: 100W, minimum power: 50W, SPS=0.80) in the high-load area; For the power supply gap (200W) of server A, non-critical server D releases 50W, non-critical server B releases 140W, and non-critical server C releases 10W.
[0027] The continuous monitoring system executes the effect and feeds back the prediction error, temperature rise deviation and SLA violation rate to the current adjustment process for the next round of optimization adjustment: According to the prediction error, temperature rise deviation and SLA violation rate, the adjustment weight of all servers is obtained; According to the adjustment weight, the adjustment weight sum is obtained; According to the adjustment weight sum, the adjusted power supply of all servers is obtained; adjust the initial power supply strategy according to the adjusted power supply power of all servers; obtain an execution effect after strategy optimization according to the adjusted initial power supply strategy; perform next round of optimization adjustment according to the execution effect after strategy optimization.
[0028] In the scheme, a dynamic closed-loop optimization mechanism is constructed. The traditional scheme cannot cope with the dynamic changes of load and temperature, and often makes lagging adjustment after problems (such as overheating and SLA violation) occur. The scheme analyzes and predicts the error, temperature rise deviation and SLA violation rate, and quantifies them as the dynamic adjustment weight of each server, and then adjusts the power supply strategy globally. An emergency response and compensation mechanism is also introduced, for example, when a server overheats, the power released by the server is redistributed in the zone, or when the SLA is exceeded, the power of non-critical tasks is reduced to prioritize core business, thereby realizing the on-demand flow and precise delivery of energy while strictly following the total power budget. This self-sensing, self-decision and self-optimization capability improves the reliability and intelligence level of the system.
[0029] Referring to Figure 4 Further, in combination with the above-mentioned device power supply intelligent control method based on a large model application, a device power supply intelligent control system based on a large model application is provided, comprising: a main control module, the main control module is used for obtaining a power supply prediction model, running data of a server cluster and available power on the power supply side, obtaining a server load index of all servers in the server cluster based on the running data of the server cluster, dividing the server cluster into server cluster partitions according to the server load index, obtaining an initial power supply strategy according to the power supply prediction model, the server cluster partitions and the server load index, obtaining an execution effect according to the initial power supply strategy, obtaining a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjusting the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate; a server partition module, configured to obtain a power stress according to power, obtain a resource utilization according to CPU utilization and GPU utilization, obtain a temperature normalization factor according to temperature, obtain a service level agreement weight factor according to a server SLA, obtain a server load index according to the power stress, the resource utilization, the temperature normalization factor and the service level agreement weight factor, obtain maximum power and average power of the server within a week, obtain an instantaneous peak adjustment factor according to the maximum power and the average power, obtain a standard deviation of resource utilization and a mean value of resource utilization of the server within the week, obtain a load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of resource utilization, adjust the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyze all adjusted server load indexes, and obtain a first threshold value and a second threshold value, and divide the server into high, medium and low load areas according to the first threshold value and the second threshold value; an initial power supply strategy adjustment module, configured to obtain an adjustment weight of all servers according to a prediction error, a temperature rise deviation and an SLA violation rate, obtain a total adjustment weight according to the adjustment weight, obtain an adjusted power supply power of all servers according to the total adjustment weight, adjust an initial power supply strategy according to the adjusted power supply power of all servers, obtain an execution effect after optimization of the strategy according to the adjusted initial power supply strategy, and perform optimization adjustment in the next round according to the execution effect after optimization of the strategy. a display module, which interacts with the main control module and is configured to output display.
[0030] The main control module specifically comprises: an information acquisition unit, configured to obtain a power supply prediction model, running data of a server cluster and available power on a power supply side, and obtain a server load index of all servers in the server cluster based on the running data of the server cluster; an initial power supply strategy unit, configured to divide the server cluster into a server cluster partition according to the server load index, obtain an initial power supply strategy according to the power supply prediction model, the server cluster partition and the server load index, obtain an execution effect according to the initial power supply strategy, obtain a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjust the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate.
[0031] The server partition module specifically comprises: a load index acquisition unit configured to acquire a power stress according to power, a resource utilization rate according to CPU utilization rate and GPU utilization rate, a temperature normalization factor according to temperature, a service level agreement weight factor according to server SLA, and a server load index according to the power stress, the resource utilization rate, the temperature normalization factor, and the service level agreement weight factor; a partition unit configured to acquire maximum power and average power of the server within a week, acquire an instantaneous peak adjustment factor according to the maximum power and the average power, acquire a standard deviation of resource utilization rate and a mean value of resource utilization rate of the server within a week, acquire a load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of resource utilization rate, adjust the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyze all adjusted server load indexes, acquire a first threshold value and a second threshold value, and divide the server into high, medium and low load zones according to the first threshold value and the second threshold value.
[0032] The adjustment initial power supply strategy module specifically comprises: a power supply power acquisition unit configured to acquire an adjustment weight of all servers according to a prediction error, a temperature rise deviation, and an SLA violation rate, acquire a total adjustment weight according to the adjustment weight, and acquire adjusted power supply power of all servers according to the total adjustment weight; an optimization unit configured to adjust the initial power supply strategy according to the adjusted power supply power of all servers, acquire an execution effect after strategy optimization according to the adjusted initial power supply strategy, and perform next round of optimization adjustment according to the execution effect after strategy optimization.
[0033] In summary, the present application has the following advantages: The present application introduces a large model for power supply prediction, and constructs a server load index by comprehensively considering multi-dimensional data such as CPU utilization rate, GPU utilization rate, power, temperature, and SLA, so as to realize fine partitioning and power supply regulation of a server cluster, improve energy efficiency utilization, enhance system stability and reliability through prediction results and weight adjustment mechanism of the large model, prevent equipment overheating through real-time temperature rise monitoring, and dynamically adjust partitioning and power supply strategy according to real-time load, so as to realize fine-grained management and elastic resource allocation at the server level.
[0034] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited to the above-mentioned embodiments, and the above-mentioned embodiments and descriptions in the specification are only the principles of the present application. Various changes and improvements can be made without departing from the spirit and scope of the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A device power supply intelligent regulation method based on a large model application, characterized in that, The application comprises the following steps: obtaining power supply prediction model, running data of server cluster and available power on the power supply side, wherein the running data comprises CPU utilization, GPU utilization, power, server SLA and temperature; obtaining server load index of all servers in the server cluster based on the running data of the server cluster; dividing the server cluster into server cluster partitions according to the server load index, wherein the server cluster partitions comprise high-load area, medium-load area and low-load area; obtaining initial power supply strategy according to the power supply prediction model, the server cluster partitions and the server load index; obtaining execution effect according to the initial power supply strategy; obtaining prediction error, temperature rise deviation and SLA violation rate according to the execution effect; adjusting the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate.
2. The device power intelligent regulation method based on large model application according to claim 1, characterized in that, The application further comprises the following steps: obtaining power pressure according to the power; obtaining resource utilization rate according to the CPU utilization and the GPU utilization; obtaining temperature normalization factor according to the temperature; obtaining service level agreement weight factor according to the server SLA; obtaining server load index according to the power pressure, the resource utilization rate, the temperature normalization factor and the service level agreement weight factor; The service level agreement weight factor is specifically as follows: wherein is a service level agreement weight factor, is the maximum allowed outage time for the server; The server load index is specifically as follows: wherein is a server load index, is a power stress, is a resource utilization, is a temperature normalization factor.
3. The device power intelligent regulation method based on large model application according to claim 1, characterized in that, The application further comprises the following steps: obtaining maximum power and average power of the server within one week; obtaining instantaneous peak adjustment factor according to the maximum power and the average power; obtaining standard deviation of resource utilization rate and mean value of resource utilization rate of the server within one week; obtaining load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of the resource utilization rate; adjusting the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor; analyzing all adjusted server load indexes to obtain first threshold value and second threshold value, wherein if the server load index of a server is greater than the first threshold value, the server is divided into the high-load area, if the server load index of a server is between the second threshold value and the first threshold value, the server is divided into the medium-load area, and if the server load index of a server is less than the second threshold value, the server is divided into the low-load area; The instantaneous peak adjustment factor is specifically as follows: wherein is a transient peak adjustment factor, is the maximum power of the server in a week, is the average power of the server in a week; The load fluctuation amplitude adjustment factor is specifically as follows: In the formula, is a load fluctuation amplitude adjustment factor, is a standard deviation of resource utilization within a week of the server, is a mean value of resource utilization within a week of the server; The adjustment of the server load index is specifically as follows: wherein is the adjusted server load index, is the unadjusted server load index.
4. The device power intelligent regulation method based on large model application according to claim 1, characterized in that, The application further comprises the following steps: normalizing the server load index of the server; obtaining server cluster partition power supply power according to the normalized server load index, wherein the server cluster partition power supply power comprises total power supply power corresponding to the load area; obtaining power supply power of all servers according to the server cluster partition power supply power and the server load index; obtaining power supply power demand of all servers based on the power supply prediction model; According to the power supply power demand of all servers, the power supply power of all servers is adjusted; According to the adjusted power supply power of all servers, an initial power supply strategy is obtained; The power supply power of the server cluster partition is specifically: In the formula, Take 1, 2, 3, , , The total power supply of high, medium and low load areas, The number of servers in the corresponding load area, The server load index of the first server in the corresponding load area, The available power of the power supply side; The power supply power of all servers is specifically: In the formula, Take 1, 2, 3, , , The total power supply power of high, medium and low load area respectively, The power supply power of the first The number of servers in the corresponding load area, The number of servers in the corresponding load area, The adjustment of the power supply power of all servers is specifically: In the formula, Take 1, 2, 3, The power supply power of the adjusted server, The power supply power of the corresponding load area server, The power supply power of the corresponding load area server, The power supply power demand of the corresponding load area server, The power supply power demand of the corresponding load area server, 5. The device power intelligent regulation method based on large model application according to claim 1, characterized in that, The adjustment of the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate specifically includes: According to the prediction error, the temperature rise deviation and the SLA violation rate, an adjustment weight of all servers is obtained; According to the adjustment weight, an adjustment weight sum is obtained; According to the adjustment weight sum, an adjusted power supply power of all servers is obtained; According to the adjusted power supply power of all servers, the initial power supply strategy is adjusted; According to the adjustment of the initial power supply strategy, an execution effect after strategy optimization is obtained; According to the execution effect after strategy optimization, the next round of optimization adjustment is performed; The adjustment weight of all servers is specifically obtained as follows: In the formula, Choose 1, 2, 3. For the corresponding load area Adjusting the weight of each server For the corresponding load area Prediction error of each server For the corresponding load area Temperature rise deviation of each server The corresponding SLA default rate for the load area; The adjusted power supply power of all servers is specifically obtained as follows: In the formula, Take 1, 2, 3, The adjusted power supply power of the server, The available power of the power supply side, The total weight of all servers.
6. The device power intelligent regulation method based on large model application according to claim 5, characterized in that, The next round of optimization adjustment according to the execution effect after strategy optimization specifically includes: The initial power supply strategy is adjusted and issued to all servers; If a server needs to urgently reduce the power supply power due to abnormal temperature rise, the released power supply power is distributed in the load area according to the server load index; If the SLA violation rate exceeds the preset threshold, a non-critical task server is obtained; The power supply power of the server in the high load area is compensated by reducing the power supply power of the non-critical task server; The system execution effect is continuously monitored, and the prediction error, the temperature rise deviation and the SLA violation rate are fed back to the current adjustment process for the next round of optimization adjustment.
7. A device power supply intelligent regulation system based on large model application, used to implement the device power supply intelligent regulation method based on large model application according to any one of claims 1-6. It includes: The main control module is used to obtain a power supply prediction model, running data of a server cluster and available power on the power supply side, obtain a server load index of all servers in the server cluster based on the running data of the server cluster, divide the server cluster into server cluster partitions according to the server load index, obtain an initial power supply strategy according to the power supply prediction model, the server cluster partitions and the server load index, obtain an execution effect according to the initial power supply strategy, obtain a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjust the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate; a server partition module, configured to acquire a power stress according to power, acquire resource utilization according to CPU utilization and GPU utilization, acquire a temperature normalization factor according to temperature, acquire a service level agreement weight factor according to a server SLA, acquire a server load index according to the power stress, the resource utilization, the temperature normalization factor and the service level agreement weight factor, acquire maximum power and average power of the server within a week, acquire an instantaneous peak adjustment factor according to the maximum power and the average power, acquire a standard deviation of resource utilization and a mean value of resource utilization of the server within the week, acquire a load fluctuation amplitude adjustment factor according to the standard deviation and the mean value of resource utilization, adjust the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyze all adjusted server load indexes, and acquire a first threshold value and a second threshold value, and divide the server into high, medium and low load areas according to the first threshold value and the second threshold value; an initial power supply strategy adjustment module, configured to acquire an adjustment weight of all servers according to a prediction error, a temperature rise deviation and an SLA violation rate, acquire a total adjustment weight according to the adjustment weight, acquire an adjusted power supply power of all servers according to the total adjustment weight, adjust an initial power supply strategy according to the adjusted power supply power of all servers, acquire an execution effect after optimization of the strategy according to the adjusted initial power supply strategy, and perform optimization adjustment in the next round according to the execution effect after optimization of the strategy; a display module, which interacts with the main control module and is configured to output display.
8. The device power intelligent regulation system based on large model application of claim 7, wherein, The main control module specifically comprises: an information acquisition unit, configured to acquire a power supply prediction model, running data of a server cluster and available power on a power supply side, acquire a server load index of all servers in the server cluster based on the running data of the server cluster; an initial power supply strategy unit, configured to divide the server cluster into a server cluster partition according to the server load index, acquire an initial power supply strategy according to the power supply prediction model, the server cluster partition and the server load index, acquire an execution effect according to the initial power supply strategy, acquire a prediction error, a temperature rise deviation and an SLA violation rate according to the execution effect, and adjust the initial power supply strategy according to the prediction error, the temperature rise deviation and the SLA violation rate.
9. The device power intelligent regulation system based on large model application of claim 7, wherein, The server partition module specifically comprises: a load index acquisition unit, configured to acquire a power stress according to power, acquire resource utilization according to CPU utilization and GPU utilization, acquire a temperature normalization factor according to temperature, acquire a service level agreement weight factor according to a server SLA, and acquire a server load index according to the power stress, the resource utilization, the temperature normalization factor and the service level agreement weight factor; The partition unit is used for obtaining the maximum power and the average power of the server within a week, obtaining an instantaneous peak adjustment factor according to the maximum power and the average power, obtaining the standard deviation of the resource utilization and the mean of the resource utilization of the server within a week, obtaining a load fluctuation amplitude adjustment factor according to the standard deviation of the resource utilization and the mean, adjusting the server load index according to the instantaneous peak adjustment factor and the load fluctuation amplitude adjustment factor, analyzing all the adjusted server load indexes, obtaining a first threshold value and a second threshold value, and dividing the server into high, medium and low load areas according to the first threshold value and the second threshold value.
10. The device power intelligent regulation system based on large model application of claim 7, wherein, The adjustment initial power supply strategy module specifically comprises: The power supply power acquisition unit is used for obtaining the adjustment weight of all the servers according to the prediction error, the temperature rise deviation and the SLA violation rate, obtaining the total adjustment weight according to the adjustment weight, and obtaining the adjusted power supply power of all the servers according to the total adjustment weight; The optimization unit is used for adjusting the initial power supply strategy according to the adjusted power supply power of all the servers, obtaining the execution effect after the strategy optimization, and performing the next round of optimization adjustment according to the execution effect after the strategy optimization.