Server power management and scheduling method and system based on deep reinforcement learning
By combining deep reinforcement learning and advanced mathematical tools, a server power management and scheduling method that can effectively handle complex states and action spaces is constructed, solving the shortcomings of existing methods in balancing energy efficiency and service quality, and achieving more efficient and stable power management.
Patent Information
- Application Number
- CN202510126836.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-27
AI Technical Summary
Existing server power management methods are difficult to effectively balance energy efficiency and service quality in the face of dynamically changing workloads, and there are many challenges in dealing with high-dimensional state spaces and complex action spaces.
A deep reinforcement learning method is adopted, combined with topology, group theory and Li group theory, to construct a deep reinforcement learning framework for server power management and scheduling. This method constructs server status representation through topology embedding algorithm, uses group theory to construct action space, uses Li group theory to construct Q function approximation, and uses transform optimization strategy and Riemann-Zeta function design to explore strategies.
It realizes effective processing of complex state space and action space, improves the approximate accuracy of the Q function, and realizes an intelligent exploration strategy, which can achieve a good balance between energy efficiency, service quality and system stability.
Smart Images

Figure CN120066776A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of server power supplies, and more specifically, to a server power management and scheduling method and system based on deep reinforcement learning. Background Art
[0002] With the rapid development of information technology, the scale and complexity of data centers have been continuously increasing, and server power management and scheduling have become a key issue that needs to be solved urgently. Traditional server power management methods mainly rely on fixed threshold strategies or simple load balancing algorithms, and these methods often perform poorly when facing dynamically changing workloads and are difficult to achieve a good balance between energy efficiency and service quality.
[0003] In recent years, with the progress of artificial intelligence technology, researchers have begun to try to apply machine learning methods to the field of server power management. Among them, the method based on reinforcement learning has received extensive attention because it can adaptively learn the optimal strategy. However, the existing reinforcement learning methods still face many challenges when dealing with high-dimensional state spaces and complex action spaces. For example, the simple Q-learning algorithm is often difficult to converge when the state space is large, and although the deep Q-network (DQN) improves the learning ability, it performs poorly when facing continuous action spaces.
[0004] In addition, the existing methods generally have the following problems: First, the representation of server states is often too simplified and it is difficult to capture the internal relationships and topological structures between states; second, the design of the action space lacks a theoretical basis and it is difficult to ensure the effectiveness and integrity of exploration; third, the approximation method of the Q function usually uses a simple neural network structure and it is difficult to ensure stability and generalization ability in complex environments; finally, the existing exploration strategies often use the simple ε-greedy method and it is difficult to achieve a good balance between exploration and exploitation.
[0005] These problems lead to the poor performance of existing methods in practical applications, especially when facing large-scale and highly dynamic data center environments, it is difficult to achieve multi-objective optimization of energy efficiency, service quality, and system stability. Therefore, there is an urgent need for new server power management and scheduling methods that can effectively handle high-dimensional state spaces, design reasonable action spaces, improve the approximation accuracy of the Q function, and implement intelligent exploration strategies. Summary of the Invention
[0006] The present invention aims to solve the above problems existing in the prior art and proposes a server power management and scheduling method and system based on deep reinforcement learning. By innovatively combining advanced mathematical tools such as topology, group theory, and Lie group theory, the method constructs a brand-new deep reinforcement learning framework that can effectively handle complex state spaces and action spaces in server power management and scheduling problems.
[0007] The present invention provides a server power management and scheduling method based on deep reinforcement learning, including:
[0008] An acquisition step, including:
[0009] Acquire the real-time load data, power consumption data, and power consumption device demand data of the server cluster;
[0010] A processing step, including:
[0011] Based on the real-time load data, use a topological embedding algorithm to construct a server state representation;
[0012] According to the server state representation, use group theory to construct an action space;
[0013] Based on the action space, use Lie group theory to construct a Q-function approximation;
[0014] According to the Q-function approximation, adopt A transformation optimization strategy;
[0015] Based on the optimization strategy, use the Riemann-Zeta function to design an exploration strategy;
[0016] An output step, including:
[0017] According to the exploration strategy, generate a server power management and scheduling plan, and output the plan.
[0018] Preferably, the acquisition step specifically includes:
[0019] Real-time collect the load metrics of the server cluster through a server monitoring system;
[0020] Obtain the power consumption data of the server cluster from a power management module;
[0021] Receive the power consumption device demand data input by a user.
[0022] Preferably, using the topological embedding algorithm to construct a server state representation specifically includes:
[0023] Define a server state set and a manifold
[0024] Construct a homeomorphic mapping φ: Map the server state from the manifold to the Euclidean space R d ;
[0025] Obtain the representation x of the server state in the Euclidean space i ∈Rd 。
[0026] Preferably, the construction of the action space using group theory specifically includes:
[0027] Define the action group G = {g j |g j : R → R d , j = 1, 2, …, m}, where each element g j represents a transformation in the Euclidean space;
[0028] Define the group action · such that g j ·x i = y i , where y i is the new state after executing the action g j .
[0029] Preferably, the construction of the Q-function approximation using Lie group theory specifically includes:
[0030] Define the Q-function Q: R d × G → R;
[0031] Parameterize the Q-function using the special orthogonal group SO(n) such that Q(x, g) = tr(R T WR);
[0032] Adopt the matrix exponential function update rule:
[0033] Preferably, the adoption of the transformation optimization strategy specifically includes:
[0034] Define the policy function π: R D → G;
[0035] Use transformation to parameterize the policy: where ad - bc = 1;
[0036] Optimize the objective function where x follows the state distribution under the policy π.
[0037] Preferably, the design of the exploration strategy using the Riemann-Zeta function specifically includes:
[0038] Define the exploration probability where ζ(s) is the Riemann-Zeta function;
[0039] Construct the exploration strategy π exp (x), execute the policy π(x) with probability 1 - p(x), and execute a random action with probability p(x).
[0040] Preferably, it further includes:
[0041] Based on the server power management and scheduling scheme, adjust the power supply of the server cluster;
[0042] Monitor the performance and energy consumption of the adjusted server cluster;
[0043] According to the monitoring results, update the deep reinforcement learning model.
[0044] Preferably, it further includes:
[0045] Set the energy consumption threshold and performance indicators;
[0046] When the energy consumption of the server cluster exceeds the energy consumption threshold or the performance is lower than the performance indicator, trigger the re - execution of the above - mentioned processing steps.
[0047] A server power management and scheduling system based on deep reinforcement learning for executing the above - mentioned method, includes:
[0048] An acquisition module, configured to acquire the real - time load data, power energy consumption data, and power consumption device demand data of the server cluster;
[0049] A processing module, configured to:
[0050] Based on the real - time load data, use the topological embedding algorithm to construct a server state representation;
[0051] According to the server state representation, use group theory to construct an action space;
[0052] Based on the action space, use Lie group theory to construct a Q - function approximation;
[0053] According to the Q - function approximation, adopt A transformation optimization strategy;
[0054] Based on the optimization strategy, use the Riemann - Zeta function to design an exploration strategy;
[0055] An output module, configured to generate a server power management and scheduling scheme according to the exploration strategy and output the scheme.
[0056] The beneficial effects of the present invention are mainly reflected in the following aspects:
[0057] First, the present invention constructs a server state representation using a topological embedding algorithm. This method can not only effectively reduce the dimension of the state space but also preserve the topological relationships between states. This enables the system to better understand and predict the overall state of the server cluster, thereby making more accurate management decisions. For example, when the load suddenly increases, the system can quickly identify the state changes of relevant servers and make corresponding adjustments, avoiding the common lag response problem in traditional methods.
[0058] Second, the present invention constructs an action space using group theory. This innovative design provides the system with rich and structured action choices. Compared with the traditional discrete action space, this method can more finely control the power management strategy of the server. For example, the system can achieve precise power regulation for individual servers instead of simple on / off operations, thereby maximizing energy efficiency while ensuring performance.
[0059] Third, the present invention constructs a Q-function approximation using Lie group theory. This method significantly improves the expressive power and learning stability of the Q-function. Compared with traditional neural network structures, the Lie group-based Q-function can better capture the internal structure of the state-action value function, thereby showing stronger generalization ability in complex environments. This means that the system can adapt to new load patterns faster and maintain stable performance during long-term operation.
[0060] In addition, the present invention adopts a transformation optimization strategy. This non-linear transformation enables the system to express complex strategies within a limited parameter space. This not only improves the expressive power of the strategy but also improves the convergence of the optimization process. In practical applications, this means that the system can find near-optimal management strategies faster and can adapt to more diverse server management scenarios.
[0061] Finally, the present invention designs an exploration strategy using the Riemann-Zeta function. This innovative method provides the system with an intelligent and adaptive exploration mechanism. Compared with the traditional ε-greedy method, the exploration strategy based on the Riemann-Zeta function can dynamically adjust the exploration probability according to the uncertainty of the current state, thereby achieving a better balance between exploration and exploitation. This enables the system to converge to a stable management strategy faster while maintaining its learning ability.
[0062] In summary, through a series of innovative mathematical tools and algorithm designs, the present invention constructs an efficient, stable, and highly adaptable server power management and scheduling system. These innovative points do not simply add up, but form an organic whole, with each part collaborating and complementing each other. For example, the state representation of topological embedding provides more meaningful inputs for the action space constructed by group theory, while the Lie group-based Q-function approximation enables learning in such a complex state-action space. The exploration strategies designed by transformation optimization and Riemann-Zeta function further improve the learning efficiency and adaptability of the system in this complex space.
[0063] This holistic innovative design enables the present invention to exhibit significant advantages in practical applications. It can not only greatly improve the energy efficiency of data centers but also ensure service quality and system stability simultaneously. This provides a brand-new solution for the intelligent and green management of modern large-scale data centers and has important theoretical significance and practical application value. Brief Description of the Drawings
[0064] Figure 1 It is a flowchart of the method of the present invention.
[0065] Figure 2 It is a logical block diagram of the acquisition module of the present invention.
[0066] Figure 3 It is a logical block diagram of the processing module of the present invention.
[0067] Figure 4 It is a logical block diagram of the output module of the present invention. Detailed Embodiments
[0068] Please refer to Figures 1 - 4 , the present invention provides a server power management and scheduling method and system based on deep reinforcement learning. This method realizes the intelligent and refined control of server cluster power management by innovatively combining deep reinforcement learning technology with advanced mathematical theory.
[0069] First, the method of the present invention includes an acquisition step, a processing step, and an output step. In the acquisition step, the system acquires real-time load data, power consumption data, and power consumption device demand data of the server cluster. These data are the basis for subsequent processing and provide the necessary input information for the deep reinforcement learning model.
[0070] In the processing step, the present invention employs a series of innovative algorithms to process the acquired data. First, based on real-time load data, a topological embedding algorithm is used to construct a server state representation. The purpose of this step is to map the complex server state into a more manageable mathematical space. Specifically, a homeomorphic mapping φ is defined:
[0071] φ:
[0072] where, is the manifold where the server state is located, and R d is the d-dimensional Euclidean space. Through this mapping, each server state s i can be represented as a point x i in the Euclidean space:
[0073] φ(s i ) = x i ,
[0074] This representation method not only preserves the topological structure of the server state but also makes subsequent mathematical processing more convenient. The topological embedding algorithm is used to map the complex server state into a more manageable mathematical space. Through this mapping, the complex state space is simplified into a form convenient for calculation. For example, in the R d space, it is convenient to calculate metrics such as the distance and similarity between states, which is crucial for subsequent decision-making processes. For instance, the load situation of the server can be evaluated by calculating the distance between CPU utilization and memory usage, thereby better allocating resources.
[0075] The load metrics of the server cluster (such as CPU usage, memory occupancy, network traffic, etc.) are collected in real-time through a server monitoring system. After preprocessing these data, the homeomorphic mapping φ is learned through a neural network (such as MLP or CNN), and the original state is mapped into the Euclidean space R d .
[0076] Group theory is used to construct the action space so that there are internal relationships between different management operations. Next, the present invention constructs the action space using group theory based on the server state representation. The innovation of this step lies in applying the abstract algebraic structure to the actual server management problem. An action group G is defined:
[0077] G = {g j |g j : R d → R d , j = 1, 2, …, m},
[0078] where, each g jRepresents a possible management action, such as adjusting the power supply of a server or reallocating tasks. The group action is defined as:
[0079] g j ·x i =y i ,
[0080] Here, y i represents the new server state after performing the action g j . The advantage of this representation is that it can capture the internal relationships between different management actions, enabling the deep reinforcement learning model to better understand and select appropriate actions.
[0081] This definition describes how actions change the system state, allowing the system to simulate and predict the effects of different management decisions. For example, the operation of adjusting the CPU frequency can be modeled as a transformation g j , which transforms the current state x i into the new state y i . This helps the system optimize resource allocation and improve energy efficiency. Data acquisition and processing:
[0082] Obtain the power consumption data of the server cluster from the power management module (such as real-time power consumption, voltage, current, etc. of each server). Use this data to train the model to identify which operations (such as reducing the CPU frequency) can reduce energy consumption while maintaining performance.
[0083] Lie group theory is used to construct the Q-function approximation to estimate the long-term reward of taking a certain action in a given state. After constructing the state representation and action space, the present invention uses Lie group theory to construct the Q-function approximation. The Q-function is a core concept in deep reinforcement learning, which estimates the long-term reward of taking a certain action in a given state. The special orthogonal group SO(n) is innovatively used to parameterize the Q-function:
[0084] Q(xg) = tr(R T WR),
[0085] where R ∈ SO(n), W is the weight matrix, and tr represents the trace of the matrix. This representation method can not only capture the complex relationships between states and actions but also ensure the smoothness and invariance of the Q-function.
[0086] The update rule of the Q-function adopts the matrix exponential function:
[0087]
[0088] Here, exp is the matrix exponential function, and ∈ is the learning rate. This update method can perform gradient descent on the manifold of the Lie group, thus ensuring the effectiveness of the update.
[0089] This parameterization method ensures the rotational invariance of the Q function, and the influence degree of different dimensions on the Q value can be flexibly controlled by adjusting the weight matrix W. For example, by adjusting W, the system can give priority to the energy consumption optimization of high-load servers, thereby improving the overall energy efficiency.
[0090] Obtain the historical operation data of the server cluster, including load changes, energy consumption changes, etc. Use this data to train the Q function model and continuously optimize the weight matrix W to adapt to different server load patterns.
[0091] Through the above steps, the method of the present invention can effectively learn and optimize the server power management strategy. In practical applications, this method can dynamically adjust the power supply according to the real-time state and load conditions of the server, thereby maximizing the energy efficiency while ensuring the service quality.
[0092] Preferably, in an embodiment of the present invention, the obtaining step can be more refined. Specifically, the load indicators of the server cluster are collected in real time through the server monitoring system, and these indicators may include CPU usage rate, memory occupancy, network traffic, etc. At the same time, the power consumption data of the server cluster is obtained from the power management module, and these data may include the real-time power consumption, voltage, current, etc. information of each server. In addition, the system also receives the power consumption device demand data input by the user, which may include information such as the expected service quality level and business peak hours.
[0093] This refined obtaining step can provide more comprehensive and accurate data support for subsequent processing. For example, by analyzing the time series of CPU usage rate, the system can predict the future load trend; by comparing the energy consumption data of different servers, the devices with low energy efficiency can be identified; by considering the user's demand data, the energy use can be optimized on the premise of ensuring the service quality.
[0094] Through this comprehensive and detailed data acquisition, the method of the present invention provides rich environmental information for the deep reinforcement learning model, thereby enabling more intelligent and efficient power management decisions. This not only improves the overall energy efficiency of the system, but also can be flexibly adjusted according to actual needs to meet the service requirements in different scenarios.
[0095] The method of the present invention adopts an innovative topological embedding algorithm when constructing the server state representation. Specifically, the algorithm first defines the server state set and the manifold The manifold can be understood as an abstract mathematical space for representing all possible server states. Next, the algorithm constructs a homeomorphic mapping φ to map the server state from the manifold Mapped to the Euclidean space R d . The core of this step lies in maintaining the topological relationship between states while simplifying the complex state space into a more manageable Euclidean space.
[0096] Preferably, in an embodiment of the present invention, the homeomorphic mapping φ can be implemented by a neural network. For example, a multi-layer perceptron (MLP) or a convolutional neural network (CNN) can be used to learn this mapping relationship. Such a design enables the system to automatically learn the most suitable state representation without manual feature design. Through this mapping, each server state s i can obtain its representation x i in the Euclidean space:
[0097] φ(s i ) = x i ,
[0098] The advantage of this representation method is that it not only preserves the topological structure of the original state space but also transforms the states into a form convenient for mathematical processing and machine learning. For example, in the space, it is convenient to calculate metrics such as the distance and similarity between states, which is crucial for subsequent decision-making processes. The method of the present invention cleverly utilizes the concept of group theory when constructing the action space. Specifically, the method defines an action group G = {g j |g j : R d →R d , j = 1, 2, …, m}, where each element g j represents a transformation in the Euclidean space. These transformations can correspond to actual server management operations, such as adjusting the CPU frequency, changing the memory allocation, or adjusting the network bandwidth, etc.
[0099] Preferably, in an embodiment of the present invention, the action group G can be designed as a Lie group, so that the continuous properties of the Lie group can be utilized to achieve smooth changes in the action space. For example, the special orthogonal group SO(n) or the special linear group SL(n) can be used to represent the actions. Such a design enables the system to make fine adjustments in the continuous action space rather than being limited to discrete action selections.
[0100] The method also defines a group action · such that g j ·x i = y i , where y i is the result of executing the action g jThe new state after that. This definition method is very intuitive and describes how actions change the system state. Through group actions, it is convenient to simulate and predict the effects of different management decisions, which provides a theoretical basis for subsequent decision optimization. In the construction process of the Q function, the method of the present invention innovatively applies Lie group theory. First, the method defines the Q function Q: R d ×G→R, which maps states and actions to a real value, representing the expected long-term reward for taking a certain action in a given state.
[0101] In a preferred embodiment of the present invention, the special orthogonal group SO(n) is used to parameterize the Q function, such that Q(x,g) = tr(R T WR). Here, R ∈ SO(n), W is the weight matrix, and tr represents the trace of the matrix. This parameterization method has several remarkable advantages: First, it ensures the rotational invariance of the Q function, which is particularly useful when dealing with high-dimensional state spaces; Second, by adjusting the weight matrix W, the influence degree of different dimensions on the Q value can be flexibly controlled. The method adopts the matrix exponential function as the update rule of the Q function:
[0102]
[0103] Here, exp is the matrix exponential function, and ∈ is the learning rate. This update method ensures that the updated R matrix still remains in the SO(n) group, thus ensuring the consistency of the Q function parameterization. At the same time, the use of the matrix exponential function enables the gradient update to be carried out in the tangent space of the Lie group, which is more reasonable theoretically and can also bring better convergence properties.
[0104] The transformation is used to parameterize the policy, enabling the system to adapt to complex server management scenarios. The method of the present invention introduces the transformation, a powerful mathematical tool, in policy optimization. First, the method defines the policy function π: R d →G, which maps states to actions. Innovatively, the method uses the transformation to parameterize the policy:
[0105]
[0106] Preferably, in an embodiment of the present invention, the characteristics of the policy can be controlled by adjusting the parameters a, b, c, d. For example, by increasing |a| and |d|, the policy can become more sensitive in certain regions of the state space; by adjusting b and c, the overall tendency of the policy can be changed. This flexible parameterization method enables the system to adapt to various complex server management scenarios.
[0107] Transformations can express complex non - linear strategies within a finite parameter space while maintaining good mathematical properties (such as conformal and bijective properties). For example, during peak server loads, by adjusting parameters a, b, c, d, the system can respond more sensitively to load changes, thereby adjusting strategies in a timely manner to optimize energy consumption.
[0108] Collect data on the power consumption device requirements input by users (such as information on the expected quality of service level, business peak hours, etc.). Use this data to train the policy function π and adjust it according to the actual operating conditions to maximize the expected long - term return J(π).
[0109] The optimization objective function of the method is defined as:
[0110]
[0111] where x follows the state distribution under policy π. This objective function measures the expected long - term return under the current policy. By maximizing this objective function, the system can continuously improve its power management strategy to achieve a better balance between energy efficiency and performance.
[0112] The Riemann - Zeta function is used to design the exploration strategy, adjust the exploration probability, and balance the relationship between exploitation and exploration. The method of the present invention adopts an innovative Riemann - Zeta function in the design of the exploration strategy. Specifically, the method defines the exploration probability where ζ(s) is the Riemann - Zeta function. The uniqueness of this design lies in using the characteristics of the Riemann - Zeta function to adjust the exploration probability.
[0113] Monitor the performance and energy consumption of the adjusted server cluster, and collect multi - dimensional monitoring data (such as service response time, request processing ability, CPU utilization, etc.). Dynamically adjust the exploration probability p(x) according to this data, enabling the system to flexibly switch between exploration and exploitation strategies in different states.
[0114] Preferably, in an embodiment of the present invention, the parameter s can be adjusted according to the specific requirements of the system. When the value of s is small, the exploration probability distribution is relatively uniform, which is conducive to extensive exploration in the entire state space; when the value of s is large, the exploration probability will concentrate in certain regions of the state space, which is conducive to in - depth exploration in known potentially high - return regions. For example, in the initial stage of the system, a smaller value of s (such as s = 2) can be selected to encourage extensive exploration; as learning progresses, the value of s can be gradually increased (such as s = 3 or 4) to concentrate exploration in potential regions. By adjusting the parameter s, a trade - off can be made between extensive exploration and concentrated exploration. This can ensure that the system can find the optimal strategy at different stages.
[0115] Based on the above exploration probability, the method constructs an exploration strategy π exp (x). This strategy executes the current optimal strategy π(x) with probability 1 - p(x) and executes a random action with probability p(x). This design cleverly balances the relationship between exploitation and exploration. Near the states with known higher rewards, the system tends to execute the current optimal strategy; while near the unknown or lower-reward states, the system is more likely to try random actions to discover potential better strategies. The method of the present invention not only focuses on the generation of strategies, but also pays attention to the feedback and adjustment after the execution of the strategies. Specifically, the method includes the step of adjusting the power supply of the server cluster based on the generated server power management and scheduling scheme. This step closely combines the theoretical model with the actual operation, ensuring the practical feasibility of the strategy.
[0116] Preferably, in an embodiment of the present invention, the adjustment of the power supply can include multiple levels. For example, at the macroscopic level, the total power consumption of the entire server cluster can be adjusted; at the mesoscopic level, the power distribution of different racks or regions can be adjusted; at the microscopic level, parameters such as the CPU frequency and memory frequency of individual servers can be finely adjusted. This multi-level adjustment strategy enables the system to optimize energy use at different granularities.
[0117] Subsequently, the method further includes the step of monitoring the performance and energy consumption of the adjusted server cluster. This step provides real-time feedback to the system, enabling the system to timely evaluate the effect of the adjustment strategy. Preferably, the monitoring metrics can include but are not limited to: service response time, request processing capacity, CPU utilization rate, memory utilization rate, network throughput, power efficiency (performance per watt), etc. These multi-dimensional monitoring data provide rich basis for subsequent strategy optimization.
[0118] Based on the monitoring results, the method further includes the step of updating the deep reinforcement learning model. This step reflects the adaptability and continuous learning ability of the method of the present invention. By continuously feeding the actual operation data back into the model, the system can dynamically adjust its decision-making strategy to adapt to the changing server load and environmental conditions.
[0119] The method of the present invention also introduces an intelligent trigger mechanism to ensure that the system always operates in the optimal state. Specifically, the method includes the steps of setting energy consumption thresholds and performance metrics. These thresholds and metrics serve as the reference benchmarks for the system operation state.
[0120] Preferably, in one embodiment of the present invention, the energy consumption threshold can be set to 80% of the rated power of the server, and the performance metric can be set such that the service response time does not exceed 100 milliseconds. These specific values can be adjusted according to the actual application scenario and requirements. For example, for high-performance computing tasks, a higher energy consumption threshold may be allowed in exchange for better performance; while for regular web services, more emphasis may be placed on energy consumption control.
[0121] The method further stipulates that when the energy consumption of the server cluster exceeds the energy consumption threshold or the performance is lower than the performance metric, the re-execution of the processing step is triggered. This design ensures that the system can respond promptly to changes in the environment. For example, when the load suddenly increases and causes a decline in service quality, the system will automatically re-evaluate the current state and generate a new management strategy; similarly, when the energy consumption rises abnormally, the system will also quickly adjust to optimize energy utilization.
[0122] Through this mechanism of dynamic adjustment and self-optimization, the method of the present invention can always maintain an efficient and stable operating state in a complex and changing server operating environment, achieving the best balance between energy consumption and performance.
[0123] The present invention also provides a server power management and scheduling system based on deep reinforcement learning. The design of this system corresponds to the above method and aims to achieve efficient and intelligent server power management.
[0124] The system includes an acquisition module 1, which is responsible for acquiring real-time load data, power consumption data, and power consumption device demand data of the server cluster. The design of the acquisition module 1 fully considers the diversity and real-time nature of the data, providing comprehensive and timely information support for subsequent processing. Preferably, the acquisition module 1 can collect data in real time through a distributed sensor network to ensure the accuracy and timeliness of the data. For example, a load monitor can be deployed on each server to collect metrics such as CPU usage rate and memory occupancy in real time; at the same time, an energy consumption monitoring device can be installed in the power supply unit to record real-time power consumption data.
[0125] The processing module 2 is the core component of the system and is responsible for performing a series of complex data processing and decision optimization tasks. First, based on the real-time load data, the processing module 2 constructs a server state representation using a topology embedding algorithm. This step is completed by the state representation sub-module 21, which realizes the conversion from the original data to an abstract mathematical representation. Preferably, the state representation sub-module 21 can adopt a deep neural network architecture, such as an autoencoder or a graph neural network, to capture the internal structure and characteristics of the server state.
[0126] Next, the action space constructor sub-module 22 in the processing module 2 constructs the action space using group theory based on the server state representation. This innovative design enables the system to rigorously define and operate possible management actions mathematically. Preferably, the action space constructor sub-module 22 can adopt a differentiable parameterization method to represent the action group, thus supporting end-to-end gradient optimization.
[0127] Based on the constructed action space, the Q-function approximation sub-module 23 constructs the Q-function approximation using Lie group theory. This step is the core of deep reinforcement learning. The Q-function approximation sub-module 23 estimates the long-term rewards of different state-action pairs through continuous learning and updating. Preferably, the Q-function approximation sub-module 23 can adopt a deep Q-network (DQN) or its variants, such as double DQN or prioritized experience replay DQN, to improve learning efficiency and stability.
[0128] The policy optimization sub-module 24 adopts a transformation optimization policy according to the Q-function approximation. This unique optimization method enables the system to express complex non-linear policies within a limited parameter space. Preferably, the policy optimization sub-module 24 can combine policy gradient methods, such as the REINFORCE algorithm or the Advantage Actor-Critic (A2C) algorithm, to further improve the effect of policy optimization.
[0129] Based on the optimized policy, the exploration policy design sub-module 25 designs the exploration policy using the Riemann-Zeta function. This innovative design provides a flexible exploration mechanism for the system, enabling it to balance the exploitation of known information and the exploration of unknown possibilities. Preferably, the exploration policy design sub-module 25 can dynamically adjust the parameters of the Riemann-Zeta function to adapt to the exploration requirements at different stages.
[0130] The output module 3 is responsible for generating and outputting the server power management and scheduling plan according to the exploration policy. The output module 3 is not just a simple data output, but also includes the visualization and interpretation functions of the plan. Preferably, the output module 3 can generate a detailed scheduling report, including information such as power configuration suggestions for each server, expected energy consumption savings, and performance impact assessment. In addition, the output module 3 can also provide an interactive graphical interface, allowing the administrator to fine-tune the output plan according to actual needs.
[0131] Through this modular design, the system of the present invention realizes the full - process automation from data acquisition to policy generation. Each module is carefully designed to be able to independently complete its specific function and seamlessly cooperate with other modules to jointly form an efficient and intelligent server power management and scheduling system. Such a system can not only significantly improve the energy efficiency of the data center but also ensure the quality of service, providing strong technical support for the operation of modern data centers. To verify the superiority of the server power management and scheduling method and its system based on deep reinforcement learning of the present invention, a series of experimental comparisons were carried out by the present invention. The following will detail the settings of the examples and comparative examples, the detection methods, and the result analysis.
[0132] Example 1 adopts the method of the present invention, including state representation by topological embedding, action space constructed by group theory, Q - function approximation constructed by Lie group theory, policy with transformation optimization, and exploration strategy based on the Riemann - Zeta function. This method is applied to a medium - sized data center composed of 100 servers, with the servers configured with dual - socket Intel Xeon processors and 128GB of memory.
[0133] Comparative Example 1 adopts the traditional fixed - threshold strategy for server power management. This strategy decides to turn on or off the servers according to a preset CPU utilization threshold, without considering the dynamic changes of the load and long - term optimization.
[0134] Comparative Example 2 adopts a simple deep Q - learning (DQN) method, using a multi - layer perceptron as the Q - network, but without adopting the special mathematical structure and optimization strategy of the present invention.
[0135] The test was continuously carried out for 30 days, during which different load scenarios were simulated, including daily fluctuations, sudden high loads, and low - load periods. The main detection indicators include energy efficiency (PUE, Power Usage Effectiveness), quality of service (QoS, represented by the average response time), system stability (represented by the standard deviation), and convergence speed (represented by the time required to reach a stable policy).
[0136] The detection methods are as follows:
[0137] 1. PUE: Calculated by dividing the total energy consumption of the data center by the energy consumption of IT equipment.
[0138] 2. QoS: Record and calculate the average response time of all requests.
[0139] 3. System stability: Calculate the standard deviation of the daily PUE within 30 days.
[0140] 4. Convergence speed: Record the number of days required for the algorithm to reach a stable policy (the change in PUE is less than 1% for 7 consecutive days).
[0141] The test results are shown in Table 1 as follows:
[0142] Table 1. Comparison Table of Test Results between Example 1 and Comparative Example 1 and Comparative Example 2
[0143] Index Example 1 Comparative Example 1 Comparative Example 2 PUE 1.15 1.42 1.28 QoS (ms) 75 120 95 Stability (σ_PUE) 0.02 0.08 0.05 Convergence rate (days) 5 N / A 12
[0144] It can be seen from the test results in Table 1 that the method of the present invention (Example 1) shows significant advantages in various indicators. The specific analysis is as follows:
[0145] 1. Energy efficiency: The PUE of Example 1 reached 1.15, which was reduced by 19.0% and 10.2% compared with Comparative Example 1 and Comparative Example 2 respectively. This indicates that the method of the present invention can more precisely regulate the server power supply and reduce unnecessary energy waste.
[0146] 2. Service quality: The average response time of Example 1 was 75 ms, which was reduced by 37.5% and 21.1% compared with Comparative Example 1 and Comparative Example 2 respectively. This shows that the method of the present invention not only saves energy but also ensures better service quality, which benefits from its accurate prediction of the load and reasonable resource allocation.
[0147] 3. System stability: The standard deviation of PUE in Example 1 was only 0.02, which was much lower than the other two methods. This reflects that the method of the present invention can maintain stable performance in the face of complex and changing load conditions and will not show large fluctuations.
[0148] 4. Convergence speed: Example 1 reached the stable strategy in only 5 days, while Comparative Example 2 required 12 days. As a fixed strategy, Comparative Example 1 has no learning process, so this index is not applicable. This indicates that the method of the present invention has a higher learning efficiency and can adapt to environmental changes faster.
[0149] These results fully prove the superiority of the method of the present invention. Its core innovation points, such as the state representation of topological embedding enabling the system to better understand the internal structure of the server state; the action space constructed by group theory providing a richer choice of management strategies; the approximation of the Q function constructed by Lie group theory improving the stability of learning; The transformation optimization strategy and the exploration strategy based on the Riemann-Zeta function have both significantly improved in terms of optimization effect and learning efficiency.
[0150] It is particularly worth noting that the method of the present invention can maintain excellent service quality while ensuring high energy efficiency, which solves the problem of energy efficiency and performance trade-off often faced by traditional methods. In addition, its fast convergence speed means that the system can quickly adapt to the new environment, which is particularly important in the dynamically changing data center scenario.
[0151] In summary, the method of the present invention has demonstrated excellent performance in practical applications, providing strong technical support for the intelligent and efficient management of modern data centers.
[0152] It should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A server power management and scheduling method based on deep reinforcement learning, characterized in that: include: The acquisition steps include: Obtain real-time load data, power consumption data, and power equipment demand data of the server cluster; Processing steps include: Based on the real-time load data, construct a server status representation using a topology embedding algorithm; constructing an action space using group theory according to the server state representation; Based on the action space, a Q function approximation is constructed using Lie group theory; According to the Q function approximation, using Change optimization strategy; Based on the optimization strategy, an exploration strategy is designed using the Riemann-Zeta function; Output steps include: According to the exploration strategy, a server power management and scheduling plan is generated and output.
2. The method according to claim 1, characterized in that The acquisition step specifically includes: Collecting the load index of the server cluster in real time through a server monitoring system; Acquire power consumption data of the server cluster from a power management module; Receive the power equipment demand data input by the user.
3. The method according to claim 1, characterized in that The use of the topology embedding algorithm to construct the server status representation specifically includes: Define server status collection and manifold Constructing a homeomorphic mapping Remove server status from the manifold Mapping to Euclidean space R d ; Get the representation x of the server state in Euclidean space i ∈R d .
4. The method according to claim 1, characterized in that: The use of group theory to construct the action space specifically includes: Define action group G = {g j |g j :R→R d ,j=1,2,…,m}, where each element g j represents a transformation in Euclidean space; Define a group action · such that g j ·x i =y i , where y i To perform action g j The new status after.
5. The method according to claim 1, characterized in that The use of Lie group theory to construct the Q function approximation specifically includes: Define the Q function Q:R d ×G→R; The Q function is parameterized using the special orthogonal group SO(n) so that Q(x,g)=tr(R T WR); Adopt matrix exponential function update rule:
6. The method according to claim 1, characterized in that The adoption The transformation optimization strategies include: Define the policy function π:R d →G; use Transform parameterization strategy: Where ad-bc=1; Optimizing the objective function Where x follows the state distribution under strategy π.
7. The method according to claim 1, characterized in that The exploration strategy designed by using the Riemann-Zeta function specifically includes: Defining exploration probability Where ζ(s) is the Riemann-Zeta function; Constructing exploration strategy π exp (x), executes policy π(x) with probability 1-p(x), and performs a random action with probability p(x).
8. The method according to claim 1, characterized in that Also includes: Adjusting the power supply of the server cluster based on the server power management and scheduling solution; Monitor the performance and energy consumption of the adjusted server cluster; According to the monitoring results, the deep reinforcement learning model is updated.
9. The method according to claim 1, characterized in that: Also includes: Set energy consumption thresholds and performance indicators; When the energy consumption of the server cluster exceeds the energy consumption threshold or the performance is lower than the performance index, the processing step is triggered to be re-executed.
10. A server power management and scheduling system based on deep reinforcement learning that executes the method according to any one of claims 1 to 9, characterized in that: include: An acquisition module is used to obtain real-time load data, power consumption data and power equipment demand data of the server cluster; Processing modules for: Based on the real-time load data, construct a server status representation using a topology embedding algorithm; constructing an action space using group theory according to the server state representation; Based on the action space, a Q function approximation is constructed using Lie group theory; According to the Q function approximation, using Change optimization strategy; Based on the optimization strategy, an exploration strategy is designed using the Riemann-Zeta function; The output module is used to generate a server power management and scheduling plan according to the exploration strategy and output the plan.
Citation Information
Patent Citations
Intelligent power module management method and system
CN117454093A
Task scheduling method and system of cloud computing cluster based on reinforcement learning
CN118409838A
Power grid topology optimization method and system based on search sorting
CN118539441A
Customer portrait key data mining method and system based on space-time big data
CN118797542A
Ray-based lightweight distributed reinforcement learning training platform design method
CN119151019A
Cited By
Singlechip intelligent power management system based on deep reinforcement learning
CN122239917A
Intelligent power management system based on deep reinforcement learning
CN122239917B