Workload Management System for Data Center Energy Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing energy consumption in high performance computing environments such as grids and clusters is challenging due to the heterogeneous nature of shared resources, multiple layers of schedulers, and the difficulty in reserving resources while maintaining performance and efficiency, leading to increased electricity usage and costs.
Innovation Solution
A system and method that monitors resource state, reports power and temperature consumption, interfaces with power management facilities, and implements intelligent policies to control power usage, including workload management software that communicates with resource managers to optimize resource allocation and reduce energy consumption by consolidating workloads, using idle servers efficiently, and scheduling tasks based on energy efficiency and cost.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If servers are kept running to handle workload requests, then service availability is maintained, but power consumption increases
Solution Approach 1:
The system dynamically adjusts server power states based on real-time workload conditions. Servers transition between active, standby, and off states depending on queue depth and predicted workload, optimizing the balance between service availability and power consumption
Solution Approach 2:
The system performs preliminary actions by pre-warming servers before predicted workload peaks and pre-cooling them before low-utilization periods. This anticipatory approach ensures servers are ready to handle requests when needed while minimizing power consumption during idle periods
2Productivity
If workload is distributed across multiple servers, then processing capacity is increased, but energy consumption increases
Solution Approach 1:
The system merges workload onto fewer servers when possible, consolidating tasks to maximize server utilization. By combining multiple workloads on single servers, the system reduces the total number of active servers and thereby decreases overall energy consumption while maintaining processing capacity
Solution Approach 2:
The system changes operational parameters by adjusting server power states (active, standby, off) and workload distribution strategies based on real-time conditions. This dynamic parameter adjustment optimizes the ratio of processing capacity to energy consumption
3Use of energy by moving object
If servers are placed in standby or off mode to save energy, then power consumption is reduced, but response time increases
Solution Approach 1:
The system performs preliminary warming of servers before predicted workload peaks, ensuring they are in active state and ready to process requests immediately. This anticipatory warming prevents response time penalties while still allowing servers to enter low-power states during confirmed idle periods
Solution Approach 2:
The system uses feedback from workload monitors and queue depth sensors to dynamically adjust server power states. When queues indicate incoming workload, the system activates servers in advance, using the queue feedback signal to trigger warming actions that prevent response time degradation
4Productivity
If more servers are activated to handle peak workload, then processing capacity is increased, but cost increases
Solution Approach 1:
The system changes the operational parameters of servers by adjusting power states and workload allocation based on predicted demand. This dynamic parameter management allows the system to maintain processing capacity during peaks while minimizing energy consumption and cost during low-utilization periods
Solution Approach 2:
The system dynamically reallocates workload and adjusts server activation based on real-time and predicted workload conditions. This dynamic management optimizes the balance between processing capacity and operational cost by activating servers only when and where needed
Data Source
AI summary
A system, method and non-transitory computer readable storage medium are disclosed for managing workload in a data center. The method includes receiving data related to at least one of a current state of workload in the compute environment at a current time and future workload scheduled to consume resources in the compute environment at a future time relative to the current time, wherein the compute environment comprises a plurality of nodes in which compute resources are reserved by a workload manager for consumption, and controlling a cooling system to selectively modify a temperature of at least one node in the compute environment based on the data.


