Predictive Server Load Distribution for Data Center Cooling Energy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face significant energy inefficiencies due to overprovisioning and peak performance designs, leading to high energy dissipation and cooling costs, with existing solutions being primarily reactive and not optimizing energy use across computing and cooling systems.
Innovation Solution
A holistic approach treating energy as a first-class resource, using predictive thermal models and global schedulers to allocate energy budgets to servers and match cooling efforts dynamically with workload and thermal conditions, integrating proactive and reactive control mechanisms to optimize energy efficiency and cost-effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If servers are designed for peak performance and overprovisioned, then computing capacity is improved, but energy efficiency deteriorates
Solution Approach 1:
The system dynamically adjusts server performance settings and workload allocation based on real-time thermal conditions and energy consumption patterns. The global scheduler continuously reconfigures task distribution across servers, and performance settings are adjusted on-the-fly rather than being fixed for peak capacity, resolving the contradiction between maintaining peak computing capacity and achieving energy efficiency.
Solution Approach 2:
The system changes operational parameters including server performance settings, workload allocation ratios, and cooling system parameters based on measured thermal conditions and energy consumption. By dynamically adjusting these parameters rather than operating at fixed peak settings, the system achieves both adequate computing capacity and improved energy efficiency.
2Productivity
If more servers are packed into data centers, then computing capacity is improved, but cooling energy consumption deteriorates
Solution Approach 1:
The system applies differentiated cooling strategies to different server racks and locations based on their specific thermal conditions and workload characteristics. Rather than uniform cooling, the system identifies hot spots and directs cooling resources locally where needed, reducing overall cooling energy consumption while maintaining adequate computing capacity across densely packed servers.
Solution Approach 2:
The system introduces thermal models and global schedulers as intermediary layers between the computing workload and physical cooling infrastructure. These intermediaries predict thermal conditions and optimize workload placement to minimize heat generation in critical areas, thereby reducing the cooling energy required to maintain computing capacity in densely packed data centers.
3Device complexity
If reactive cooling solutions are used, then thermal management is simplified, but energy efficiency deteriorates
Solution Approach 1:
The system uses thermal models to predict future thermal conditions and proactively adjusts cooling settings and workload allocation before overheating occurs. Rather than simply reacting to temperature sensor feedback, the predictive approach anticipates thermal trends and takes preventive action, improving energy efficiency by avoiding excessive cooling while maintaining thermal management without significantly increasing system complexity.
4Loss of energy
If servers operate below peak load, then energy consumption is reduced, but energy efficiency deteriorates
Solution Approach 1:
The system consolidates workloads to keep active servers operating continuously at or near peak load levels, maximizing their energy efficiency. Rather than allowing servers to idle at low utilization, the global scheduler continuously allocates tasks to maintain high utilization of active resources, ensuring that energy is consumed efficiently when computing work is performed while allowing idle servers to be powered down or cooled less aggressively.
Data Source
AI summary
A method for controlling a data center, comprising a plurality of server systems, each associated with a cooling system and a thermal constraint, comprising: a concurrent physical condition of a first server system; predicting a future physical condition based on a set of future states of the first server system; dynamically controlling the cooling system in response to at least the input and the predicted future physical condition, to selectively cool the first server system sufficient to meet the predetermined thermal constraint; and controlling an allocation of tasks between the plurality of server systems to selectively load the first server system within the predetermined thermal constraint and selectively idle a second server system, wherein the idle second server system can be recruited to accept tasks when allocated to it, and wherein the cooling system associated with the idle second server system is selectively operated in a low power consumption state.

