Data Center Load Distribution With Predictive Thermal Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face a significant energy crisis due to inefficient energy management and cooling systems, leading to high utility costs and environmental impact.
Innovation Solution
A holistic approach is taken to manage data centers as cyberphysical systems, where cooling solutions are coordinated with computing level solutions for energy management. This involves using fast thermal models, global schedulers, and modified server operating system kernels to allocate energy budgets and optimize cooling efforts dynamically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If servers are overprovisioned with high capacity power supplies and designed for peak performance, then peak capacity and performance are improved, but energy efficiency deteriorates because servers operate far from peak loading levels most of the time
Solution Approach 1:
The patent implements dynamic workload scheduling that adapts to real-time thermal conditions and energy states. The system continuously monitors server thermal states and dynamically redistributes workloads between servers, adjusting power-performance settings based on current operating conditions rather than static peak provisioning. This allows the system to maintain peak capacity when needed while operating at efficient points during normal conditions.
Solution Approach 2:
The system changes operational parameters by dynamically adjusting power-performance settings of server components (CPU frequency, voltage, disk spin speeds) based on actual workload demands and thermal conditions. Instead of maintaining fixed peak-capacity parameters, the system modulates these parameters to match real-time requirements, improving energy efficiency while preserving peak capacity availability when truly needed.
2Productivity
If more servers are packed into a given physical space using smaller form factors, then processing capacity is improved, but energy dissipation and cooling requirements worsen
Solution Approach 1:
The patent applies local quality by implementing rack-level and server-level thermal management rather than uniform data center-wide cooling. The system monitors thermal conditions locally at each rack and server, then applies targeted cooling strategies and workload redistribution specific to hot spots. This allows high-density packing while managing energy dissipation locally where it occurs most intensely.
Solution Approach 2:
The system segments the data center into manageable units (racks, server groups) with independent thermal monitoring and control. By dividing the overall system into smaller controllable segments, the patent can optimize energy dissipation and cooling for each segment based on its specific thermal characteristics and workload patterns, rather than treating the entire data center as a single homogeneous unit.
3Ease of operation
If reactive energy management solutions are used that adjust cooling and power settings based on sensor feedback, then energy management responsiveness is improved, but thermal lag causes delayed response to sudden workload changes
Solution Approach 1:
The patent implements preliminary action by proactively anticipating thermal conditions and workload patterns before they fully manifest. The system uses predictive models and historical data to forecast thermal states and proactively adjusts cooling and workload distribution in advance, rather than merely reacting to current sensor readings. This reduces thermal lag by acting before the full thermal impact occurs.
Solution Approach 2:
The system employs multi-layered feedback mechanisms that combine real-time sensor data with predictive modeling. Rather than relying solely on reactive feedback from temperature sensors, the patent integrates predictive feedback from workload analysis and thermal models to anticipate future states. This hybrid feedback approach maintains responsiveness while compensating for thermal lag through predictive adjustments.
Data Source
AI summary
A method for controlling a data center, comprising a plurality of server systems, each associated with a cooling system and a thermal constraint, comprising: a concurrent physical condition of a first server system; predicting a future physical condition based on a set of future states of the first server system; dynamically controlling the cooling system in response to at least the input and the predicted future physical condition, to selectively cool the first server system sufficient to meet the predetermined thermal constraint; and controlling an allocation of tasks between the plurality of server systems to selectively load the first server system within the predetermined thermal constraint and selectively idle a second server system, wherein the idle second server system can be recruited to accept tasks when allocated to it, and wherein the cooling system associated with the idle second server system is selectively operated in a low power consumption state.

