Data Center Power Sizing Using Statistical Multiplexing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face significant challenges in efficiently utilizing their power budgets due to underutilization, conservative equipment ratings, variable load, and statistical effects, leading to increased costs and inefficiencies in powering and cooling massive computing systems.
Innovation Solution
A method for designing and managing data centers involves determining a design power density, calculating an oversubscription ratio, and optimizing the spatial layout to maximize power utilization, including monitoring and adjusting CPU utilization, job scheduling, and implementing power management techniques like CPU voltage scaling to reduce peak power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the power capacity is sized to meet the maximum power draw (sum of peak power draws of all computers), then reliability is improved, but the power capacity is significantly over-provisioned leading to underutilization and increased costs
Solution Approach 1:
The patent changes the parameter from using peak power draw to using expected power draw, which is a statistically derived value representing typical operating conditions. This parameter change allows the power capacity to be sized appropriately for actual usage patterns rather than worst-case scenarios, resolving the contradiction between reliability and utilization.
Solution Approach 2:
The system uses monitoring and measurement of actual power consumption patterns to self-adjust the sizing criteria. By continuously measuring power draws and calculating expected values based on observed behavior, the system automatically determines appropriate power capacity without external intervention, balancing reliability with efficient utilization.
2Reliability
If conservative equipment ratings are used for power capacity planning, then reliability is improved, but power utilization efficiency deteriorates due to excessive capacity provisioning
Solution Approach 1:
The patent implements feedback loops where power consumption data is continuously monitored, measured, and fed back into the sizing calculation. This feedback mechanism allows the system to learn from actual operating conditions and adjust the expected power draw calculations, replacing conservative estimates with data-driven values that maintain reliability while improving productivity.
Solution Approach 2:
The system performs preliminary measurements and monitoring during a setup phase to establish baseline power consumption patterns before finalizing the power capacity sizing. This preliminary action allows the system to gather real-world data and calculate expected power draws before committing to a fixed power capacity configuration.
3Reliability
If peak power draw is used for each computer, then individual computer reliability is ensured, but aggregate power utilization is poor due to statistical multiplexing effects
Solution Approach 1:
The patent merges individual computer power draw measurements into an aggregate expected power draw calculation for the entire data center. By combining multiple individual measurements and applying statistical analysis, the system captures the multiplexing effect where not all computers peak simultaneously, thus reducing wasted power capacity while maintaining individual computer reliability.
Solution Approach 2:
The patent transforms the parameter from individual peak power draw to aggregate expected power draw, which incorporates statistical relationships between multiple computers. This parameter change accounts for the fact that peak loads do not occur simultaneously across all systems, reducing energy waste while ensuring reliability.
4Loss of energy
If power capacity is reduced below maximum power draw, then cost efficiency is improved, but the risk of exceeding power capacity increases
Solution Approach 1:
The system uses self-monitoring and statistical analysis of actual power consumption patterns to determine safe power capacity levels. By continuously measuring and learning from operational data, the system automatically identifies the expected power draw that maintains adequacy while optimizing cost efficiency, without requiring external validation or conservative margins.
Solution Approach 2:
The patent implements feedback mechanisms where power consumption data is continuously collected and used to adjust the expected power draw calculations. This feedback ensures that the reduced power capacity remains adequate by comparing actual usage against predicted usage and adjusting accordingly, maintaining reliability while improving cost efficiency.
Data Source
AI summary
A method for use in deploying computers into a data center includes calculating in a computer an expected peak power draw for a plurality of computers. The expected peak power draw for the plurality of computers is less than a sum of individual expected peak power draws for each computer from the plurality of computers.


