Data Center Control Hierarchy for Neural Network Power Flow
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face challenges in accommodating workload variations and energy efficiency, with current control designs being overly complex and lacking scalability and technology reusability, particularly due to the integration of different control systems for cooling, power, and IT systems, and the need for AI/ML models that require extensive training and tuning.
Innovation Solution
A control hierarchy design for data centers with three levels (load, source, and intermediate) that includes a power flow optimizer using neural networks to predict power requirements based on thermal and load data, and a resource optimizer to configure power sources, enabling decoupled control while maintaining system integration, and utilizing AI/ML models for optimized computing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If separate control modules are used for cooling systems, power systems, and IT systems, then each module can be designed and operated independently, but the overall system integration becomes extremely complicated and lacks scalability
Solution Approach 1:
The control system is divided into three hierarchical levels: load-level controllers (managing individual racks), intermediate-level controllers (managing zones or floors), and source-level controllers (managing power sources). Each level operates semi-independently with standardized communication interfaces, enabling modular design while achieving systematic integration through the hierarchical structure.
Solution Approach 2:
The patent implements a universal control framework where controllers at different levels can manage multiple types of equipment (IT loads, cooling systems, power sources) through standardized protocols. This multi-functional approach allows the same control architecture to handle diverse systems without requiring separate specialized integration for each module type.
2Measurement precision
If AI/ML models are trained extensively for one data center cluster, then the model performance is optimized for that specific cluster, but the model requires significant retraining when deployed to different clusters
Solution Approach 1:
The system performs preliminary actions by collecting and preprocessing data from multiple clusters during the training phase, creating a more generalized model that anticipates variations across different data center environments. This preliminary data aggregation and model training on diverse datasets reduces the need for extensive retraining when deploying to new clusters.
Solution Approach 2:
The patent employs transfer learning and fine-tuning techniques where the core model structure and parameters are trained on one cluster, then adapted to other clusters by adjusting only specific parameters rather than retraining the entire model. This allows the model to maintain high accuracy across different clusters with minimal retraining effort.
3Reliability
If control systems are tightly integrated to achieve organic operation, then system coordination is improved, but the control design becomes overly complicated
Solution Approach 1:
The control system is divided into three hierarchical levels: load-level controllers (managing individual racks), intermediate-level controllers (managing zones or floors), and source-level controllers (managing power sources). Each level operates semi-independently with standardized communication interfaces, enabling modular design while achieving systematic integration through the hierarchical structure.
Solution Approach 2:
The intermediate-level controller acts as a mediator between load-level and source-level controllers, coordinating information flow and control decisions across the system. This intermediary layer simplifies the overall design by providing a structured communication bridge that reduces direct complexity between higher and lower levels.
4Productivity
If data centers deploy more servers to accommodate increasing workload requirements, then processing capacity is improved, but energy consumption and operational costs increase
Solution Approach 1:
The control system continuously monitors power consumption, thermal conditions, and workload levels across all racks and dynamically adjusts power distribution and cooling based on real-time feedback. This feedback mechanism enables the system to optimize energy usage by matching power delivery and cooling capacity to actual workload demands, preventing energy waste while maintaining processing capacity.
Solution Approach 2:
The patent implements dynamic power management where the system continuously adapts power allocation to workload demands. Controllers adjust power distribution in real-time based on changing conditions, enabling the data center to maintain high processing capacity when needed while reducing power consumption during lower-demand periods, thus resolving the contradiction between productivity and energy use.
Data Source
AI summary
A data center system includes a load section having an array of electronic racks, a thermal management system, and a power flow optimizer. The power flow optimizer is configured to determine a load power requirement of the load section based on workload data of the electronic racks and thermal data of the thermal management system. The data center system further includes a resource section having a number of power sources to provide power to the load section. The resource section includes a resource controller to configure and select at least some of the power sources to provide power to the load section based on the load power requirement provided by the power flow optimizer. The power flow optimizer includes a power flow neural network (NN) model to predict, based on the thermal data and the load data, an amount of power that IT clusters and the thermal management system need.


