Data Center Resource Forecasting via ML Telemetry Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data centers face challenges in efficiently managing and forecasting resource utilization across tens of thousands of assets, given the variability in workloads and the lead times associated with asset deployment and replacement.
Innovation Solution
A method and system that involve receiving data center asset telemetry information, processing it to generate performance indicator metrics, and using these metrics to train a resource utilization machine learning model. This model then provides forecasts on data center asset resource utilization, enabling proactive management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If data center operators manually monitor and manage resource utilization across tens of thousands of assets, then they can maintain control over resource allocation, but the complexity and time required for accurate forecasting increases significantly
Solution Approach 1:
The patent introduces machine learning models as intermediary systems that process telemetry data and generate forecasting predictions. These models act as mediators between raw data collection and decision-making, automatically analyzing patterns and predicting future resource utilization without requiring manual intervention for each asset.
Solution Approach 2:
The system enables self-service forecasting by automatically collecting telemetry data, processing it through machine learning models, and generating resource utilization predictions without human intervention. The automated pipeline continuously monitors assets and provides forecasting results, freeing operators from manual monitoring tasks.
2Reliability
If data center operators increase the frequency of resource monitoring to improve forecasting accuracy, then they can better predict resource needs, but the computational resources and processing time required increase
Solution Approach 1:
The patent applies partial action by monitoring and processing only the most critical telemetry parameters and assets that have the greatest impact on forecasting accuracy. The machine learning models are trained to identify and focus on key performance indicators and assets that contribute most to resource utilization predictions, reducing unnecessary processing of less significant data.
Solution Approach 2:
The system performs preliminary action by pre-processing and filtering telemetry data before it reaches the machine learning models. Data is aggregated, cleaned, and transformed into relevant features in advance, reducing the computational burden during actual forecasting operations and enabling more frequent monitoring with lower energy consumption.
3Adaptability or versatility
If data center operators deploy more machine learning models to handle tens of thousands of assets, then forecasting coverage improves, but the device complexity and implementation difficulty increases
Solution Approach 1:
The patent implements universality by developing a single machine learning framework that can handle multiple asset types and forecasting scenarios. The system uses a unified approach to data collection, processing, and prediction that works across diverse data center assets, eliminating the need for separate specialized models for each asset category and reducing overall system complexity.
Solution Approach 2:
The system applies segmentation by dividing the large-scale forecasting problem into manageable components: data collection at the asset level, processing at the cluster level, and aggregation at the data center level. This hierarchical segmentation allows the system to handle tens of thousands of assets through modular processing stages, reducing implementation complexity while maintaining comprehensive coverage.
Data Source
AI summary
A system, method, and computer-readable medium for performing a data center monitoring and management operation. The data center monitoring and management operation includes: receiving data center asset telemetry information; processing the data center asset telemetry information to provide a performance indicator metric; training a resource utilization machine learning model using the performance indicator metric; and, using the resource utilization machine learning model to provide a data center asset resource utilization result.


