Datacenter Utilization Prediction via Hierarchical ML Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Datacenters face challenges in accurately predicting hardware utilization, especially at higher levels, due to inaccuracies in non-operating system-generated utilization data, which affects resource allocation, capacity planning, and security operations, leading to over-provisioning and sub-optimal client relationships.
Innovation Solution
A hierarchy of machine learning models is trained to predict hardware utilization at various levels of a datacenter, using non-OS sources like out-of-band sensors and network traffic analysis, to provide accurate and comprehensive forecasts without diverting processing power from client use, preserving 'bare-metal' guarantees.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If operating system tools are used to acquire server utilization information, then utilization data can be obtained, but server security is reduced and server device performance is negatively affected
Solution Approach 1:
The patent introduces an intermediary system that collects hardware utilization information directly from hardware components (CPU, memory, storage, network interfaces) without requiring operating system involvement. This intermediary layer acts as a mediator between the monitoring system and hardware resources, obtaining utilization data through hardware-level sensors and counters while preserving operating system integrity and security.
2Loss of information
If operating system tools are used to acquire server utilization information, then utilization data can be obtained, but server device performance is negatively affected
Solution Approach 1:
The monitoring system uses an intermediary approach that collects utilization data directly from hardware components through dedicated sensors and hardware-level interfaces. This eliminates the need for operating system to execute monitoring instructions, thereby preventing any CPU cycle consumption or bandwidth usage that would otherwise impact server performance and client workloads.
3Loss of information
If utilization monitoring is performed on bare-metal servers, then hardware utilization can be tracked, but access to operating system-generated utilization information is prohibited
Solution Approach 1:
For bare-metal servers, the patent implements an intermediary monitoring system that operates independently of any operating system. The system directly interfaces with hardware components through hardware-level sensors and management interfaces, collecting utilization information without requiring or accessing an operating system environment. This approach maintains full client control over the server while enabling the provider to monitor hardware utilization.
Solution Approach 2:
The patent replaces the traditional software-based utilization monitoring mechanism (which requires operating system involvement) with a hardware-based monitoring system. Instead of using software agents or operating system tools to collect utilization data, the system uses hardware-level sensors, counters, and management interfaces to directly measure CPU utilization, memory usage, storage I/O, and network traffic, thereby eliminating the need for operating system access.
4Reliability
If datacenter resources are over-provisioned to prevent high utilization spikes, then service reliability is improved, but cost increases
Solution Approach 1:
The patent implements preliminary action by using machine learning models to predict future hardware utilization trends and potential high-utilization events before they occur. The system analyzes historical utilization patterns, workload characteristics, and external factors to forecast future resource demands, enabling the datacenter to proactively allocate resources based on predicted needs rather than relying on conservative over-provisioning strategies.
Solution Approach 2:
The system employs continuous feedback loops where actual utilization data is constantly collected from hardware components, compared against predictions from machine learning models, and used to refine future predictions. This feedback mechanism enables dynamic resource allocation adjustments, allowing the datacenter to optimize resource capacity in real-time based on actual utilization patterns rather than static over-provisioning decisions.
Data Source
AI summary
Embodiments use a hierarchy of machine learning models to predict datacenter behavior at multiple hardware levels of a datacenter without accessing operating system generated hardware utilization information. The accuracy of higher-level models in the hierarchy of models is increased by including, as input to the higher-level models, hardware utilization predictions from lower-level models. The hierarchy of models includes: server utilization models and workload/OS prediction models that produce predictions at a server device-level of a datacenter; and also top-of-rack switch models and backbone switch models that produce predictions at higher levels of the datacenter. These models receive, as input, hardware utilization information from non-OS sources. Based on datacenter-level network utilization predictions from the hierarchy of models, the datacenter automatically configures its hardware to avoid any predicted over-utilization of hardware in the datacenter. Also, the predictions from the hierarchy of models can be used to detect anomalies of datacenter hardware behavior.


