IT Resource Tuning With Driver-Metric Threshold Forecasting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional resource tuning systems fail to effectively utilize correlations between performance metrics and driver metrics to predict and prevent system failures, often leading to untimely and ineffective resource allocation.
Innovation Solution
A system that identifies correlations between performance metrics and driver metrics, uses extrapolation algorithms to determine driver metric thresholds, and predicts potential system failures, enabling proactive resource tuning by allocating resources before threshold breaches.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If conventional resource tuning systems are used, then resource allocation is performed, but the allocation is untimely and ineffective because systems fail to utilize correlations between performance metrics and driver metrics
Solution Approach 1:
The system performs preliminary actions by identifying correlations between driver metrics and performance metrics in advance, training extrapolation algorithms on historical data, and establishing threshold relationships before actual system failures occur. This enables the system to predict when performance thresholds will be breached and allocate resources proactively rather than reactively, making resource allocation timely and effective
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring both driver metrics (external factors) and performance metrics (system responses), comparing actual performance against predicted thresholds, and using this feedback to refine extrapolation algorithms and improve future predictions. This closed-loop feedback ensures resource allocation decisions are based on accurate, continuously improved correlations
2Reliability
If resource allocation is delayed until system failures occur, then resources are allocated to address failures, but system downtime increases and performance deteriorates
Solution Approach 1:
The system allocates resources in advance by predicting future performance threshold breaches based on correlated driver metrics and trained extrapolation algorithms. When the system detects that a driver metric is approaching a threshold that would cause performance degradation, resources are allocated beforehand to prevent the failure, thereby maintaining system availability and avoiding downtime
Solution Approach 2:
The system takes preliminary anti-action by identifying and counteracting the effects of driver metrics before they cause performance degradation. By predicting the impact of external factors (such as traffic patterns, seasonal variations, or operational changes) on system performance, the system proactively adjusts resource allocation to prevent the negative effects from occurring, thus maintaining system availability
3Reliability
If the system monitors only performance metrics without considering external driver metrics, then monitoring is simple, but the system cannot predict failures caused by external factors
Solution Approach 1:
The system segments monitoring into two distinct but correlated components: driver metrics (external factors) and performance metrics (system responses). By separately tracking and analyzing each type of metric and then establishing correlations between them through trained extrapolation algorithms, the system gains failure prediction capability while keeping the monitoring structure organized and manageable rather than creating an undifferentiated complex monitoring system
Solution Approach 2:
The system introduces trained extrapolation algorithms as intermediaries that connect driver metrics to performance metrics. These algorithms serve as mediators that learn the relationships between external factors and system responses, enabling the system to predict failures caused by external factors without requiring direct monitoring of every possible failure mode, thus balancing prediction capability with manageable complexity
Data Source
AI summary
Described techniques determine performance metric values of a performance metric characterizing a performance of a system resource of an information technology (IT) system, and determine driver metric values of a driver metric characterizing an occurrence of an event that is at least partially external to the system resource. A correlation analysis may confirm a potential correlation between the performance metric values and the driver metric values as a correlation. A graph relating the performance metric to the driver metric may be generated. A plurality of extrapolation algorithms may be trained to obtain a plurality of trained extrapolation algorithms using a first subset of data points of the graph, and the plurality of trained extrapolation algorithms may be validated using a second subset of data points of the graph. A driver metric threshold corresponding to the performance metric threshold may be determined using a validated extrapolation algorithm.


