Automated Service Time Estimation via Clustering Regression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Estimating service time in IT systems is challenging due to the lack of available measurements and the need for invasive techniques, and existing methods require manual intervention and are prone to errors, especially when dealing with anomalous behavior and multiple system configurations.
Innovation Solution
A method combining density-based clustering, clusterwise regression, and a refinement procedure, using structural regression models and orthogonal regression to accurately estimate service time, reducing the need for manual intervention and improving accuracy by accounting for errors in both workload and utilization measurements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If service time measurements are obtained through invasive techniques such as benchmarking, load testing, profiling, application instrumentation or kernel instrumentation, then measurement precision is improved, but device complexity and ease of operation worsen
Solution Approach 1:
The patent uses aggregate measurements (workload and utilization) as intermediaries to indirectly estimate service time without requiring direct invasive measurements. These intermediaries are easily obtainable from system monitors and serve as proxies that capture the necessary information while avoiding the complexity of instrumentation.
Solution Approach 2:
The patent replaces the mechanical/invasive measurement system (instrumentation, benchmarking) with a computational/statistical system that uses readily available aggregate data and regression analysis to estimate service time, thereby eliminating the need for complex physical or software instrumentation.
2Ease of operation
If a single regression model is used to estimate service time from workload and utilization, then ease of operation is improved, but measurement precision worsens due to anomalous behavior and multiple system configurations
Solution Approach 1:
The patent segments the system's operational space into multiple working zones or clusters, each with its own regression model. This segmentation allows the system to capture different behaviors under various configurations and workloads, improving accuracy while maintaining automated operation through cluster-based model selection.
Solution Approach 2:
The patent implements a dynamic regression modeling approach where the system automatically adapts between different regression models based on the current operating conditions. Instead of a static single model, the system dynamically selects or combines models based on workload characteristics and system state, improving precision without requiring manual intervention.
3Measurement precision
If manual intervention is used to detect regression models and classify observation samples, then measurement precision is improved, but productivity and ease of operation worsen
Solution Approach 1:
The patent implements self-service through automated clustering algorithms that automatically detect working zones, select appropriate regression models, and classify observation samples without human intervention. The system performs capacity planning and service time estimation autonomously, maintaining high accuracy while dramatically improving productivity and eliminating manual effort.
Solution Approach 2:
The patent uses feedback mechanisms where the system continuously monitors aggregate measurements, compares observed behavior against multiple regression models, and automatically adjusts model selection and parameters. This closed-loop feedback enables automated detection and classification with precision comparable to or exceeding manual methods.
4Ease of operation
If aggregate measurements are used instead of service time measurements, then ease of operation is improved, but measurement precision worsens
Solution Approach 1:
The patent transforms the estimation problem by changing the parameters used for calculation. Instead of directly measuring service time, the system uses readily available aggregate parameters (workload and utilization) and applies regression analysis to derive service time estimates. This parameter transformation maintains ease of operation while achieving acceptable precision through statistical relationships.
Data Source
AI summary
Embodiments provide a method for upgrading resources in a system including normalizing a collected dataset, scattering data from the normalized dataset, obtaining a plurality of clusters based on the scattered data, discarding one or more clusters from the plurality of clusters with less than a percentage of a total number of observations, in each cluster, performing clusterwise regression and obtaining linear sub-clusters in a defined number, reducing one or more sub-clusters including applying a refinement procedure, removing one or more sub-clusters that fit to outliers and merging pairs of clusters that fit an equivalent model, updating one or more clusters with the reduced sub-clusters, removing one or more globular clusters, reducing a number of clusters with the refinement procedure, and de-normalizing one or more results.


