Storage Response Time Prediction via Workload Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting response times of storage systems are inefficient as they often lack accurate forecasting, leading to undersized infrastructure that may not meet required performance parameters, resulting in suboptimal storage system configurations.
Innovation Solution
A method involving the creation of training examples from telemetry data, which are then used in unsupervised and supervised learning processes to cluster and predict response times based on system characteristics and workload features, allowing for precise determination of necessary storage system configurations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional response time prediction methods are used, then the prediction process is simpler, but the prediction accuracy is insufficient leading to undersized infrastructure
Solution Approach 1:
The patent segments the prediction problem by creating multiple specialized learning processes, each trained on specific clusters of workload characteristics. This segmentation allows each model to focus on particular patterns, improving overall prediction accuracy while maintaining manageable complexity through modular architecture
Solution Approach 2:
The patent applies preliminary action by pre-processing telemetry data to extract relevant features and pre-clustering workload characteristics before prediction. This preparation work is done in advance during training, enabling the prediction system to make accurate predictions without complex real-time analysis
2Reliability
If accurate response time prediction is achieved through clustering and multiple learning processes, then infrastructure sizing is optimized, but the system complexity increases
Solution Approach 1:
The system divides the complex prediction task into multiple specialized learning processes, each handling specific workload clusters. This segmentation improves reliability by ensuring each model is expert in its domain, while the modular structure keeps overall system complexity manageable
Solution Approach 2:
The patent changes parameters by transforming raw telemetry data into engineered features and organizing them into discrete clusters. This parameter transformation simplifies the input space for each learning process, improving reliability without requiring excessively complex models
3Productivity
If traditional infrastructure sizing is used, then the deployment is faster and simpler, but the performance requirements are not met
Solution Approach 1:
The patent performs preliminary action by training the learning processes and clustering algorithms in advance using historical telemetry data. This upfront investment in model training enables rapid, accurate predictions during deployment, improving productivity without significant time loss during actual use
Solution Approach 2:
The system creates copies of learning processes for different workload clusters, allowing parallel prediction capabilities. This copying approach enables comprehensive performance analysis without sequential processing delays, improving productivity while the models are trained once and reused
Data Source
AI summary
Training examples are created from telemetry data, in which each training example engineered features derived from the telemetry data, storage system characteristics about the storage system that processed the workload associated with the telemetry data, and the response time of the storage system while processing the workload. The training examples are provided to an unsupervised learning process which assigns the training examples to clusters. Training examples of each cluster are used to train/test a separate supervised learning process for the cluster, to cause each supervised learning process to learn a regression between independent variables (system characteristics and workload features) and a dependent variable (storage system response time). To determine a response time of a proposed storage system, the proposed workload is used to select one of the clusters, and then the trained learning process for the selected cluster is used to determine the response time of the proposed storage system.


