Storage Response Time Prediction via Workload Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for predicting response times of storage systems are inefficient as they often lack accurate forecasting, leading to undersized infrastructure that may not meet required performance parameters, resulting in suboptimal storage system configurations.

Innovation Solution

A method involving the creation of training examples from telemetry data, which are then used in unsupervised and supervised learning processes to cluster and predict response times based on system characteristics and workload features, allowing for precise determination of necessary storage system configurations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional response time prediction methods are used, then the prediction process is simpler, but the prediction accuracy is insufficient leading to undersized infrastructure

Engineering Contradiction:
Improveresponse time prediction accuracyVSAvoidprediction system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the prediction problem by creating multiple specialized learning processes, each trained on specific clusters of workload characteristics. This segmentation allows each model to focus on particular patterns, improving overall prediction accuracy while maintaining manageable complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-processing telemetry data to extract relevant features and pre-clustering workload characteristics before prediction. This preparation work is done in advance during training, enabling the prediction system to make accurate predictions without complex real-time analysis

Inventive Principle:
Principle #10Preliminary action

2Reliability

If accurate response time prediction is achieved through clustering and multiple learning processes, then infrastructure sizing is optimized, but the system complexity increases

Engineering Contradiction:
Improveinfrastructure performance guaranteeVSAvoidlearning process architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system divides the complex prediction task into multiple specialized learning processes, each handling specific workload clusters. This segmentation improves reliability by ensuring each model is expert in its domain, while the modular structure keeps overall system complexity manageable

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes parameters by transforming raw telemetry data into engineered features and organizing them into discrete clusters. This parameter transformation simplifies the input space for each learning process, improving reliability without requiring excessively complex models

Inventive Principle:
Principle #35Parameter changes

3Productivity

If traditional infrastructure sizing is used, then the deployment is faster and simpler, but the performance requirements are not met

Engineering Contradiction:
Improvestorage system response timeVSAvoidprediction and planning time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by training the learning processes and clustering algorithms in advance using historical telemetry data. This upfront investment in model training enables rapid, accurate predictions during deployment, improving productivity without significant time loss during actual use

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of learning processes for different workload clusters, allowing parallel prediction capabilities. This copying approach enables comprehensive performance analysis without sequential processing delays, improving productivity while the models are trained once and reused

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11915153B2Workload-oriented prediction of response times of storage systems
Publication Date: 2024.02.27 EMC IP HLDG CO LLC
  • US11915153B2 patent drawing
  • US11915153B2 patent drawing
  • US11915153B2 patent drawing

AI summary

Training examples are created from telemetry data, in which each training example engineered features derived from the telemetry data, storage system characteristics about the storage system that processed the workload associated with the telemetry data, and the response time of the storage system while processing the workload. The training examples are provided to an unsupervised learning process which assigns the training examples to clusters. Training examples of each cluster are used to train/test a separate supervised learning process for the cluster, to cause each supervised learning process to learn a regression between independent variables (system characteristics and workload features) and a dependent variable (storage system response time). To determine a response time of a proposed storage system, the proposed workload is used to select one of the clusters, and then the trained learning process for the selected cluster is used to determine the response time of the proposed storage system.