Workload Skew Prediction Model for Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges in predicting and managing workload skew across different storage tiers, leading to inefficient data movement and capacity planning, as existing methods lack accurate models to anticipate and adapt to dynamically changing workloads.
Innovation Solution
A method is developed to predict cumulative skew curves using machine learning regression algorithms, which generate destination cumulative skew curves based on observed data and source skew curves, allowing for optimized data movement and capacity planning across storage tiers by identifying features such as total area, read/write activity, and time-based characteristics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If traditional data storage systems use fixed data movement granularity, then system simplicity is maintained, but workload skew prediction accuracy deteriorates
Solution Approach 1:
The patent applies dynamics by making the data movement granularity adaptive rather than fixed. The system dynamically adjusts the granularity of data movement operations based on observed workload patterns and skew characteristics. This allows the storage system to optimize data migration efficiency for different workload types while maintaining manageable system complexity through automated adaptation.
Solution Approach 2:
The patent changes the parameter of data movement granularity from a static value to a dynamically adjustable parameter. By modifying this parameter based on workload conditions, the system achieves better workload skew prediction accuracy without permanently increasing complexity, as the changes are driven by automated monitoring and adaptation mechanisms.
2Productivity
If manual capacity planning is used, then planning flexibility is maintained, but time consumption and efficiency worsen
Solution Approach 1:
The patent implements self-service by enabling the storage system to automatically perform capacity planning based on monitored workload patterns. The system autonomously analyzes skew curves, predicts future workload distributions, and generates capacity planning recommendations without requiring manual intervention. This dramatically reduces both the time required and increases the efficiency of capacity planning while maintaining flexibility through adaptive algorithms.
Solution Approach 2:
The patent uses feedback mechanisms where the system continuously monitors actual workload patterns and compares them against predictions. This feedback loop allows the capacity planning process to automatically adjust based on real-world performance data, improving accuracy over time and reducing the need for manual re-planning, thereby increasing efficiency and reducing time loss.
3Productivity
If data storage systems lack workload prediction models, then system simplicity is maintained, but data movement optimization deteriorates
Solution Approach 1:
The patent applies preliminary action by implementing workload prediction models that forecast future skew patterns before actual data movement is needed. By predicting workload distributions in advance, the system can proactively optimize data movement operations, pre-position data on appropriate storage tiers, and prevent performance degradation before it occurs, thereby improving productivity without excessive complexity.
Solution Approach 2:
The patent replaces manual or rule-based data movement mechanisms with intelligent prediction models. Instead of relying on simple thresholds or fixed policies, the system uses learned patterns from historical workload data to make sophisticated predictions about future skew, enabling more effective optimization while managing complexity through automated machine learning approaches.
Data Source
AI summary
Described are techniques that determine cumulative skew curves. A first model is determined that generates a predicted destination cumulative skew curve for a specified data set in a destination data storage system having a destination data movement granularity. The predicted destination cumulative skew curve is predicted by the first model in accordance with one or more inputs including a source cumulative skew curve for the specified data set in a source data storage system that uses a source data movement granularity. The source cumulative skew curve for the specified data set is determined based on observed data. First processing is performed using the first model. The first model generates as an output the predicted destination cumulative skew curve. The first processing includes providing the one or more inputs to the first model. Also described is how to generate the first model.


