Workload Skew Determination Using Exponential Curve Fitting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face challenges in efficiently managing workload skew across different storage tiers, leading to suboptimal performance and capacity utilization, as they lack effective methods to dynamically assess and adjust workload distribution based on performance rankings.
Innovation Solution
A method is introduced to determine workload skew by measuring and analyzing workload and capacity data across multiple storage tiers, using exponential curve fitting to model performance and adjust data movements between tiers, thereby optimizing capacity planning and response times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is distributed across multiple storage tiers without workload skew analysis, then storage capacity is utilized, but system performance deteriorates due to inefficient data placement
Solution Approach 1:
The system continuously monitors workload measurements and capacity measurements across storage tiers, uses exponential curve fitting to model workload skew, and automatically generates data movement recommendations. This closed-loop feedback mechanism enables the system to adapt to changing workload patterns and optimize data placement dynamically, resolving the contradiction between improving performance and managing complexity.
Solution Approach 2:
The system changes the parameter of data placement by analyzing workload skew through exponential curve fitting and generating specific data movement recommendations. By dynamically adjusting which data portions are moved between storage tiers based on modeled workload patterns, the system optimizes performance without requiring manual configuration management.
2Loss of time
If manual data movement between storage tiers is performed, then capacity planning is simplified, but response time increases due to lack of dynamic adjustment
Solution Approach 1:
The system performs self-service by automatically measuring workload, modeling workload skew using exponential curve fitting, and generating data movement recommendations without requiring manual intervention. This automation eliminates the trade-off between manual control and dynamic response, as the system autonomously optimizes data placement based on real-time workload patterns.
Solution Approach 2:
The system performs preliminary action by proactively modeling workload skew and generating data movement recommendations before performance degradation occurs. By continuously analyzing workload patterns and predicting optimal data placement, the system prevents response time increases rather than reacting to them.
3Adaptability or versatility
If storage tiers are configured without workload-based optimization, then system complexity is reduced, but capacity utilization deteriorates
Solution Approach 1:
The system segments the storage system into multiple storage tiers with different performance characteristics and applies workload-based optimization to each tier independently. By dividing the storage system into manageable segments and analyzing workload skew for each, the system achieves high adaptability without overwhelming complexity, as each tier can be optimized separately based on its specific workload patterns.
Data Source
AI summary
Determining cumulative workload skew is described. Measurements for one or more logical devices are determined. The set of measurements include, for each of N storage tiers, a workload measurement identifying workload directed to the single tier, and a capacity measurement identifying an amount of data stored in the single tier. N points may be determined using the measurements. Each point corresponds to a different storage tier and has a first coordinate identifying a cumulative percentage of data portions stored in the storage tier and all other tiers having a higher performance ranking than the one storage tier, and a second coordinate denoting an aggregated percentage of total workload directed to the foregoing cumulative percentage of data portions. A curve representing a cumulative workload skew may be determined using these N points and a point of origin.


