Storage Data Reduction Management via ML Workload Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In storage systems, resource contention among workloads leads to performance issues due to ineffective data reduction applications, making it difficult for administrators to determine which workloads are consuming resources inefficiently, resulting in high latency and wasteful resource usage.
Innovation Solution
A data reduction management engine utilizing machine learning models to identify high latency and recommend whether data reduction should be applied to specific storage volumes, based on performance metrics and resource utilization, converting inefficiently utilizing storage volumes from 'thick' to 'non-thick' to optimize resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data reduction is applied to all storage volumes, then storage space savings are improved, but resource contention and latency increase due to ineffective applications
Solution Approach 1:
The system changes the parameter of data reduction application from a universal setting to a selective setting based on workload characteristics. Machine learning models analyze workload parameters (I/O patterns, data types, access frequencies) and dynamically adjust whether data reduction should be applied, transforming the system from static to adaptive parameter configuration.
Solution Approach 2:
The system enables workloads to self-determine their data reduction needs through machine learning analysis. Each workload is evaluated independently by the ML models which predict the effectiveness of data reduction for that specific workload, allowing the system to self-optimize without manual administrator intervention for each case.
2Productivity
If data reduction is applied to storage volumes, then storage efficiency is improved, but resource contention increases making it difficult to identify inefficient workloads
Solution Approach 1:
The system implements feedback loops where machine learning models continuously monitor workload performance and resource usage. The models receive feedback about actual data reduction effectiveness and resource consumption patterns, then use this feedback to refine predictions and identify which workloads are causing inefficient resource contention, making detection straightforward through the ML analysis.
Solution Approach 2:
The machine learning models act as intermediaries between the workloads and the storage system resources. Instead of administrators directly monitoring complex resource contention among multiple workloads, the ML models mediate by analyzing and interpreting the interactions, identifying which specific workloads are causing inefficiencies, and providing clear attribution of resource contention sources.
3Quantity of substance
If data reduction is applied without selective analysis, then storage space is saved, but processing resources are wasted on ineffective reductions
Solution Approach 1:
The system applies partial action by selectively enabling data reduction only for workloads where it will be effective, rather than applying it universally. The machine learning models identify the subset of workloads that benefit from data reduction and apply the technique only to those cases, avoiding the excessive action of applying data reduction to all workloads regardless of effectiveness.
Solution Approach 2:
The system performs preliminary analysis using machine learning models to predict which workloads will benefit from data reduction before actually applying the data reduction technique. This preliminary evaluation prevents wasted processing resources by identifying ineffective candidates in advance, ensuring that data reduction processing is only applied where it will produce actual storage savings.
Data Source
AI summary
In some examples, a system identifies resource contention for a resource in a storage system, and determines that a workload collection that employs data reduction is consuming the resource. The system identifies relative contributions to consumption of the resource attributable to storage volumes in the storage system, where the workload collection that employs data reduction are performed on data of the storage volumes. The system determines whether storage space savings due to application of the data reduction for a given storage volume of the storage volumes satisfy a criterion, and in response to determining that the storage space savings for the given storage volume do not satisfy the criterion, the system indicates that the data reduction is not to be applied for the given storage volume.


