ML-Based Storage Event Detection for Workload Performance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Information processing systems face challenges in efficiently identifying and addressing performance-impacting events affecting workloads on storage systems, often resulting in false positives and difficulties in balancing unique storage resource demands across multiple workloads.
Innovation Solution
A method utilizing machine learning algorithms, specifically a neural network, to monitor performance data, detect potential performance-impacting events, classify them as true positives or false positives, and adjust storage resource provisioning accordingly, thereby improving workload performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional monitoring methods are used to detect performance-impacting events, then event detection capability is provided, but false positives increase and accuracy decreases
Solution Approach 1:
A machine learning classifier acts as an intermediary between performance data collection and event detection. The classifier processes raw performance metrics and visualizations, learning to distinguish true performance-impacting events from false positives caused by normal workload variations. This intermediary layer filters out false positives while preserving true event detection capability.
Solution Approach 2:
The system implements feedback mechanisms where performance data is continuously monitored, classified, and used to refine future detections. The machine learning model learns from historical performance patterns and feedback loops, improving its ability to accurately identify true performance-impacting events while reducing false positives over time.
2Reliability
If storage resources are increased to handle all workloads, then workload performance is maintained, but resource allocation efficiency decreases
Solution Approach 1:
The system dynamically adjusts storage resource provisioning based on real-time performance event detection. Instead of static over-provisioning, the machine learning-based detection system enables dynamic resource allocation that responds to actual performance needs, maintaining workload performance while optimizing resource utilization efficiency.
Solution Approach 2:
The system changes provisioning parameters based on detected performance events. When true performance-impacting events are identified, the system adjusts storage resource parameters accordingly, rather than maintaining fixed provisioning levels. This enables efficient resource allocation that adapts to actual workload performance requirements.
Data Source
AI summary
A method includes monitoring a given workload running on one or more storage systems to obtain performance data, detecting a given potential performance-impacting event affecting the given workload based at least in part on a given portion of the obtained performance data, and generating a visualization of at least the given portion of the obtained performance data. The method also includes providing the generated visualization as input to a machine learning algorithm, utilizing the machine learning algorithm to classify the given potential performance-impacting event as one of (i) a true positive event affecting performance of the given workload and (ii) a false positive event corresponding to one or more changes in storage resource utilization by the given workload, and modifying provisioning of storage resources of the one or more storage systems responsive to classifying the given potential performance-impacting event as a true positive event affecting performance of the given workload.


