Storage QOS Monitoring with Bully Volume Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems face challenges in efficiently monitoring and analyzing Quality of Service (QOS) performance, particularly in identifying and addressing incidents related to resource usage and response times across multiple storage volumes, which can lead to suboptimal performance and service level violations.
Innovation Solution
A method and system that collect QOS data from storage volumes, generate expected ranges for future performance, filter potential victim and bully storage volumes based on response time deviations and IOPS, and provide remediation plans to address resource contention and performance issues.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If continuous monitoring of QOS data for multiple storage volumes is implemented, then performance incidents can be identified timely, but system complexity and computational overhead increase
Solution Approach 1:
The monitoring system segments the storage environment into multiple storage volumes and further divides monitoring into phases: data collection phase, analysis phase, and incident identification phase. Each storage volume is monitored independently with its own QOS parameters, allowing parallel processing and reducing overall system complexity while maintaining comprehensive coverage
Solution Approach 2:
The system performs preliminary actions by collecting QOS data continuously and generating expected performance ranges before incidents occur. Historical QOS data is analyzed in advance to establish baseline performance characteristics, enabling proactive incident detection rather than reactive response, thus improving reliability without proportionally increasing complexity
2Measurement precision
If detailed analysis of response time and IOPS for each storage volume is performed, then accurate incident identification is achieved, but processing time and computational resources increase
Solution Approach 1:
The system applies partial action by focusing analysis only on storage volumes that deviate from their expected QOS ranges. Instead of analyzing all volumes equally, it identifies victims (volumes with abnormal performance) and bullies (volumes causing resource contention) selectively, reducing processing time while maintaining measurement precision for critical cases
Solution Approach 2:
Different levels of analysis depth are applied to different storage volumes based on their current performance state. Volumes within expected ranges receive minimal monitoring, while those showing deviations undergo detailed analysis of response time, IOPS, and resource usage patterns, optimizing the balance between precision and processing time
3Productivity
If resource contention and bully storage volumes are identified, then performance optimization is enabled, but additional monitoring and analysis overhead is created
Solution Approach 1:
The system implements feedback mechanisms where QOS monitoring data flows back to identify victims and bullies, which then triggers targeted analysis. This feedback loop enables automatic detection of resource contention without continuous full-system analysis, optimizing productivity while managing computational energy through event-driven processing rather than constant monitoring
Data Source
AI summary
Methods and systems for monitoring quality of service (QOS) data for a plurality of storage volumes are provided. QOS data is collected for the plurality of storage volumes and includes a response time in which each of the plurality of storage volumes respond to an input/output (I/O) request. An expected range for future QOS data based on the collected QOS data is generated. The process then determines a deviation of each potential bully storage volume of a resource used by any victim storage volume, where the deviation of each bully storage volume is based on a number of current I/O requests (IOPS) that are processed by each potential bully storage volume, a forecasted value of TOPS and a predicted upper threshold TOPS value for each potential bully storage volume; and filters the potential bully storage volumes based on an impact of each potential bully storage volume.


