Proactive Storage Array Issue Prevention via Cloud Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Managing disparate issues across varying storage arrays in data centers is challenging, as existing corrective measures often impact performance and may not prevent problems proactively.
Innovation Solution
A system that receives data from storage arrays, detects problem signatures indicative of specific issues, and automatically deploys corrective measures without user intervention if the issue violates operational policies, using a cloud-based storage array services provider to preempt and correct issues by upgrading software, firmware, or adjusting performance parameters.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If corrective measures are deployed to prevent storage array problems, then reliability is improved, but performance deteriorates during deployment
Solution Approach 1:
The system performs preliminary actions by detecting problem signatures and deploying corrective measures before actual failures occur. The monitoring system continuously analyzes event data to identify patterns indicative of impending failures, allowing preventive maintenance to be performed proactively rather than reactively, thus improving reliability without requiring performance-degrading corrective actions during operation
Solution Approach 2:
The system implements feedback mechanisms by continuously monitoring storage array events, comparing them against known problem signatures, and adjusting corrective measures based on the detected patterns. This closed-loop approach allows the system to learn from actual failure data and refine its predictive capabilities, enabling more accurate and less intrusive preventive actions
2Reliability
If monitoring and detection systems are enhanced to identify problems earlier, then reliability is improved, but device complexity increases
Solution Approach 1:
The monitoring system achieves universality by using a single event collection framework that handles multiple storage array types and failure modes through unified problem signature patterns. Rather than implementing separate monitoring mechanisms for each storage array type, the system uses generic event parsing and pattern matching capabilities that work across diverse storage configurations, reducing overall system complexity while maintaining comprehensive monitoring
Solution Approach 2:
The system changes parameters by transforming raw event data into normalized problem signatures that capture essential failure patterns independent of specific storage array implementations. By abstracting failure conditions into standardized signature patterns rather than maintaining detailed type-specific monitoring logic, the system reduces complexity while improving detection accuracy across heterogeneous storage environments
3Productivity
If automatic corrective measures are deployed without user intervention, then productivity is improved, but ease of operation deteriorates
Solution Approach 1:
The system implements self-service by automatically detecting problem signatures and deploying corrective measures without requiring user intervention. The monitoring system autonomously analyzes event data, identifies failures based on predefined signatures, and executes appropriate corrective actions, allowing the storage array system to self-diagnose and self-heal, thus improving availability while reducing operational burden on users
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Proactively providing corrective measures for storage arrays includes: receiving data from a storage array, the data including one or more events; detecting, in dependence upon a problem signature, one or more events from the data indicative of a particular problem, where the problem signature comprises a specification of a pattern of events indicative of the particular problem experienced by at least one other storage array; determining whether the particular problem violates an operational policy of the storage array, the operational policy specifying at least one requirement for an operational metric of the storage array; and if the particular problem violates the operational policy of the storage array, deploying automatically without user intervention one or more corrective measures to prevent the storage array from experiencing the particular problem.