Availability Score Calculation for Storage System Failure Management
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current storage systems lack an effective method to determine the availability score based on diverse resource types, which hinders the decision-making process for failure operations, such as failover or failback, leading to potential service outages and data access disruptions.
Innovation Solution
A computer program and method that calculates an availability score by aggregating the availability of various resource types within a storage system, considering both the number of available resources and recovery events, and transmits this score to a failure manager to determine the appropriate course of action, such as initiating a failover or failback operation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a storage system uses traditional failure detection methods, then the system can detect failures, but the decision-making process for failure operations is insufficient leading to potential service outages
Solution Approach 1:
The patent replaces traditional mechanical failure detection mechanisms with an automated availability scoring system that calculates scores based on multiple resource types (storage, compute, network, etc.) and uses predefined thresholds to automatically trigger failover operations, eliminating manual decision-making complexity
Solution Approach 2:
The system changes the parameter of failure detection from binary (failure/normal) to a continuous availability score that aggregates multiple resource availability metrics, enabling more nuanced decision-making while maintaining automated operation through threshold-based triggers
2Reliability
If the system performs frequent failover operations to maintain availability, then service continuity is improved, but system stability deteriorates due to excessive operations
Solution Approach 1:
The system implements feedback control by continuously monitoring availability scores, comparing them against thresholds, and only triggering failover operations when thresholds are breached, thereby maintaining service continuity while preventing excessive operations through controlled response to actual system state changes
Solution Approach 2:
The patent applies partial action by using threshold-based triggers that activate failover only when availability scores fall below predetermined levels, rather than continuously switching systems, thus maintaining adequate service continuity while avoiding the instability of excessive failover operations
3Measurement precision
If the system monitors multiple resource types in detail, then availability assessment accuracy is improved, but the complexity of resource monitoring increases
Solution Approach 1:
The system achieves universality by creating a unified availability scoring mechanism that aggregates multiple resource type metrics (storage, compute, network, etc.) into a single comprehensive score, maintaining high measurement precision while reducing monitoring complexity through consolidated evaluation
4Stability of the object's composition
If the system delays failover operations to assess full system health, then premature failover is avoided, but service disruption time increases
Solution Approach 1:
The system applies preliminary action by pre-calculating availability scores and establishing threshold triggers before failures occur, enabling rapid automated response when thresholds are breached without requiring time-consuming post-failure assessment, thus avoiding both premature and delayed failover
Data Source
AI summary
Provided are a computer program product, system, and method for determining an availability score based on available resources of different resource types in a storage system to determine whether to perform a failure operation for the storage system. Information is maintained indicating availability of a plurality of storage system resources for a plurality of resource types. An availability score is calculated as a function of a number of available resources of the resource types. Information on the availability score is transmitted to a failure manager. The failure manager uses the transmitted availability information to determine whether to initiate a storage system failure mode for the storage system.


