Availability Score Calculation for Storage System Failure Management

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current storage systems lack an effective method to determine the availability score based on diverse resource types, which hinders the decision-making process for failure operations, such as failover or failback, leading to potential service outages and data access disruptions.

Innovation Solution

A computer program and method that calculates an availability score by aggregating the availability of various resource types within a storage system, considering both the number of available resources and recovery events, and transmits this score to a failure manager to determine the appropriate course of action, such as initiating a failover or failback operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a storage system uses traditional failure detection methods, then the system can detect failures, but the decision-making process for failure operations is insufficient leading to potential service outages

Engineering Contradiction:
Improvesystem availabilityVSAvoidfailure decision complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical failure detection mechanisms with an automated availability scoring system that calculates scores based on multiple resource types (storage, compute, network, etc.) and uses predefined thresholds to automatically trigger failover operations, eliminating manual decision-making complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes the parameter of failure detection from binary (failure/normal) to a continuous availability score that aggregates multiple resource availability metrics, enabling more nuanced decision-making while maintaining automated operation through threshold-based triggers

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the system performs frequent failover operations to maintain availability, then service continuity is improved, but system stability deteriorates due to excessive operations

Engineering Contradiction:
Improveservice continuityVSAvoidsystem stability
Core Design Contradiction:
ReliabilityVSStability of the object's composition

Solution Approach 1:

The system implements feedback control by continuously monitoring availability scores, comparing them against thresholds, and only triggering failover operations when thresholds are breached, thereby maintaining service continuity while preventing excessive operations through controlled response to actual system state changes

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies partial action by using threshold-based triggers that activate failover only when availability scores fall below predetermined levels, rather than continuously switching systems, thus maintaining adequate service continuity while avoiding the instability of excessive failover operations

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If the system monitors multiple resource types in detail, then availability assessment accuracy is improved, but the complexity of resource monitoring increases

Engineering Contradiction:
Improveavailability assessment accuracyVSAvoidresource monitoring complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system achieves universality by creating a unified availability scoring mechanism that aggregates multiple resource type metrics (storage, compute, network, etc.) into a single comprehensive score, maintaining high measurement precision while reducing monitoring complexity through consolidated evaluation

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Stability of the object's composition

If the system delays failover operations to assess full system health, then premature failover is avoided, but service disruption time increases

Engineering Contradiction:
Improvefailover timing accuracyVSAvoidservice disruption time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The system applies preliminary action by pre-calculating availability scores and establishing threshold triggers before failures occur, enabling rapid automated response when thresholds are breached without requiring time-consuming post-failure assessment, thus avoiding both premature and delayed failover

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9703619B2Determining an availability score based on available resources of different resource types in a storage system to determine whether to perform a failure operation for the storage system
Publication Date: 2017.07.11 MAPLEBEAR INC
  • US9703619B2 patent drawing
  • US9703619B2 patent drawing
  • US9703619B2 patent drawing

AI summary

Provided are a computer program product, system, and method for determining an availability score based on available resources of different resource types in a storage system to determine whether to perform a failure operation for the storage system. Information is maintained indicating availability of a plurality of storage system resources for a plurality of resource types. An availability score is calculated as a function of a number of available resources of the resource types. Information on the availability score is transmitted to a failure manager. The failure manager uses the transmitted availability information to determine whether to initiate a storage system failure mode for the storage system.