Availability Calculation for Storage Node Maintenance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In large-scale SDS systems with multiple nodes, frequent maintenance due to individual node failures leads to increased maintenance time and operational costs, making it challenging to maintain system availability and determine when redundant configurations are compromised.

Innovation Solution

An information processing system that calculates the availability level based on configuration and failure information, allowing for proactive maintenance scheduling to minimize system downtime and reduce operational expenses by continuing I/O processing only when necessary.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If maintenance work is performed each time a failure occurs in one node, then system reliability is maintained, but maintenance time and operational costs increase significantly

Engineering Contradiction:
Improvesystem reliabilityVSAvoidmaintenance time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary assessment of failure impact by calculating availability levels before maintenance is executed. The processor evaluates configuration information and failure information to determine whether immediate maintenance is necessary or if the system can continue operating with reduced availability, thereby avoiding unnecessary maintenance interruptions.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention introduces availability level as a dynamic parameter that changes based on system configuration and failure patterns. By monitoring this parameter and comparing it against thresholds, the system adapts maintenance decisions rather than following fixed rules, allowing maintenance to be deferred when availability remains acceptable.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If maintenance work is performed frequently to maintain reliability, then system availability is preserved, but productivity decreases due to repeated interruptions

Engineering Contradiction:
Improvesystem availabilityVSAvoidsystem productivity
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements feedback by continuously monitoring failure information and recalculating availability levels after each failure event. This feedback loop enables dynamic adjustment of maintenance timing, allowing the system to maintain adequate availability while minimizing productivity loss by avoiding unnecessary maintenance interruptions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The maintenance strategy transitions from static (fixed schedule or immediate response) to dynamic, where maintenance timing is adjusted based on real-time system state. The processor dynamically determines whether to proceed with maintenance or allow continued operation by evaluating current availability levels against thresholds.

Inventive Principle:
Principle #15Dynamics

3Productivity

If the system continues operation with failed nodes, then productivity is maintained, but reliability deteriorates when redundant configuration is compromised

Engineering Contradiction:
Improvesystem operation continuityVSAvoidredundant configuration integrity
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system takes preliminary anti-action by predicting when reliability will deteriorate to unacceptable levels. By calculating availability levels and comparing against thresholds before actual failure occurs, the system prevents reliability collapse by scheduling maintenance proactively rather than reactively.

Inventive Principle:
Principle #9Preliminary anti-action

Solution Approach 2:

The availability level calculation acts as an intermediary between system operation and maintenance decisions. This intermediary metric translates complex system state (configuration + failures) into a simple threshold comparison that guides whether to maintain productivity by continuing operation or protect reliability by performing maintenance.

Inventive Principle:
Principle #24Intermediary (Mediator)

4Reliability

If monitoring is performed to determine redundant configuration status, then reliability can be maintained, but system complexity increases

Engineering Contradiction:
Improveredundant configuration monitoringVSAvoidmonitoring system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The monitoring function is implemented as a self-service capability within the existing storage nodes. Each storage node's processor independently evaluates configuration information and failure information to calculate availability levels, eliminating the need for separate complex monitoring infrastructure while maintaining reliable monitoring of redundant configuration status.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11221935B2Information processing system, information processing system management method, and program thereof
Publication Date: 2022.01.11 HITACHI VANTARA LTD
  • US11221935B2 patent drawing
  • US11221935B2 patent drawing
  • US11221935B2 patent drawing

AI summary

An object of the present invention is to provide a redundant information processing system that can continue its operation without stopping as much as possible. To achieve this, the information processing system includes multiple storage nodes including processors, memories, and storage devices, as well as a network for connecting the storage nodes. The processor of at least one storage node performs the steps of: obtaining the configuration information of the information processing system; obtaining first failure information related to a first failure occurred in the information processing system; calculating the availability level of the information processing system when a second failure further occurs, based on the configuration information and the first failure information; and controlling the operation of the system based on the availability level. Then, the processor outputs a notification.