Storage Controller Failure Recovery Policy for RAID Groups

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In storage systems, unnecessary drive replacements occur due to misidentification of failure causes, leading to wasted costs and reduced performance, as the system currently lacks the ability to differentiate between drive failures and other fault sources.

Innovation Solution

A storage system with a storage controller that manages RAID groups and employs policy management information to specify appropriate failure recovery processing based on the type of command error, allowing for targeted recovery and cost-performance optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the system performs comprehensive failure cause analysis before replacement, then replacement accuracy improves, but system performance and operating rate decrease

Engineering Contradiction:
Improvefailure cause identification accuracyVSAvoidsystem operating rate
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary failure analysis using policy management information and error type classification before initiating replacement procedures. This preliminary action identifies whether the failure originates from the drive unit itself or external factors, allowing accurate replacement decisions to be made in advance without delaying system operations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces policy management information as an intermediary mechanism that mediates between failure detection and replacement execution. This intermediary layer classifies errors into different types (drive unit failures vs. external failures) and determines appropriate recovery actions, preventing unnecessary replacements while maintaining system performance.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system replaces drives immediately upon failure detection, then system reliability improves, but operational cost increases due to unnecessary replacements

Engineering Contradiction:
Improvesystem reliabilityVSAvoidreplacement cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The system implements feedback mechanisms that continuously monitor error types and update replacement decisions based on policy management information. When errors are classified as external failures rather than drive unit failures, the feedback loop prevents unnecessary replacements, reducing costs while maintaining reliability through targeted recovery actions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the parameter of replacement timing from immediate to conditional based on error type classification. By modifying this parameter according to policy management information and failure analysis results, the system avoids unnecessary replacements caused by external factors while ensuring timely replacement when drive unit failures are confirmed.

Inventive Principle:
Principle #35Parameter changes

3Speed

If the system implements no failure analysis and replaces drives immediately, then replacement speed improves, but cost increases due to misidentification of failure causes

Engineering Contradiction:
Improvereplacement speedVSAvoidwasted replacement cost
Core Design Contradiction:
SpeedVSLoss of energy

Solution Approach 1:

The system performs preliminary classification of error types using policy management information before replacement. This preliminary action quickly distinguishes between drive unit failures requiring replacement and external failures not requiring replacement, maintaining fast replacement speed while preventing wasted costs through accurate failure cause identification.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10509700B2Storage system and storage management method
Publication Date: 2019.12.17 HITACHI VANTARA LTD
  • US10509700B2 patent drawing
  • US10509700B2 patent drawing
  • US10509700B2 patent drawing

AI summary

A storage system has a storage controller and a RAID group. The storage controller has policy management information such that one failure recovery process among a plurality of differing failure recovery processes is associated with each RAID group, and when an error in a command issued to a RAID group is detected, the failure recovery process associated with the RAID group to which the command was issued is specified on the basis of the policy management information, and the specified failure recovery process is executed.