Autonomic RAID Controller Error Recovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing storage system maintenance procedures require user intervention for error recovery processes, leading to extended downtime, reduced drive lifecycle, and negative impacts on system performance and availability.

Innovation Solution

An autonomic method for managing error recovery procedures in RAID systems, where a resource controller identifies and schedules drive error recovery processes based on predefined operational goals, minimizing user intervention and optimizing system reliability, redundancy, and performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If user intervention is required to instigate error recovery procedures, then drive maintenance can be performed, but system downtime increases and drive lifecycle decreases

Engineering Contradiction:
Improvedrive maintenanceVSAvoidsystem downtime
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables drives to self-diagnose and self-repair by automatically detecting when error recovery procedures are needed and executing them without user intervention. The controller monitors drive health metrics and autonomously initiates format units, table rebuilds, and other maintenance procedures, allowing the system to service itself and eliminate downtime associated with manual intervention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary monitoring and detection of drive conditions that indicate upcoming failures or performance degradation. By identifying these conditions early and automatically scheduling error recovery procedures before critical failures occur, the system prevents extended downtime and maintains continuous operational availability while extending drive lifecycle.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If error recovery procedures are executed manually, then drive health can be restored, but system availability decreases

Engineering Contradiction:
Improvedrive health restorationVSAvoidsystem availability
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The autonomous error recovery system allows drives to restore their own health automatically by detecting when procedures like format units or table rebuilds are needed and executing them without taking the system offline. This self-service capability maintains system availability while ensuring drive health is restored, eliminating the need for manual intervention that would cause downtime.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system ensures continuous useful action by automatically managing error recovery procedures in a way that minimizes disruption to RAID array operations. The controller schedules and executes drive maintenance procedures seamlessly, maintaining system availability and ensuring that productivity is not compromised while drive health is continuously monitored and restored as needed.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If error recovery procedures are performed on drives, then drive reliability improves, but drive lifecycle decreases

Engineering Contradiction:
Improvedrive reliabilityVSAvoiddrive lifecycle
Core Design Contradiction:
ReliabilityVSDuration of action of stationary object

Solution Approach 1:

The system performs preliminary detection of drive conditions that indicate future failures and schedules error recovery procedures proactively before critical failures occur. By addressing minor issues early through automatic format units, table rebuilds, and other maintenance procedures, the system prevents catastrophic failures that would end drive lifecycle, thereby extending drive life while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The autonomous error recovery system extends drive lifecycle by enabling drives to self-maintain through automatic execution of recovery procedures. This continuous self-service approach addresses wear and potential failures progressively rather than allowing cumulative degradation, thereby extending the operational life of drives while maintaining high reliability levels throughout the extended lifecycle.

Inventive Principle:
Principle #25Self-service

4Extent of automation

If automated error recovery procedures are implemented, then user intervention is reduced, but system complexity increases

Engineering Contradiction:
Improveerror recovery automationVSAvoidsystem complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The controller is designed with multi-functionality to handle both normal RAID array operations and autonomous error recovery procedures within a single integrated system. The controller universally manages drive monitoring, health assessment, procedure selection, execution, and verification, eliminating the need for separate dedicated systems and reducing overall system complexity despite the advanced automation capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9940211B2Resource system management
Publication Date: 2018.04.10 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US9940211B2 patent drawing
  • US9940211B2 patent drawing
  • US9940211B2 patent drawing

AI summary

A resource system comprises a plurality of resource elements and a resource controller connected to the resource elements and operating the resource elements according to a predefined set of operational goals. A method of operating the resource system comprises the steps of identifying error recovery procedures that could be executed by the resource elements, categorizing each identified error recovery procedure in relation to the predefined set of operational goals, detecting that an error recovery procedure is to be performed on a specific resource element, deploying one or more actions in relation to the resource elements according to the categorization of the detected error recovery procedure, and performing the detected error recovery procedure on the specific resource element.