RAID Rebuild Management via Distributed Storage Agents

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In shared-storage RAID environments, the lack of explicit ownership over RAID volumes can lead to delayed or prevented rebuilds of degraded volumes, and failure recognition among nodes is not immediate, hindering automatic RAID recovery.

Innovation Solution

A storage management agent is implemented in each server node to monitor RAID volumes, pausing to determine if a rebuild is initiated and taking remedial action if not, and initiating task transfers if a node fails, ensuring rebuild and failover tasks are completed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If exclusive ownership over RAID volumes is not established, then resource sharing flexibility is improved, but rebuild responsibility and automatic recovery capability deteriorate

Engineering Contradiction:
Improveresource sharing flexibilityVSAvoidautomatic recovery capability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system implements a feedback mechanism where storage management agents continuously monitor RAID volume status and automatically trigger rebuild operations when degradation is detected. This closed-loop feedback ensures that even without exclusive ownership, the system maintains reliable automatic recovery by detecting and responding to failures promptly.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The storage management agents enable the system to perform self-service by automatically initiating and monitoring rebuild operations without requiring manual intervention or explicit ownership assignment. The system monitors its own state and takes corrective action autonomously, resolving the contradiction between flexible sharing and reliable recovery.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual intervention is required for RAID rebuild, then control precision is improved, but productivity and response time deteriorate

Engineering Contradiction:
Improvecontrol precisionVSAvoidrebuild response time
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The storage management agents automate the rebuild process by automatically detecting degraded RAID volumes and initiating reconstruction without requiring manual intervention. This self-service capability maintains precise control over the rebuild process while dramatically improving response time and productivity by eliminating human delay.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary monitoring and preparation by continuously tracking RAID volume health status. When degradation is detected, the system is already positioned to immediately initiate the rebuild process, reducing response time while maintaining controlled, precise execution of the recovery operation.

Inventive Principle:
Principle #10Preliminary action

3Stability of the object's composition

If node failure is not immediately recognized, then system stability is improved, but reliability and recovery speed deteriorate

Engineering Contradiction:
Improvesystem stabilityVSAvoidfailure recognition speed
Core Design Contradiction:
Stability of the object's compositionVSReliability

Solution Approach 1:

Storage management agents implement continuous feedback monitoring of node status and RAID volume health. This real-time feedback enables immediate recognition of node failures while maintaining system stability through proactive detection and automated response, preventing the delays associated with passive failure detection.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary monitoring and status assessment before failures manifest as critical issues. By continuously tracking node health and RAID volume status, the system is prepared to immediately respond to failures, maintaining both stability and rapid recovery capability.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If storage management agents monitor all RAID volumes, then reliability is improved, but device complexity increases

Engineering Contradiction:
Improvefailure monitoring coverageVSAvoidagent implementation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the monitoring function by implementing independent storage management agents on each node, each responsible for monitoring RAID volumes within its scope. This segmentation distributes complexity across multiple simple, identical units rather than requiring a single complex centralized system, improving reliability coverage while managing implementation complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7536586B2System and method for the management of failure recovery in multiple-node shared-storage environments
Publication Date: 2009.05.19 DELL PROD LP
  • US7536586B2 patent drawing
  • US7536586B2 patent drawing
  • US7536586B2 patent drawing

AI summary

A storage architecture and method for managing the operation of a network in a RAID environment is provided in which a storage management agent is included in each server node of the network. The storage management agents monitor the status of the drives of the storage array in shared storage. If a storage management agent identifies a failed drive, the storage management agent monitors the rebuild of the degraded RAID volume. During the rebuild of a degraded RAID volume, the storage management agent determines if a server node has failed, and, if required, initiates the transfer of the RAID rebuild tasks of the failed server node to another server node.