RAID Controller Rebuild Isolation via Dynamic Assignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

RAID systems face increased load and access delays when executing rebuild processes for failed disks, affecting other RAID groups and the entire system, as the rebuild process burdens the RAID controller, leading to inhibited access to healthy groups.

Innovation Solution

A storage management system that dynamically reassigns control apparatuses to manage data redundancy, allowing a separate control apparatus to handle rebuild processes for affected groups without impacting other RAID groups by switching the control assignment and utilizing an information processing apparatus to manage the reassignment and rebuild process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a rebuild process is executed for a failed disk in a RAID group, then data redundancy is restored, but the load on the RAID controller increases and access to other RAID groups is inhibited

Engineering Contradiction:
Improvedata redundancyVSAvoidaccess speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system segments the RAID controller functionality by separating the rebuild process execution from normal I/O operations. When a disk failure is detected, the affected RAID group is isolated and the rebuild process is executed independently, preventing it from impacting other RAID groups. This segmentation allows normal I/O operations to continue on healthy RAID groups while the rebuild process runs on the affected group alone.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The rebuild process is extracted from the normal RAID controller operation path. When a disk failure occurs, the system extracts the affected RAID group from the pool of actively managed groups and directs it to a dedicated rebuild execution unit or isolated processing path. This extraction ensures that the resource-intensive rebuild operation does not contend with normal I/O operations for system resources.

Inventive Principle:
Principle #2Taking out (Extraction)

2Device complexity

If multiple RAID groups are managed by a single RAID controller, then device complexity is reduced, but executing rebuild processes adversely affects other controlled RAID groups

Engineering Contradiction:
Improvecontroller quantityVSAvoidsystem performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system dynamically adjusts the operational mode of the RAID controller based on system conditions. During normal operation, multiple RAID groups are managed by a single controller to maintain simplicity. However, when a disk failure is detected, the controller dynamically switches to a specialized rebuild execution mode that isolates the affected RAID group, temporarily transforming the controller's behavior to prevent performance degradation of other groups.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system introduces an intermediary management layer that sits between the RAID controller and the RAID groups. This intermediary monitors the health status of disks and dynamically redirects rebuild operations for failed disks to dedicated rebuild execution units or isolated processing paths, preventing the rebuild process from directly impacting normal I/O operations on other RAID groups.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS8972777B2Method and system for storage management
Publication Date: 2015.03.03 FUJITSU LTD
  • US8972777B2 patent drawing
  • US8972777B2 patent drawing
  • US8972777B2 patent drawing

AI summary

Multiple storage apparatuses are provided, at least part of which are individually incorporated into one of storage groups. Each of multiple control apparatuses is configured to, when assigned one or more of the storage groups each including one or more of the storage apparatuses, control data storage by storing data designating each assigned storage group redundantly in the storage apparatuses of the assigned storage group. An information processing apparatus is configured to, when a storage group with data redundancy being lost is detected, make a change in control apparatus assignment for the storage groups in such a manner that a storage group different from the detected storage group is not assigned to a control apparatus with the detected storage group assigned thereto. Subsequently, the information processing apparatus causes the control apparatus to execute a process of restoring the data redundancy of the detected storage group.