RAID Controller Rebuild Isolation via Dynamic Assignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
RAID systems face increased load and access delays when executing rebuild processes for failed disks, affecting other RAID groups and the entire system, as the rebuild process burdens the RAID controller, leading to inhibited access to healthy groups.
Innovation Solution
A storage management system that dynamically reassigns control apparatuses to manage data redundancy, allowing a separate control apparatus to handle rebuild processes for affected groups without impacting other RAID groups by switching the control assignment and utilizing an information processing apparatus to manage the reassignment and rebuild process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a rebuild process is executed for a failed disk in a RAID group, then data redundancy is restored, but the load on the RAID controller increases and access to other RAID groups is inhibited
Solution Approach 1:
The system segments the RAID controller functionality by separating the rebuild process execution from normal I/O operations. When a disk failure is detected, the affected RAID group is isolated and the rebuild process is executed independently, preventing it from impacting other RAID groups. This segmentation allows normal I/O operations to continue on healthy RAID groups while the rebuild process runs on the affected group alone.
Solution Approach 2:
The rebuild process is extracted from the normal RAID controller operation path. When a disk failure occurs, the system extracts the affected RAID group from the pool of actively managed groups and directs it to a dedicated rebuild execution unit or isolated processing path. This extraction ensures that the resource-intensive rebuild operation does not contend with normal I/O operations for system resources.
2Device complexity
If multiple RAID groups are managed by a single RAID controller, then device complexity is reduced, but executing rebuild processes adversely affects other controlled RAID groups
Solution Approach 1:
The system dynamically adjusts the operational mode of the RAID controller based on system conditions. During normal operation, multiple RAID groups are managed by a single controller to maintain simplicity. However, when a disk failure is detected, the controller dynamically switches to a specialized rebuild execution mode that isolates the affected RAID group, temporarily transforming the controller's behavior to prevent performance degradation of other groups.
Solution Approach 2:
The system introduces an intermediary management layer that sits between the RAID controller and the RAID groups. This intermediary monitors the health status of disks and dynamically redirects rebuild operations for failed disks to dedicated rebuild execution units or isolated processing paths, preventing the rebuild process from directly impacting normal I/O operations on other RAID groups.
Data Source
AI summary
Multiple storage apparatuses are provided, at least part of which are individually incorporated into one of storage groups. Each of multiple control apparatuses is configured to, when assigned one or more of the storage groups each including one or more of the storage apparatuses, control data storage by storing data designating each assigned storage group redundantly in the storage apparatuses of the assigned storage group. An information processing apparatus is configured to, when a storage group with data redundancy being lost is detected, make a change in control apparatus assignment for the storage groups in such a manner that a storage group different from the detected storage group is not assigned to a control apparatus with the detected storage group assigned thereto. Subsequently, the information processing apparatus causes the control apparatus to execute a process of restoring the data redundancy of the detected storage group.


