RAID Hot Spare Pre-Population for Rebuild Time Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In RAID systems, the rebuild time for a failed disk drive can be excessively long, leading to potential data loss if another disk drive fails during the rebuild process, especially with larger disk drives where rebuild times can take days, and existing methods do not efficiently utilize controller resources for proactive data protection.

Innovation Solution

Implementing a method that uses a dedicated hot spare with a copyback process to proactively write data to the hot spare during low controller usage, allowing for rapid replacement and minimizing data loss by incorporating intelligence in the RAID controller to manage I/O requests and prioritize rebuilds, thereby reducing rebuild time and ensuring data integrity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a hot spare is used to replace a failed disk drive, then data protection is improved, but rebuild time becomes excessively long for large disk drives

Engineering Contradiction:
Improvedata protectionVSAvoidrebuild time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by proactively copying data from active disk drives to the hot spare before failures occur. The system continuously monitors controller resource availability and initiates copy operations during low-utilization periods, so that when a failure occurs, the hot spare is already populated with current data or can be rapidly populated, dramatically reducing rebuild time from days to hours or minutes.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the hot spare sits idle until failure, then controller resources are conserved, but valuable time is lost during rebuild operations

Engineering Contradiction:
Improvecontroller resource efficiencyVSAvoidrebuild time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system dynamically adjusts hot spare population operations based on real-time controller resource availability. The patent implements a monitoring mechanism that detects when controller utilization is low and automatically initiates data copy operations to the hot spare. This dynamic approach allows the system to utilize controller resources efficiently during normal operation while rapidly responding to failure events, resolving the contradiction between resource conservation and rebuild speed.

Inventive Principle:
Principle #15Dynamics

3Reliability

If rebuild operations consume significant controller resources, then data protection is maintained, but system performance and other I/O operations are degraded

Engineering Contradiction:
Improvedata protectionVSAvoidsystem performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

By performing data copy operations to the hot spare in advance during periods of low controller utilization, the system prepares redundancy without impacting normal I/O performance. When a failure occurs, the rebuild operation can proceed using the pre-populated hot spare or can be rapidly initiated without competing for resources with active I/O operations, thus maintaining both data protection and system performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements periodic monitoring of controller resource usage and performs hot spare population operations during identified low-utilization windows. This periodic approach allows the system to gradually populate the hot spare without creating sustained resource contention, distributing the I/O load over time rather than concentrating it during failure events, thereby maintaining overall system performance.

Inventive Principle:
Principle #19Periodic action

4Quantity of substance

If disk drive size increases to provide more storage capacity, then storage capacity is improved, but rebuild time increases proportionally

Engineering Contradiction:
Improvestorage capacityVSAvoidrebuild time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent addresses the rebuild time issue for large-capacity drives by proactively population the hot spare with data from active drives before failures occur. This preliminary action decouples the rebuild time from the disk drive capacity, as the hot spare is already prepared with the necessary data or can be rapidly populated using parallel transfer mechanisms, reducing the impact of large drive sizes on rebuild duration.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8417989B2Method and system for extra redundancy in a raid system
Publication Date: 2013.04.09 AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
  • US8417989B2 patent drawing
  • US8417989B2 patent drawing
  • US8417989B2 patent drawing

AI summary

A system and method of creating an extra redundancy in a RAID system is disclosed. In one embodiment, one or more RAID arrays are created. Each RAID array comprises a plurality of disk drives. Further, a respective dedicated hot spare is created for each RAID array. Furthermore, data is copied from each RAID array to the respective dedicated hot spare using a copyback process based on a predetermined controller usage threshold value.