RAID Drive Risk Redistribution to Reduce Data Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current RAID systems, particularly RAID-5, experience significant data loss due to drive failures combined with media errors, highlighting the need for methods to reduce such incidents and provide better reporting and statistics to encourage transitions to more robust RAID levels like RAID-6.

Innovation Solution

A method that identifies higher and lower risk storage drives within RAID arrays and swaps them to distribute risk more evenly, using a reporting module to document and mitigate data loss by potentially converting RAID-5 arrays to RAID-6 arrays, thereby reducing data loss incidents.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If RAID-5 is used to provide data redundancy, then storage capacity is increased, but data loss risk increases when combined with media errors

Engineering Contradiction:
Improvestorage capacityVSAvoiddata loss risk
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent changes the RAID level parameter from RAID-5 to RAID-6, which adds an additional parity value. This parameter change increases the reliability by enabling the system to withstand two simultaneous drive failures instead of one, while maintaining the same storage capacity through optimized striping configurations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements preemptive risk assessment by identifying drives with elevated failure risks before failures occur. The system proactively redistributes data and recalculates risk profiles, cushioning against potential data loss by preparing remediation strategies in advance rather than reacting after failures happen.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Device complexity

If storage drives are concentrated in fewer RAID arrays, then management complexity is reduced, but risk distribution worsens

Engineering Contradiction:
Improvemanagement complexityVSAvoidrisk distribution
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent implements dynamic risk assessment and redistribution by continuously monitoring drive health metrics, failure rates, and operational characteristics. The system automatically recalculates risk profiles and redistributes drives across RAID arrays based on current risk levels, creating a dynamic balancing act between management simplicity and risk distribution that adapts to changing conditions.

Inventive Principle:
Principle #15Dynamics

3Reliability

If RAID-6 is used instead of RAID-5, then data protection is improved, but storage capacity decreases

Engineering Contradiction:
Improvedata protectionVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent optimizes the striping configuration parameters when transitioning from RAID-5 to RAID-6 by adjusting stripe width, block size, and drive distribution patterns. These parameter changes minimize the overhead impact of the additional parity value, recovering more storage capacity compared to standard RAID-6 implementations while maintaining enhanced data protection capabilities.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11442826B2Reducing incidents of data loss in raid arrays having the same raid level
Publication Date: 2022.09.13 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11442826B2 patent drawing
  • US11442826B2 patent drawing
  • US11442826B2 patent drawing

AI summary

A method for reducing incidents of data loss in redundant arrays of independent disks (RAIDs) having the same RAID level is disclosed. In one embodiment, such a method identifies, in a data storage environment, a set of RAIDs having a common RAID level. The method also identifies, in the set of RAIDs, higher risk storage drives having a failure risk above a threshold and lower risk storage drives having a failure risk below the threshold. The method swaps, within the RAIDs, higher risk storage drives with lower risk storage drives to more evenly distribute higher risk storage drives across the RAIDs. A corresponding system and computer program product are also disclosed.