RAID Drive Preemptive Failure Testing via Data Mirroring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing RAID systems lack a method to preemptively test and maintain drive redundancy, leading to increased likelihood of data loss due to drive failures, especially after power loss events.

Innovation Solution

A method for preemptive failure testing in RAID arrays, which involves selecting a drive, mirroring its data to spare storage, and then testing the drive to identify potential failures, allowing for proactive replacement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If drives are tested by powering off and on again, then latent problems are exposed, but drive failure likelihood increases

Engineering Contradiction:
Improvedetection of latent drive problemsVSAvoiddrive failure likelihood
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system performs preliminary testing actions by powering drives off and on before they fail in production. This preliminary action exposes latent problems early when drives are still functional and can be replaced, rather than waiting for failures to occur during normal operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a cushioning effect by having spare drives ready and using RAID redundancy to protect against the increased failure risk during testing. The redundancy acts as a buffer that allows aggressive testing without risking data loss, as failed drives can be replaced from spare inventory.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

2Quantity of substance

If multiple drives from the same batch are used, then system cost is reduced, but likelihood of simultaneous failures increases

Engineering Contradiction:
Improvedrive batch uniformityVSAvoidsimultaneous failure risk
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The system implements feedback by continuously monitoring drive health metrics and testing results. Drives that fail preliminary testing are identified and replaced before being deployed to production, preventing defective drives from entering the array. This feedback loop breaks the correlation between drives from the same batch by filtering out failures early.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Testing is performed preliminarily on drives before they are installed in the RAID array. This preliminary screening action identifies and removes defective drives from batches before deployment, preventing simultaneous failures that would occur if multiple defective drives from the same batch were deployed together.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If drive testing is performed in production environment, then system resilience improves, but testing complexity increases

Engineering Contradiction:
Improvesystem resilience to power lossVSAvoidtesting process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system performs self-service testing where the RAID controller automatically manages the testing process, including selecting which drives to test, powering them off and on, and monitoring for failures. This automation reduces the complexity burden on operators while maintaining high reliability through continuous self-testing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements periodic testing of drives in the RAID array, cycling through drives at scheduled intervals to power them off and on. This periodic action maintains system resilience without requiring continuous complex monitoring, as testing is performed in regular, manageable cycles rather than continuously.

Inventive Principle:
Principle #19Periodic action

4Reliability

If drives are tested frequently, then preemptive failure detection improves, but system downtime increases

Engineering Contradiction:
Improvepreemptive failure detection capabilityVSAvoidsystem downtime for testing
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs partial testing by selecting a subset of drives to test at any given time rather than testing all drives simultaneously. This partial action approach maintains preemptive detection capability while distributing testing over time, preventing excessive downtime that would occur if all drives were tested at once.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

Testing is performed periodically on a rotating basis across different drives in the array. Instead of testing all drives at once or continuously, the system cycles through drives in periodic intervals, maintaining detection capability while minimizing overall system downtime through time-distributed testing.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12292806B2Testing drives in a redundant array of independent disks (raid)
Publication Date: 2025.05.06 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12292806B2 patent drawing
  • US12292806B2 patent drawing
  • US12292806B2 patent drawing

AI summary

A method, computer program product, and computer system are provided for testing drives in a redundant array of independent disks (RAID) array. The method includes: mirroring data from a selected drive to be tested in a RAID array to spare storage space in the RAID array; and, once the data is successfully mirrored, testing the selected drive to identify a preemptive failure of the selected drive. The RAID may be a traditional RAID (TRAID) array and the spare space may be a spare physical drive independent of array drive members. The RAID array may alternatively be a distributed RAID (DRAID) array and the spare space may be spare capacity spread through the array drive members.