Cold Spare Disk Power Cycling for RAID Reliability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

RAID systems face data loss due to simultaneous disk failures, as hot spare disks, continuously powered on, often fail at similar times to the disks they replace, leading to vulnerability during failover.

Innovation Solution

Configuring spare disks as cold spares, which remain powered down in standby mode, are individually powered on and tested at predetermined intervals, and then returned to standby, reducing wear and increasing their effective lifespan.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If hot spare disks are continuously powered on, then failover capability is immediately available, but disk lifespan is reduced due to continuous wear

Engineering Contradiction:
Improvefailover capabilityVSAvoiddisk lifespan
Core Design Contradiction:
ReliabilityVSDuration of action of moving object

Solution Approach 1:

The patent applies dynamics by transitioning spare disks from a static continuously-powered state to a dynamic state where they alternate between powered-on and powered-off states. The system dynamically adjusts power states based on whether failover is needed, thereby reducing unnecessary wear while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the operational parameter of disk power state from continuous operation to periodic operation. By controlling the power state parameter (on/off) based on operational needs, the system extends disk lifespan while maintaining failover capability when required.

Inventive Principle:
Principle #35Parameter changes

2Duration of action of moving object

If cold spare disks remain powered down, then disk lifespan is extended, but failover response time is increased

Engineering Contradiction:
Improvedisk lifespanVSAvoidfailover response time
Core Design Contradiction:
Duration of action of moving objectVSLoss of time

Solution Approach 1:

The patent applies preliminary action by periodically powering on cold spare disks to perform self-tests and health checks before actual failover is needed. This advance preparation ensures that when failover is required, the spare disk is already verified to be functional, reducing the risk of secondary failures.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements periodic action by scheduling regular power-on intervals for cold spare disks. During these periodic intervals, disks are activated for testing and then powered down again, creating a rhythm that balances lifespan extension with readiness verification.

Inventive Principle:
Principle #19Periodic action

3Reliability

If cold spare disks are periodically powered on and tested, then disk functionality is verified, but energy consumption increases

Engineering Contradiction:
Improvedisk functionalityVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by powering on cold spare disks only for the minimum necessary duration to perform health checks and self-tests. Rather than keeping them continuously powered on, the system activates them partially (intermittently) just enough to verify functionality, then powers them down to conserve energy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10360116B2Disk preservation and failure prevention in a raid array
Publication Date: 2019.07.23 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10360116B2 patent drawing
  • US10360116B2 patent drawing
  • US10360116B2 patent drawing

AI summary

Methods, computer systems, and computer program products for configuring a redundant array of independent disks (RAID) array by a processor device, include, within a RAID array, configuring spare failover disks to run as cold spares, such that the cold spare disks stay in a powered-down standby mode, wherein each cold spare disk is powered on individually at predetermined intervals, tested, and powered back down to standby mode.