SSD Wearout Staggering via Write Rate Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional RAID techniques are ineffective in managing wearout failures in Solid State Drive (SSD) arrays, as SSD failures often occur simultaneously due to wearout, leading to high probabilities of data loss during replacement periods.

Innovation Solution

Implementing a method where a controller monitors the write rate of SSDs, predicts their mean time to failure, and issues notifications to replace disks before wearout, staggering the wearout periods of storage disks to prevent simultaneous failures and data loss.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If RAID techniques are applied to SSD arrays, then data protection is improved, but the system becomes vulnerable to simultaneous wearout failures

Engineering Contradiction:
Improvedata protectionVSAvoidsimultaneous wearout failures
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary actions by monitoring write rates and predicting MTTF before wearout occurs. The controller issues advance notifications to replace disks before they fail, preventing simultaneous failures from occurring during replacement operations. This proactive approach transforms the passive RAID protection into an active prevention system.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements continuous feedback by monitoring write rates and calculating predicted MTTF for each SSD. Based on this feedback, the controller dynamically determines replacement timing and issues notifications at optimal moments. This closed-loop control ensures replacements are scheduled when wearout is predicted but before simultaneous failures can occur.

Inventive Principle:
Principle #23Feedback

2Reliability

If SSDs are replaced before wearout, then data loss is prevented, but the operational life of SSDs is reduced

Engineering Contradiction:
Improvedata loss preventionVSAvoidoperational life of SSDs
Core Design Contradiction:
ReliabilityVSDuration of action of moving object

Solution Approach 1:

The system performs preliminary replacement actions based on predicted MTTF calculations. By replacing disks before actual wearout occurs but after predicting their remaining life, the system optimizes the balance between preventing data loss and maximizing SSD operational life. The preliminary notification system allows scheduled replacements that prevent catastrophic failures while minimizing premature replacements.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If wearout of SSDs is accelerated for testing, then failure characteristics are improved, but the actual service life is reduced

Engineering Contradiction:
Improvefailure characteristicsVSAvoidservice life
Core Design Contradiction:
Measurement precisionVSDuration of action of moving object

Solution Approach 1:

The system uses feedback from monitored write rates to calculate predicted MTTF, which then guides replacement scheduling. This feedback mechanism allows the system to understand actual wear patterns without accelerating them, maintaining service life while enabling precise prediction of failure characteristics for better planning.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10620847B2Shifting wearout of storage disks
Publication Date: 2020.04.14 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10620847B2 patent drawing
  • US10620847B2 patent drawing
  • US10620847B2 patent drawing

AI summary

Technical solutions are described to forestall data loss caused by wearout of storage disks in an array of storage disks in a storage system by monitoring a rate of writes for a first storage disk in the array and determining a mean time to failure of the first storage disk. A start time is determined based on the mean time to failure, a number of storage disks in the array, and a time to replace a storage disk in the array. At the start time, a notification is issued as an alert to replace the first storage disk to forestall data loss caused by wearout of a second storage disk in conjunction with a wearout of the first storage disk.