RAID System with Nonvolatile Memory for Proactive Rebuild
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
RAID rebuild operations in systems with nonvolatile memory devices or solid-state drives are time-consuming and negatively impact I/O performance, as they require significant time and resources during data restoration after a failure.
Innovation Solution
A RAID system with a nonvolatile memory device and a RAID controller that monitors failure probabilities of memory chips, performing a first rebuild operation by storing data from multiple failing chips in a spare memory chip and a second rebuild operation using the spare data to restore individual failing chips, thereby reducing the time and I/O overhead during the rebuild process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a traditional RAID rebuild operation is performed after disk failure, then data integrity is restored, but the rebuild process takes a long time and considerably affects I/O performance
Solution Approach 1:
The patent performs preliminary rebuild operations by proactively identifying memory chips with high failure probability through monitoring parameters such as program/erase cycle counts. When a memory chip is identified as potentially failing, the system reconstructs its data in advance and stores it in a spare memory chip before the actual failure occurs. This preliminary action eliminates the need for time-consuming rebuild operations when failures actually occur, thus resolving the contradiction between maintaining data integrity and minimizing rebuild time.
2Reliability
If a traditional RAID rebuild operation is performed after disk failure, then data integrity is restored, but I/O performance of the entire RAID system is considerably affected
Solution Approach 1:
The system performs data reconstruction and storage in spare memory chips during periods when the identified memory chips are still functional. By completing these rebuild operations in advance, the system ensures that when actual failures occur, data can be quickly restored without initiating resource-intensive rebuild processes that would degrade I/O performance, thus maintaining both data integrity and system productivity.
3Loss of time
If proactive rebuild operations are performed on memory chips with high failure probability, then rebuild time is reduced, but additional I/O operations are required during the proactive rebuild process
Solution Approach 1:
The system monitors parameters such as program/erase cycle counts to identify memory chips approaching failure thresholds. By changing the operational state of identified chips to read-only mode and prioritizing their data migration, the system performs proactive rebuilds during periods of lower system demand. This parameter-based approach allows the system to manage I/O operations more efficiently, reducing the energy cost of proactive rebuilds while still achieving faster rebuild times when failures occur.
Data Source
AI summary
A redundant array of inexpensive disks (RAID) system including nonvolatile memory and an operating method of the same is provided. A nonvolatile memory device implemented as a RAID and including a plurality of first memory chips, which store data chunks, and a second memory chip, in which spare memory regions are defined. A RAID controller controls RAID operations and a rebuild operation of the nonvolatile memory device. The RAID controller monitors a failure probability of each of the first memory chips, and in response to detecting a failure probability of two or more first memory chips that satisfies a predefined threshold value, a first rebuild on data stored in each of the first memory chips is performed to store the data in the second memory chip. A second rebuild on data stored in the first memory chip having the failure using data stored in the second memory chip.


