Storage Array Device Recovery via Data Copying
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for managing storage device failures in RAID arrays are processor-intensive, leading to saturation and potential discarding of expensive storage drives with non-fatal errors, which can be recoverable.
Innovation Solution
A method to identify and recover storage devices with non-fatal errors by rebuilding the storage array with a spare device, copying data from the spare to the potentially failed device, and reconfiguring it to fix errors, allowing for validation and reuse without disrupting operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a storage device is immediately swapped out when errors are detected, then the storage array operations are not disrupted, but the potentially recoverable storage device is discarded and expensive resources are lost
Solution Approach 1:
The patent implements a recovery process where storage devices initially marked for discarding are instead recovered through data copying from spare devices, validation testing, and conditional reintegration into the storage array, thereby preventing premature loss of expensive storage resources
Solution Approach 2:
The patent performs preliminary data copying from spare storage devices to potentially failed devices before final validation and reintegration, ensuring that recovery operations are prepared in advance and do not disrupt ongoing storage array operations
2Reliability
If a storage array is rebuilt using processor-intensive operations, then failed storage devices are replaced and data is restored, but the processors become saturated and I/O requests are delayed
Solution Approach 1:
The patent performs data copying operations in advance during off-peak periods or concurrently with minimal impact, so that when storage devices need to be reintegrated, the data is already prepared and validation can proceed quickly without saturating processors during critical I/O operations
Solution Approach 2:
The patent uses data copying from spare storage devices to potentially failed devices as the core recovery mechanism, enabling restoration of data integrity through relatively low-overhead copy operations rather than intensive reconstruction algorithms
3Reliability
If spare storage devices are used frequently to replace failed devices, then the storage array maintains reliability, but the spare devices experience excessive wear and their useful life is reduced
Solution Approach 1:
The patent recovers potentially failed storage devices through validation and reintegration into the storage array, reducing the frequency with which spare devices must be deployed and thereby extending the operational lifespan of spare storage resources
Solution Approach 2:
The patent enables potentially failed devices to serve themselves through automated validation and reintegration processes, reducing the need for continuous intervention from spare devices and optimizing the utilization lifecycle of spare storage resources
4Reliability
If storage devices with non-fatal errors are discarded immediately, then the storage array maintains high reliability, but expensive storage resources are lost and recovery opportunities are missed
Solution Approach 1:
The patent implements a comprehensive recovery process for storage devices with non-fatal errors, including data copying from spares, validation testing, conditional reintegration, and monitoring, thereby recovering expensive storage resources that would otherwise be prematurely discarded
Solution Approach 2:
The patent uses feedback from validation processes and error monitoring to determine whether potentially failed storage devices should be reintegrated into the array or permanently discarded, enabling data-driven decisions that balance reliability with resource conservation
Data Source
AI summary
Provided are a computer program product, system, and method for recovering storage devices in a storage array having errors. A determination is made to replace a first storage device in a storage array with a second storage device. The storage array is rebuilt by including the second storage device in the storage array and removing the first storage device from the storage array resulting in a rebuilt storage array. The first storage device is recovered from errors that resulted in the determination to replace. Data is copied from the second storage device included in the rebuilt storage array to the first storage device. The recovered first storage device is swapped into the storage array to replace the second storage device in response to copying the data from the second storage device to the first storage device.


