Disk Array Controller Error Counting and Warning Mode
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing disk array systems with redundancy face inefficiencies in data restoration due to high error rates in disk drives, leading to prolonged processing times and potential failure acceleration when a disk drive is recognized as likely to fail, which can disrupt data rebuilding and access operations.
Innovation Solution
An array controller with error counters, a failure estimation unit, and a mode-setting unit that sets a disk drive in a warning mode to reduce access frequency and maintain it as a member of the array, making it less accessible than other drives, thereby elongating its lifespan and preventing premature failure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a disk drive recognized as likely to fail is accessed in the same manner as normal disk drives during rebuild process, then data restoration can be performed, but the time when the disk drive actually fails is accelerated
Solution Approach 1:
The patent applies local quality by treating the recognized-at-risk disk drive differently from normal disk drives. Specifically, the controller restricts access to the at-risk drive only to rebuild operations while allowing normal drives to be accessed freely for both rebuild and host I/O operations. This differentiated access policy preserves the at-risk drive longer while still enabling necessary data restoration.
Solution Approach 2:
The patent implements preliminary action by proactively identifying disk drives at risk of failure through error counting before actual failure occurs. The controller monitors error rates and preemptively restricts access to drives exceeding error thresholds, preventing further degradation and potential failure while maintaining system functionality through alternative access paths.
2Productivity
If a disk drive recognized as likely to fail is accessed frequently, then data restoration process can proceed, but the disk drive may fail before rebuild is completed
Solution Approach 1:
The controller proactively identifies drives at risk through error monitoring before failure occurs, enabling preemptive protection measures. By restricting access to at-risk drives only for rebuild operations and preventing other access patterns, the system ensures rebuild completion while avoiding premature drive failure that would interrupt the process.
Solution Approach 2:
The patent uses the controller as an intermediary that manages access to the at-risk disk drive. The controller acts as a gatekeeper, allowing only necessary rebuild operations to access the vulnerable drive while blocking other access patterns that could accelerate failure. This intermediary control enables safe completion of rebuild processes.
3Reliability
If a disk drive is made less accessible to elongate its lifespan, then premature failure is prevented, but data access speed may be reduced
Solution Approach 1:
The patent applies local quality by implementing selective access restrictions only on identified at-risk disk drives while leaving normal drives fully accessible. The controller differentiates between safe and unsafe drives, applying access limitations only where necessary. This targeted approach preserves overall system performance while protecting vulnerable components.
Solution Approach 2:
The system uses redundancy and copying mechanisms to mitigate the impact of restricted access. Data from at-risk drives is restored to safe drives or spare drives, creating copies that can serve future access requests. This copying approach maintains data availability and access speed while protecting the original at-risk drive from further stress.
Data Source
AI summary
According to one embodiment, a read/write control unit controls read/write access to at least two disk drives that provide a disk array. Error counters are provided for the respective disk drives for counting respective numbers of errors if the errors occur when the disk drives are accessed. A failure estimation unit detects, as a disk drive which is very likely to fail, a disk drive included in the disk array and having a high error occurrence degree, based on the numbers of errors counted by the error counters. A mode-setting unit sets the detected disk drive in a particular mode in which the detected disk drive is maintained as a member of the disk array and is made more inaccessible than the remaining disk drive of the disk array.


