Storage Control Method for Disk Redundancy Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional disk array subsystems face issues with data redundancy loss during recovery processing, where the separation of a failed disk drive occurs before data recovery is complete, leading to a loss of redundancy and increased risk of data loss.
Innovation Solution
A storage control method and system that maintain redundancy by allowing data recovery from a suspect disk drive to a spare disk drive under the same device adapter, with data copying and rebuilding processes managed to ensure continuous operation and minimal redundancy loss.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data recovery is executed from a failed disk drive to a spare disk drive, then data redundancy is maintained, but the separation of the failed disk drive occurs before data recovery completion, causing loss of redundancy
Solution Approach 1:
The system performs preliminary actions by completing data recovery to the spare disk drive before separating the failed disk drive. The control unit waits for data recovery completion and only then executes the separation process, ensuring redundancy is maintained throughout the entire recovery process and eliminating the redundancy loss period that occurs when separation happens prematurely.
2Productivity
If the failed disk drive is separated immediately, then replacement and repair can begin, but data recovery is not complete and redundancy is lost
Solution Approach 1:
The system executes data recovery as a preliminary action before disk separation. The control unit completes the data transfer from the failed disk to the spare disk, ensuring all data is safely recovered, and only then proceeds to separate the failed disk. This eliminates the risk of data loss while maintaining the ability to replace and repair the failed disk.
3Reliability
If data recovery is executed while maintaining redundancy, then data safety is improved, but the complexity of the recovery process increases
Solution Approach 1:
The control unit acts as an intermediary that manages the recovery process between the failed disk drive and the spare disk drive. It coordinates data transfer operations, monitors recovery progress, and executes separation only when recovery is complete. This intermediary management simplifies the overall process by providing centralized control and eliminating the need for complex coordination between multiple components.
Data Source
AI summary
In case an error statistics of one of the disk drives exceeds a predetermined threshold, the disk is determined as a suspect disk drive. A recovery mode is set successively. During the time when a setting of the recovery mode is in progress and no access is made from a host 16 in this time, the address range of the suspect disk drive is specified. At the same time, a processing is started in that the data of the suspect disk is copied to a spare disk 34 sequentially to recover the data. The data of the suspect disk drive is copied to the spare disk drive 34 to recover the data when the address range of the suspect disk drive does not correspond to the write failure address range of a management table 48. The data of a normal disk drive is copied to the spare disk drive 34 to recover the data when the address range of the suspect disk drive corresponds to the write failure address range of the management table 48. Upon the completion of the recovery of the data, the suspect disk drive 32 is separated and replaced with the spare disk drive 34.


