Distributed Storage Volume Recovery During Site Failure
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current distributed storage systems face challenges in automatically creating and maintaining distributed storage volumes during site failures, array failures, or inter-site network failures, which requires manual administrator intervention and is error-prone.
Innovation Solution
A method and system that automatically detect array unavailability, create a local part of the distributed volume, export it to an available array, continue processing I/O requests, and upon array reavailability, create a remote part and perform consistency processing to ensure compliance with distributed storage requirements, using a RAID system to maintain data consistency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If distributed storage volumes are not automatically created during site failure, array failure or inter-site network failure, then manual administrator intervention is required which is error-prone, but the system cannot maintain automated operation and data availability
Solution Approach 1:
The system automatically detects array unavailability and creates local parts of distributed volumes without administrator intervention. The storage system self-manages the entire workflow including detecting failures, creating local volume parts, exporting to available arrays, and performing consistency processing when arrays return.
Solution Approach 2:
The system performs preliminary actions by automatically creating local parts of distributed volumes on available arrays before the failed arrays return. This preliminary creation allows I/O processing to continue immediately rather than waiting for manual intervention or array recovery.
2Ease of operation
If manual workflow steps are required to create volumes on available arrays during failures, then the process is complex and error-prone, but automation is not implemented
Solution Approach 1:
Multiple manual operations are merged into a single automated process. The system combines failure detection, volume creation, data export, I/O processing continuation, and consistency management into one unified automated workflow that executes without administrator intervention.
Solution Approach 2:
The complex manual workflow steps are extracted from the administrator's responsibilities and transferred to the storage system's automated processes. The system independently handles detecting array unavailability, creating local volume parts, managing exports, and performing consistency processing.
3Reliability
If distributed storage processing stops during array unavailability, then data consistency is maintained, but I/O requests cannot be processed
Solution Approach 1:
The distributed volume is segmented into local and remote parts. When remote arrays are unavailable, the system creates and uses local volume parts that can independently process I/O requests. This segmentation allows continued operation with local data while maintaining the ability to restore full distributed functionality when arrays return.
Solution Approach 2:
Local volume parts act as intermediaries between I/O requests and the unavailable remote arrays. These local parts enable continued I/O processing by providing local storage access, while consistency processing later reconciles any differences when remote arrays become available again.
4Reliability
If consistency processing is not performed automatically when arrays return, then compliance with distributed storage requirements may be violated, but manual intervention is required
Solution Approach 1:
The system implements feedback by automatically detecting when previously unavailable arrays return to availability. Upon detecting array reavailability, the system triggers consistency processing to compare and reconcile data between local and remote volume parts, ensuring compliance with distributed storage requirements.
Solution Approach 2:
The system performs preliminary consistency processing automatically when arrays return, before requiring any manual intervention. This ensures that data compliance is verified and restored immediately upon array recovery, maintaining the integrity of the distributed storage system.
Data Source
AI summary
A system and method are provided for processing to create distributed volume in a distributed storage system during a failure that has partitioned the distributed volume (e.g. an array failure, a site failure and/or an inter-site network failure). In an embodiment, the system described herein may provide for continuing distributed storage processing in response to I/O requests from a source by creating the local parts of the distributed storage during the failure, and, when the remote site or inter-site network return to availability, the remaining part of the distributed volume is automatically created. The system may include an automatic rebuild to make sure that all parts of the distributed volume are consistent again. The processing may be transparent to the source of the I/O requests.


