Disk-Level Cloning for Fast Mirror Resynchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for re-establishing mirror protection in high-availability storage systems after a permanent disk failure are time-consuming, typically requiring a level-0 resync that can take days or weeks to complete.
Innovation Solution
The use of disk-level cloning to quickly re-establish mirror protection by attaching cloned disks to the failed node and performing a level-1 resync based on a base file system snapshot, significantly reducing the time required for resynchronization to minutes or hours.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a level-0 resync is performed to re-establish mirror protection after disk failure, then mirror protection is restored, but the resynchronization process takes days or weeks to complete
Solution Approach 1:
The patent applies preliminary action by creating and maintaining a base file system snapshot before disk failure occurs. When a disk fails, this pre-existing snapshot enables a level-1 resync instead of requiring a level-0 resync, significantly reducing recovery time from days/weeks to minutes/hours while still restoring mirror protection
Solution Approach 2:
The patent uses copying by creating a disk-level clone from the base file system snapshot. This clone serves as a replacement for the failed disk, allowing the mirror protection to be re-established through a faster level-1 resync process rather than the slower level-0 resync, thus resolving the time contradiction
2Loss of time
If disk-level cloning is used to re-establish mirror protection, then resynchronization time is reduced to minutes or hours, but the process requires additional cloning operations
Solution Approach 1:
The system performs self-service by automatically detecting disk failures and initiating the cloning and resync process without manual intervention. The base file system snapshot is automatically utilized to create the necessary disk-level clones, reducing operational complexity despite the advanced cloning operations involved
Solution Approach 2:
The base file system snapshot acts as an intermediary that enables the fast recovery process. This pre-existing snapshot serves as the source for creating disk-level clones, mediating between the failed disk and the recovery process, thereby enabling reduced resynchronization time while managing complexity through a structured intermediate step
Data Source
AI summary
Systems and methods for performing a fast resynchronization of a mirrored aggregate of a distributed storage system using disk-level cloning are provided. According to one embodiment, responsive to a failure of a disk of a plex of the mirrored aggregate utilized by a high-availability (HA) pair of nodes of a distributed storage system, disk-level clones of the disks of a healthy plex may be created external to the distributed storage system and attached to the degraded HA partner node. After detection of the cloned disks by the degraded HA partner node, mirror protection may be efficiently re-established by assimilating the cloned disks within the failed plex and then resynchronizing the mirrored aggregate by performing a level-1 resync of the failed plex with the healthy plex based on a base file system snapshot of the healthy plex. In this manner, a more time-consuming level-0 resync may be avoided.


