Multi-site Storage Snapshot Retention and Failover Reconfiguration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-site distributed storage systems face challenges in maintaining common snapshots between secondary and tertiary storage sites during a primary storage site failure, leading to time-consuming and bandwidth-intensive baseline data transfers for resumption of protection configurations.
Innovation Solution
Implementing an asynchronous mirroring policy to ensure common snapshots are maintained between secondary and tertiary storage sites, enabling automatic unplanned failover and realignment of protection configurations without baseline data transfers, and automatically reconfiguring replication relationships post-failure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If baseline data transfer is performed between secondary and tertiary storage sites after primary storage failure, then data consistency is restored, but network bandwidth consumption increases and recovery time extends
Solution Approach 1:
The system performs preliminary actions by maintaining asynchronous replication relationships and snapshot copies between secondary and tertiary storage sites before the primary storage failure occurs. This preliminary setup ensures that when failover happens, data can be restored without requiring extensive baseline transfers, thus reducing network bandwidth consumption during recovery.
Solution Approach 2:
The patent utilizes snapshot copies as lightweight replicas of the original data. Instead of transferring complete baseline data between storage sites, the system creates and maintains snapshot copies that can be quickly activated during failover, significantly reducing the network bandwidth required for data consistency restoration.
2Loss of time
If manual user intervention is used to restore operations after storage site failure, then system complexity is reduced, but recovery time increases
Solution Approach 1:
The system implements self-service capabilities through automatic failover mechanisms. When the primary storage site fails, the secondary storage site automatically takes over without requiring manual user intervention. The system also automatically initiates realignment and reconfiguration of protection configurations, enabling rapid recovery while maintaining high automation levels.
Solution Approach 2:
The system employs feedback mechanisms to monitor the health status of storage sites and automatically trigger failover procedures when failures are detected. This closed-loop control enables the system to respond to failures autonomously, reducing recovery time without sacrificing automation.
3Reliability
If synchronous replication is used from primary to secondary storage site, then data availability is improved, but network bandwidth consumption increases
Solution Approach 1:
The system applies different replication strategies to different data paths based on local requirements. Synchronous replication is used from primary to secondary storage site to ensure data availability, while asynchronous replication is used from primary to tertiary storage site to reduce network bandwidth consumption. This localized optimization allows the system to achieve reliability where needed while minimizing bandwidth usage elsewhere.
Solution Approach 2:
The system dynamically changes replication parameters based on operational conditions. The replication mode (synchronous or asynchronous) and update schedules are adjusted according to the specific data path and requirements, allowing optimization of both data availability and network bandwidth consumption for different replication relationships.
4Productivity
If protection configuration realignment is performed automatically after failover, then operational continuity is maintained, but system complexity increases
Solution Approach 1:
The system implements a universal realignment mechanism that automatically handles multiple configuration tasks after failover. The same automated process manages both the activation of protection configurations and the reconfiguration of replication relationships, reducing the need for separate manual intervention steps and maintaining operational continuity despite the inherent complexity.
Data Source
AI summary
Multi-site distributed storage systems and computer-implemented methods are described for providing common snapshot retention and automatic fanout reconfiguration for an asynchronous leg after a failure event that causes a failover from a primary storage site to a secondary storage site. A computer-implemented method comprises providing an asynchronous replication relationship with an asynchronous update schedule from one or more storage objects of the first storage node to one or more replicated storage objects of the third storage node, creating a snapshot copy of the one or more storage objects of the first storage node, transferring the snapshot copy to the third storage node based on an asynchronous mirror policy, and intercepting the snapshot create operation on the primary storage site and transferring the snapshot copy to the second storage node to provide a common snapshot between the second storage node and the third storage node.


