Metro Cluster Failover via Distributed Data Manager
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data storage systems face inefficiencies and disruptions during failover between different physical sites, as they require complex transfer of host data and site-specific configuration data, making failover time-consuming and disruptive.
Innovation Solution
Implementing a distributed data manager (DDM) in the IO stack of both sites to provide LUN virtualization, synchronous data mirroring, and cache coherency, allowing seamless failover by preserving virtual LUN IDs and transferring configuration and site-specific data within the virtualized environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional failover between different physical sites is implemented, then data storage system can resume operation after failure, but the process is time-consuming and disruptive due to complex transfer of host data and site-specific configuration data
Solution Approach 1:
The system performs preliminary actions by continuously maintaining synchronized copies of configuration data and host data at the remote site before failure occurs. The remote storage processor remains in a standby state with pre-loaded data, so when failure occurs, failover can occur immediately without waiting for data transfer.
Solution Approach 2:
The invention creates and maintains copies of configuration data and host data at the remote site through continuous data synchronization. These copies are kept identical to the primary site data, allowing the remote storage processor to assume operation immediately upon failure without requiring data transfer during failover.
2Reliability
If conventional failover between different physical sites is implemented, then data storage system can resume operation after failure, but the process is complex and disruptive
Solution Approach 1:
The distributed data manager acts as an intermediary that abstracts and simplifies the failover process. It automatically manages data synchronization, maintains data consistency between sites, and coordinates the failover procedure, eliminating the need for complex manual configuration transfers and reducing operational complexity.
Solution Approach 2:
The system implements a universal failover mechanism that handles both configuration data and host data through the same distributed data manager infrastructure. This multi-functional approach unifies the failover process, making it as simple as local SP failover despite occurring across different physical sites.
3Ease of operation
If LUN virtualization with preserved virtual LUN IDs is implemented across sites, then seamless failover is enabled, but data consistency and accessibility must be maintained across distributed systems
Solution Approach 1:
The distributed data manager implements feedback mechanisms that continuously monitor and verify data consistency between the primary and remote sites. It tracks data synchronization status and ensures that virtual LUN data remains consistent across sites, maintaining reliability while enabling seamless failover through preserved virtual LUN IDs.
Data Source
AI summary
A technique for supporting failover between SPs at different physical sites includes operating a distributed data manager (DDM) in an IO stack of both a first SP at a first site and a second SP at a second site. The DDMs of the first and second SPs cooperatively function to provide LUN virtualization that preserves virtual LUN IDs such that the first SP and the second SP can each access the same virtualized LUNs using the same virtual LUN IDs. In the event of a failure at the first site, the second SP at the second site may access the virtualized LUNs originally accessed by the first SP, including those storing configuration and site-specific data for the first site, as if those LUNs were local to the second SP.


