Flash Copy for Disaster Recovery Testing Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems for disaster recovery (DR) testing face challenges in maintaining data consistency across multiple clusters, as existing flash copy solutions are limited to single-node consistency, and cannot create a composite consistency point in time, leading to unpredictable replication states and misleading results during DR testing.
Innovation Solution
The system accesses a snapshot of data on DR clusters only when it is consistent with production clusters before a designated time-zero, allowing multiple hosts to mount virtual tapes with the same identifier, and manages snapshot ownership independently of live virtual tape ownership, enabling DR testing with accurate, time-zero consistent data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional flash copy solutions are used in grid architecture, then single-node consistency can be achieved, but composite consistency across multiple clusters cannot be created
Solution Approach 1:
The system segments the grid architecture into multiple DR families, where each DR family represents a group of clusters that can be independently flashed to a consistent state. This segmentation allows flash operations to be performed on specific clusters or sites within the grid, enabling composite consistency across multiple clusters while maintaining the ability to handle large amounts of data through distributed flash operations
Solution Approach 2:
The patent introduces DR families as an intermediary layer between individual clusters and the flash copy operation. A DR family acts as a logical container that groups clusters/sites, allowing the system to manage consistency across multiple clusters through a unified interface while performing distributed flash operations on the underlying clusters
2Productivity
If copies continue after DR testing starts, then data replication remains ongoing, but misleading results are obtained because data would not be available in a real disaster
Solution Approach 1:
The system performs a flash operation before DR testing begins, creating a point-in-time consistent snapshot of all data across the DR clusters. This preliminary action freezes the data state at a specific moment, ensuring that the testing environment reflects what would be available in a real disaster scenario while allowing production replication to continue uninterrupted
3Ease of operation
If virtual tape ownership is protected with one host at a time, then data access control is maintained, but DR testing with multiple hosts cannot be performed
Solution Approach 1:
The patent segments virtual tape ownership by creating distinct ownership contexts for production and DR environments. Production hosts maintain ownership of live virtual tapes, while DR hosts gain ownership of flashed snapshots through the DR family mechanism. This segmentation allows multiple hosts to simultaneously access different representations of the same data without conflict
Solution Approach 2:
The system creates copied representations of virtual tapes through flash operations. Instead of allowing direct access to the same virtual tape by multiple hosts, the system creates snapshot copies that can be independently mounted by DR hosts. These copies preserve the data state at a specific point in time while enabling concurrent access by multiple testing hosts
4Productivity
If production data continues to replicate to DR clusters, then backup readiness is maintained, but data consistency at time-zero cannot be ensured
Solution Approach 1:
The system performs a flash operation as a preliminary action before DR testing to capture a consistent snapshot of all data across DR clusters at time-zero. This flash operation freezes the data state across all clusters in the DR family, ensuring consistency while allowing production replication to continue afterward without affecting the tested snapshot
Data Source
AI summary
In one embodiment, a system includes a processor and logic integrated with and/or executable by the processor, the logic being configured to cause the processor to access a snapshot of data stored on one or more DR clusters within a disaster recovery (DR) family only when the snapshot was made consistent with respect to data on one or more production clusters within the DR family before a time-zero, and perform DR testing using the snapshot, wherein the time-zero represents a time selected to simulate a disaster, wherein the DR family includes one or more DR clusters accessible to a DR host and the one or more production clusters accessible to a production host, and wherein the DR host is configured to replicate data from the one or more production clusters to the one or more DR clusters.


