Cascading Snapshot Replication for 3-Site Data Protection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data protection systems require system shutdown during backups, limit recovery points in time, and are resource-intensive, leading to potential data loss and prolonged recovery times, especially with remote backup storage locations.
Innovation Solution
The asynchronous cascading replication process creates consistent snapshots across multiple storage clusters, allowing for efficient data transmission and replication without waiting for cycle completion, enabling immediate data transmission and marking of complete cycles, thus reducing overhead and enabling faster recovery.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is backed up to remote storage locations, then data protection coverage is improved, but replication time increases significantly
Solution Approach 1:
The patent segments the replication process into multiple independent target sites (first target site and second target site) that receive data in parallel from the source site. Each target site independently receives and processes replication data, eliminating the sequential bottleneck where data must traverse through one target to reach another. This segmentation directly addresses the time penalty of remote replication while maintaining comprehensive data protection coverage.
Solution Approach 2:
The patent creates empty containers at target sites in advance before data replication begins. This preliminary preparation eliminates setup delays during the actual replication process and allows data to be immediately written to pre-configured storage structures upon arrival, reducing overall replication time while ensuring data is protected across multiple remote locations.
2Reliability
If traditional backup systems are used, then data is stored periodically, but system shutdown is required during backup operations
Solution Approach 1:
The patent implements continuous data replication that operates without requiring system shutdowns. Data is replicated from the source site to multiple target sites in real-time or near-real-time as changes occur, allowing the production system to remain fully operational throughout the backup process. This eliminates the interruption to business operations while maintaining robust data protection.
3Reliability
If multiple backup storage systems are employed, then data protection coverage is improved, but resource consumption increases
Solution Approach 1:
The patent segments the replication workload across multiple independent target sites, where each site receives a portion of the replication data in parallel. This distributes the processing and storage resources required across the system, preventing any single site from becoming a resource bottleneck. The segmented approach achieves comprehensive data protection coverage while optimizing resource utilization across the distributed storage infrastructure.
4Stability of the object's composition
If sequential replication to remote sites is performed, then data consistency is maintained, but recovery time increases
Solution Approach 1:
The patent segments the replication architecture into multiple parallel target sites that simultaneously receive consistent replication data from the source. Each target site maintains data consistency independently through its own reception and processing of the same replication stream, eliminating the sequential chain that delays recovery. This segmented parallel architecture enables faster recovery by allowing any target site to serve as a recovery source without waiting for other sites in the sequence.
Data Source
AI summary
In one aspect a data replication process in a storage system includes creating, at a first target site, an empty container in a storage system. The empty container matches a container at a source site in response to initiation of an asynchronous data replication process. An aspect also includes transmitting a command to a second target site to create a container at the second target site. The first target site performs the asynchronous data replication process, which includes scanning the data upon receipt from the source site for a first target replication cycle and transmitting the scanned data to the container at the second target site for a second target replication cycle.


