Cross-Zone Block Storage Failover with Synchronous Replication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing network-based block storage devices are susceptible to wide-scale failures such as power outages or natural disasters due to lack of off-site replication, which can result in data loss and service interruptions, as traditional off-site backups are not immediately usable and may take hours or days to restore.
Innovation Solution
Implementing cross-zone replication across isolated availability zones, where data is synchronously replicated across multiple computing systems, allowing each volume to be independently functional and up-to-date, reducing the risk of data loss and performance impact.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional off-site backups are used for block storage devices, then data protection against wide-scale failures is improved, but service restoration time increases significantly (hours or days)
Solution Approach 1:
The system performs preliminary actions by continuously maintaining synchronized secondary volumes in advance before failures occur. These pre-positioned secondary volumes are kept ready and can immediately assume the primary role when a failure happens, eliminating the need for time-consuming restoration processes associated with traditional backups.
Solution Approach 2:
The invention creates and maintains exact copies (secondary volumes) of the primary block storage device in different availability zones. These copies are continuously synchronized and can immediately replace the primary volume upon failure, providing both data protection and rapid service restoration without the delays inherent in traditional backup restoration.
2Reliability
If cross-zone replication is implemented, then service availability and resilience to wide-scale failures are improved, but system complexity increases
Solution Approach 1:
The system segments the block storage device into multiple independent volumes distributed across different availability zones. Each volume can operate independently, and the failure of one zone does not affect others. This segmentation provides resilience while managing complexity through modular, zone-isolated architecture.
Solution Approach 2:
The system introduces an intermediary mechanism (the secondary volume in a different availability zone) that mediates between the primary volume and potential failures. This intermediary maintains synchronization and can assume the primary role, simplifying the failover process while improving availability across zone boundaries.
3Reliability
If synchronous replication across multiple zones is performed, then data integrity and immediate usability are improved, but network bandwidth consumption increases
Solution Approach 1:
The system applies local quality by optimizing replication behavior based on the specific characteristics of each availability zone and workload. Synchronous replication is performed selectively for critical data operations, while leveraging local storage capabilities within zones to reduce unnecessary network traffic while maintaining data integrity.
Data Source
AI summary
The present disclosure generally relates to creating virtualized block storage devices whose data is replicated across isolated computing systems to lower risk of data loss even in wide-scale events, such as natural disasters. The virtualized device can include at least two volumes, each of which is implemented in a distinct computing system. In the case of a failed volume, a new volume can be created and populated with data from the surviving volume. During population, new writes can continue to be replicated to the new volume. The population process can write data from the surviving volume to the new volume “under” new writes, such that the population process does not overwrite data included in the new writes.


