Data Replication System for Multi-Region Large Dataset Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer storage file platforms do not scale to multiple regions, require small replication sizes, and are limited by latency issues when transferring and recovering large data sets across distant sites.
Innovation Solution
A system and method for replicating file portions across multiple remote geographical locations, allowing for rapid data transfer and recovery by dividing files into consistent-sized chunks, using a central replication system with geographically distributed storage systems that support file system replication, block replication, and accelerated traffic, enabling parallel data transfer and global rebuilds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional file systems with RAID are used, then data protection is achieved within single disk array environment, but the system cannot scale to multiple regions and replication is limited to small sizes
Solution Approach 1:
The patent segments data into discrete blocks that can be independently replicated across multiple geographic regions. Instead of treating data as a monolithic structure in traditional file systems, the system divides it into replicable units that can be distributed to multiple locations, enabling both protection and multi-region scalability simultaneously
Solution Approach 2:
The patent introduces a geographic dimension to data storage by enabling replication across multiple physical locations. This adds a spatial dimension to the traditional single-location storage model, allowing data to be protected through geographic distribution rather than just disk-level redundancy
2Reliability
If replication points are created in traditional file systems, then data can be replicated, but the replication size is limited and requires short distance between sites
Solution Approach 1:
The patent segments data into blocks that can be independently replicated, allowing large datasets to be replicated efficiently. Each block can be sent to multiple locations simultaneously, enabling large replication sizes without the latency penalties that would result from trying to replicate entire files or datasets as single units
Solution Approach 2:
Instead of replicating only small portions of data, the system enables replication of large datasets by processing them in manageable blocks. This partial action approach allows the system to handle large quantities of data while maintaining performance through parallel block replication
3Adaptability or versatility
If data is transferred across distant sites in traditional systems, then multi-region storage is achieved, but latency issues limit transfer speed
Solution Approach 1:
The patent segments data into blocks that can be sent in parallel across multiple geographic locations simultaneously. This segmentation enables the system to overcome latency by sending multiple small blocks in parallel rather than waiting for sequential transfer of large datasets, thereby maintaining high transfer speeds across distant sites
Solution Approach 2:
The system maintains continuous data transfer operations by parallelizing block replication across multiple locations. Instead of stopping to wait for latency-prone sequential transfers, the system continuously sends blocks to multiple destinations simultaneously, keeping the transfer pipeline full and maximizing throughput
4Reliability
If traditional RAID is used for data protection, then data can be rebuilt if a disk is lost, but the system cannot provide global rebuild capability across multiple regions
Solution Approach 1:
The patent segments data into blocks stored across multiple geographic locations, enabling global rebuild capability. When a disk fails, the system can reconstruct data by retrieving blocks from multiple different geographic locations rather than relying on a single local RAID array, providing both protection and global scalability
Solution Approach 2:
The system provides universal data recovery capability by enabling rebuild operations across multiple geographic locations. Any block can be recovered from any location where it is replicated, making the recovery process universal rather than location-specific, thus providing both local and global rebuild capability
Data Source
AI summary
Systems for rapidly transferring and, as needed, recovering large data sets and methods for making and using the same. In various embodiments, the system advantageously can allow data to be transferred in larger sizes, wherein data may be easily recovered from multiple regions and wherein latency is no longer an issue, among other things.


