Metadata Replication for Storage Migration Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data migration between geographically distributed storage systems is inefficient due to high network throughput requirements, leading to long migration processes and interference with new data replication.
Innovation Solution
The technology employs metadata replication, specifically virtual data chunks, to minimize data transfer between storage systems, allowing newer storage systems to convert virtual chunks into real data chunks by reading from locally coupled legacy systems, thus reducing actual data replication and ensuring data consistency through checksum verification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If actual data copying is performed between geographically distributed storage systems, then data migration completeness is improved, but network throughput requirements increase and migration time extends
Solution Approach 1:
The patent extracts only the essential metadata (virtual data chunks) from the full data set for replication between geographically distributed storage systems. By replicating only virtual chunk metadata rather than actual data, the system achieves data migration completeness through subsequent local data transfer from legacy systems, while dramatically reducing network throughput requirements and migration time.
Solution Approach 2:
The patent performs preliminary replication of virtual data chunk metadata to the target storage system before actual data migration. This preliminary action establishes the data structure and metadata in advance, allowing the target system to then efficiently import actual data from local legacy systems without requiring extensive network data transfer during the main migration process.
2Reliability
If actual data copying is performed between remote storage systems, then data consistency is improved, but network throughput consumption increases
Solution Approach 1:
The patent extracts and replicates only virtual data chunk metadata between remote storage systems, eliminating the need to transfer actual data over the network. This ensures data consistency is maintained through metadata synchronization while dramatically reducing network throughput consumption compared to traditional full data copying approaches.
Solution Approach 2:
The patent introduces virtual data chunks as an intermediary representation that mediates between source and target storage systems. Instead of directly copying actual data across the network, the system uses virtual chunk metadata as an intermediary that can be replicated efficiently, with actual data then imported locally from legacy systems to the target storage system.
3Reliability
If extensive data replication is performed between storage systems, then data protection is improved, but replication traffic increases
Solution Approach 1:
The patent extracts only the critical virtual data chunk metadata for replication between storage systems, maintaining data protection through metadata synchronization and checksum verification. This approach ensures data integrity and protection while minimizing replication traffic by excluding actual data from the replication process.
Solution Approach 2:
The patent changes the replication parameter from copying actual data to copying virtual data chunk metadata. This parameter change maintains the essential function of data protection and consistency verification through checksums, while dramatically reducing the volume of replication traffic between storage systems.
Data Source
AI summary
The described technology is generally directed towards replicating metadata representing a virtual data structure corresponding to replicated legacy data instead of the actual data for the data structure. Once virtual chunks are replicated to a remote, newer storage system, the corresponding legacy data is locally read into the virtual chunks to transform the virtual chunks into real data chunks of the remote newer storage system. A checksum can be replicated for the remote newer storage system to evaluate the consistency of the data. Efficient data storage migration is thus accomplished in a replicated environment based on relatively negligible replication traffic between two remote locations, while still assuring the consistency of migrated data.


