Data Archiving via Aggregated Objects in Distributed Storage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data archiving from replicated networked distributed storage systems (NDSS) to erasure coded NDSS is sequential and does not effectively utilize the replicated nature of data chunks, leading to inefficiencies in storage space and recovery processes.
Innovation Solution
The proposed method involves aggregating co-located data chunks into data objects, which are then written in parallel to erasure coded storage nodes, maintaining metadata mapping to ensure correct reassembly and utilizing dispersed object stores as a secondary tier for primary enterprise block storage, optimizing data archiving by leveraging the replicated nature and reducing storage overhead.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If sequential data archiving is used from replicated NDSS to erasure coded NDSS, then data can be archived, but storage space is not optimized and recovery processes are inefficient
Solution Approach 1:
The patent segments data into chunks and further divides them into sub-chunks, organizing them into data objects with specific structures. This segmentation enables parallel processing during archiving operations, significantly improving productivity while optimizing storage space through efficient data organization and deduplication.
Solution Approach 2:
The patent merges multiple data chunks into consolidated data objects, combining original data chunks with their replicas. This merging reduces redundant storage and optimizes space utilization while maintaining data integrity and enabling faster archiving through batch operations.
2Reliability
If replicated data chunks are stored separately in different storage nodes, then data reliability is improved, but storage overhead increases
Solution Approach 1:
The patent creates data objects that include both original data chunks and their replicas, effectively copying data for redundancy. However, it optimizes storage by organizing these copies in a structured manner that reduces overall overhead compared to traditional replication methods, while maintaining data reliability through the replicated structure.
Solution Approach 2:
The patent creates composite data objects that combine original data chunks with their replicas in a unified structure. This composite approach allows for efficient storage management, reducing overhead while maintaining the reliability benefits of replication through the integrated object structure.
3Productivity
If conventional sequential archiving is used, then implementation is simple, but archiving efficiency is low
Solution Approach 1:
The patent introduces dynamic parallel processing capabilities to the archiving system, allowing multiple data objects to be processed simultaneously. This dynamic approach significantly improves archiving efficiency while managing complexity through structured data organization and metadata management that coordinates the parallel operations.
Solution Approach 2:
The patent introduces data objects as intermediary structures between source data and destination storage. These objects serve as mediators that organize and manage the archiving process, enabling parallel operations while maintaining system coherence. The intermediary structure manages complexity by providing a standardized interface for parallel processing.
Data Source
AI summary
A set of data chunks stored in a first data storage system is accessed. The set of data chunks includes original data chunks and replicated data chunks respectively corresponding to the original data chunks. A given original data chunk and the corresponding replicated data chunk are stored in separate storage nodes of the first data storage system. For each of at least a subset of storage nodes of the first data storage system, unique ones of the original data chunks and the replicated data chunks stored on the storage node are aggregated to form a data object. The data objects thereby formed collectively represent a given data volume. Each of the data objects is stored in separate storage nodes of a second data storage system.


