Tape Storage Deduplication via Clique Segregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional tape storage systems face challenges in efficient data reliability and storage cost due to limitations in duty cycles and the inefficiency of traditional RAID and erasure coding methods, which require loading multiple tapes and result in high storage costs and latency issues.
Innovation Solution
Implementing a data storage system that employs restricted-deduplication assisted replication by segregating data into cliques, selecting compatible tape plexes for deduplication, and writing replicas across multiple tape plexes to balance read/write operations and reduce latency, while maintaining data reliability and storage efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional RAID or erasure coding methods are used for data reliability on tape storage systems, then data reliability is improved, but storage costs increase and latency increases due to requiring multiple tape loads
Solution Approach 1:
The patent divides data into segments or chunks and distributes them across multiple tape drives simultaneously. This segmentation allows parallel processing of data reliability operations without requiring sequential tape loads, thereby reducing latency while maintaining data reliability through distributed storage across multiple segments
Solution Approach 2:
The system performs preliminary deduplication and data preparation operations before writing to tape drives. By pre-processing data to identify and eliminate duplicates, the system reduces the actual data volume that needs to be written and stored, which decreases the time required for subsequent read operations and reduces latency for data retrieval
2Reliability
If traditional RAID or erasure coding methods are used for data reliability on tape storage systems, then data reliability is improved, but storage costs increase
Solution Approach 1:
The patent extracts and removes duplicate data segments before writing to tape storage. By identifying and eliminating redundant data copies through deduplication operations, the system reduces the total volume of data that needs to be stored, thereby lowering storage costs while maintaining data reliability through selective retention of unique data segments
Solution Approach 2:
The system changes the parameter of data representation by storing deduplicated references instead of complete duplicate copies. This parameter change from storing full data segments to storing references or pointers to existing segments significantly reduces storage requirements while maintaining data integrity and reliability through the reference structure
3Quantity of substance
If tape drives are used for sequential storage to reduce costs, then storage efficiency is improved, but throughput decreases and latency increases due to slow seek times
Solution Approach 1:
The patent segments data into multiple chunks that can be written to and read from different tape drives simultaneously. This segmentation enables parallel I/O operations, where multiple tape drives work concurrently on different data segments, thereby increasing overall system throughput despite the inherently sequential nature of individual tape drive operations
Solution Approach 2:
The system transitions from single-dimensional sequential access to multi-dimensional parallel access by distributing data across multiple tape drives. This dimensional change allows the system to exploit parallelism in the storage architecture, where data can be accessed from multiple tape drives at the same time, effectively increasing throughput beyond what a single sequential tape drive could achieve
Data Source
AI summary
A method and system for deduplicating data for a data storage system using similarity determinations are described. A tape library is arranged in a hierarchy of tape groups and tape plexes. Tape groups are an admin visible entity and are comprised of multiple tape plexes (at least equal to the number of replicas in a tape group). Tape plexes in turn comprise multiple tape cartridges. Data files and objects received within a time period are initially staged in a disk cache where they are logically segregated into cliques based on their expected deduplication ratios. These cliques are then evaluated for the amount of duplication they have with data existing in tape plexes. Based on the number of replicas being written, the top few tape plexes are selected from within the tape group. The cliques are deduplicated with data on the selected tape plexes, compressed, and written to tape.


