Deduplication Storage Retention via Hash-Based File Hierarchy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data protection systems face challenges in efficiently managing and retaining data snapshots across storage arrays, particularly in deduplication environments, where data redundancy and retention policies are complex and costly.
Innovation Solution
A method and apparatus for generating a protection file system in a deduplication storage array, which includes taking snapshots of production volumes, creating a file hierarchy with hashes of data, and adding retention indicators to each hash, allowing for efficient data retention and migration to less expensive retention storage when no longer needed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data replication is used to protect production site data, then data protection reliability is improved, but storage cost and system complexity increase
Solution Approach 1:
The patent creates a copy of production site data through snapshot generation, where the snapshot contains hashes of the original data. This copying mechanism provides data protection while using deduplication to minimize actual storage requirements, thus reducing both cost and complexity compared to traditional replication methods
Solution Approach 2:
The patent transforms data into hash values as a parameter change, allowing for efficient comparison and deduplication. By converting data to hash parameters, the system can identify duplicate content across snapshots and production data, reducing storage requirements while maintaining data integrity verification capabilities
2Reliability
If snapshots are stored in protection storage, then data recovery capability is improved, but storage cost increases
Solution Approach 1:
The patent converts snapshot data into hash values, transforming the data parameter from raw content to a compact representation. This parameter change enables deduplication where identical data blocks across different snapshots are recognized and stored only once, significantly reducing storage cost while maintaining the ability to recover any snapshot version
Solution Approach 2:
The system creates hash-based copies of snapshot data rather than storing full duplicates. By storing hashes and using deduplication, the system maintains multiple snapshot versions for recovery purposes while minimizing actual storage consumption through intelligent data redundancy elimination
3Quantity of substance
If deduplication is implemented in storage array, then storage efficiency is improved, but data retention management complexity increases
Solution Approach 1:
The patent segments data into discrete blocks and represents each block by its hash value. This segmentation approach allows the deduplication system to independently manage individual data blocks across multiple snapshots, making retention management more granular and controllable despite the increased complexity
4Reliability
If retention policies are enforced for all snapshot data, then data integrity is improved, but storage cost and processing overhead increase
Solution Approach 1:
The patent applies retention indicators selectively to specific hash values based on their reference counts and retention policy requirements. Rather than uniformly enforcing retention on all snapshot data, the system identifies which specific data blocks need retention protection and applies retention markers only to those, reducing unnecessary storage costs while maintaining data integrity for protected data
Data Source
AI summary
In one aspect, a method includes generating a protection file system in a deduplication storage array, generating a snapshot of a production volume in the deduplication storage array including hashes of data in the snapshot, generating a first file hierarchy for the hashes of the data in the snapshot in the protection file system and adding a retention indicator to each hash in the first file hierarchy.


