Deduplication Storage Retention via Hash-Based File Hierarchy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data protection systems face challenges in efficiently managing and retaining data snapshots across storage arrays, particularly in deduplication environments, where data redundancy and retention policies are complex and costly.

Innovation Solution

A method and apparatus for generating a protection file system in a deduplication storage array, which includes taking snapshots of production volumes, creating a file hierarchy with hashes of data, and adding retention indicators to each hash, allowing for efficient data retention and migration to less expensive retention storage when no longer needed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data replication is used to protect production site data, then data protection reliability is improved, but storage cost and system complexity increase

Engineering Contradiction:
Improvedata protection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent creates a copy of production site data through snapshot generation, where the snapshot contains hashes of the original data. This copying mechanism provides data protection while using deduplication to minimize actual storage requirements, thus reducing both cost and complexity compared to traditional replication methods

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms data into hash values as a parameter change, allowing for efficient comparison and deduplication. By converting data to hash parameters, the system can identify duplicate content across snapshots and production data, reducing storage requirements while maintaining data integrity verification capabilities

Inventive Principle:
Principle #35Parameter changes

2Reliability

If snapshots are stored in protection storage, then data recovery capability is improved, but storage cost increases

Engineering Contradiction:
Improvedata recovery capabilityVSAvoidstorage cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent converts snapshot data into hash values, transforming the data parameter from raw content to a compact representation. This parameter change enables deduplication where identical data blocks across different snapshots are recognized and stored only once, significantly reducing storage cost while maintaining the ability to recover any snapshot version

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system creates hash-based copies of snapshot data rather than storing full duplicates. By storing hashes and using deduplication, the system maintains multiple snapshot versions for recovery purposes while minimizing actual storage consumption through intelligent data redundancy elimination

Inventive Principle:
Principle #26Copying

3Quantity of substance

If deduplication is implemented in storage array, then storage efficiency is improved, but data retention management complexity increases

Engineering Contradiction:
Improvestorage efficiencyVSAvoiddata retention management complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments data into discrete blocks and represents each block by its hash value. This segmentation approach allows the deduplication system to independently manage individual data blocks across multiple snapshots, making retention management more granular and controllable despite the increased complexity

Inventive Principle:
Principle #1Segmentation

4Reliability

If retention policies are enforced for all snapshot data, then data integrity is improved, but storage cost and processing overhead increase

Engineering Contradiction:
Improvedata integrityVSAvoidstorage cost
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies retention indicators selectively to specific hash values based on their reference counts and retention policy requirements. Rather than uniformly enforcing retention on all snapshot data, the system identifies which specific data blocks need retention protection and applies retention markers only to those, reducing unnecessary storage costs while maintaining data integrity for protected data

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10108356B1Determining data to store in retention storage
Publication Date: 2018.10.23 EMC IP HLDG CO LLC
  • US10108356B1 patent drawing
  • US10108356B1 patent drawing
  • US10108356B1 patent drawing

AI summary

In one aspect, a method includes generating a protection file system in a deduplication storage array, generating a snapshot of a production volume in the deduplication storage array including hashes of data in the snapshot, generating a first file hierarchy for the hashes of the data in the snapshot in the protection file system and adding a retention indicator to each hash in the first file hierarchy.