Lightweight Deduplication via Extent and Content Indices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deduplication processes are computationally intensive and inefficient on resource-constrained devices or when dealing with large datasets, and they fail to minimize data transfer effectively over Wide Area Networks (WANs).

Innovation Solution

A lightweight deduplication system that uses a combination of extent and content indices to identify and eliminate duplicate data segments, employing full and shortened fingerprints for aligned and non-aligned segments, and caching techniques to optimize performance on resource-limited devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional deduplication processes are used on resource-constrained devices, then data duplication elimination can be achieved, but computational overhead becomes excessive and performance deteriorates

Engineering Contradiction:
Improvededuplication effectivenessVSAvoidcomputational overhead
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent segments data into fixed-size blocks and processes them independently. Each block is assigned a unique identifier (fingerprint) that can be compared without processing the entire data set. This segmentation allows resource-constrained devices to handle deduplication in manageable chunks rather than as a monolithic computational task.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses fingerprint copies (hash values) of data blocks instead of comparing the actual data blocks directly. These fingerprint copies are stored in index structures that can be efficiently searched. By working with compact fingerprint representations rather than full data blocks, the system dramatically reduces memory and computational requirements while maintaining deduplication effectiveness.

Inventive Principle:
Principle #26Copying

2Measurement precision

If full fingerprints are used for all segments, then deduplication accuracy is maintained, but memory consumption and processing time increase significantly

Engineering Contradiction:
Improvededuplication accuracyVSAvoidmemory consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies different fingerprint lengths to different segments based on their alignment status. Aligned segments (those at the same position in different data sets) use shorter fingerprints for quick comparison, while unaligned segments use full fingerprints for accurate matching. This local differentiation optimizes the balance between accuracy and resource usage.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses a two-stage approach where shorter fingerprints are used for initial screening and partial matching, and full fingerprints are applied only when needed for final verification. This partial application of full fingerprints reduces overall memory consumption while maintaining deduplication accuracy where it matters most.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If data is replicated over WAN without optimization, then data redundancy is created, but bandwidth utilization suffers and transfer time increases

Engineering Contradiction:
Improvedata replication integrityVSAvoidbandwidth utilization
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent performs deduplication operations before data replication over the WAN. By identifying and eliminating duplicate blocks in advance using fingerprint comparison, the system ensures that only unique data blocks are transmitted across the network. This preliminary deduplication action significantly reduces the volume of data requiring WAN transmission.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces physical data comparison (mechanical system) with cryptographic fingerprint comparison. Instead of transmitting and comparing full data blocks over the WAN, the system transmits compact fingerprints that can be quickly compared to identify duplicates. This substitution dramatically reduces network bandwidth requirements while maintaining replication integrity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Loss of substance

If resource-limited devices perform deduplication on large datasets, then storage space can be saved, but processing speed and efficiency deteriorate

Engineering Contradiction:
Improvestorage space reductionVSAvoidprocessing efficiency
Core Design Contradiction:
Loss of substanceVSProductivity

Solution Approach 1:

The patent divides large data sets into fixed-size blocks that can be processed independently and in parallel. This segmentation allows resource-limited devices to process data in manageable units, improving throughput and enabling better utilization of limited computational resources while maintaining overall processing efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces resource-intensive full data block comparisons with efficient fingerprint hash computations. These fingerprint copies are computed once per block and then reused for multiple comparison operations. This approach dramatically improves processing efficiency on resource-limited devices while still achieving effective deduplication and storage space reduction.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11321278B2Light-weight index deduplication and hierarchical snapshot replication
Publication Date: 2022.05.03 RUBRIK INC
  • US11321278B2 patent drawing
  • US11321278B2 patent drawing
  • US11321278B2 patent drawing

AI summary

A lightweight deduplication system can perform resource efficient data deduplication using an extent index and a content index. The extent index can store full fingerprints of data segments to be deduplicated and the content index can store shortened versions of the full fingerprints. The system can alternate between the extent and content indexes, and cache portions of the indices to perform lightweight data deduplication. Further, the system can be configured with an efficient heuristic approach for selecting content index data lookups for chains of volumes for deduplication, such as a long chain of snapshots.