Selective Deduplication in Storage Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data storage systems face challenges in efficiently using resources during inline deduplication, as they must perform deduplication operations in real-time, which can impact performance and require significant computational resources, while post-processing deduplication delays storage space savings until after data is initially stored.

Innovation Solution

Implementing selective deduplication by determining the probability of deduplication for each data object based on its characteristics and only performing deduplication on objects with a high probability of benefit, allowing resources to be focused on those that will yield storage space savings, thereby reducing the need for additional storage space and maintaining system performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If inline deduplication is performed on all data objects, then storage space savings are achieved, but system performance deteriorates due to computational overhead

Engineering Contradiction:
Improvestorage spaceVSAvoidsystem performance
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent applies partial action by performing deduplication only on selected data objects rather than all data objects. A selection criterion is used to identify which data objects should undergo deduplication, thereby reducing computational overhead while still achieving storage space savings on the most beneficial objects.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent changes the parameter of deduplication application from universal to selective by introducing a selection criterion. This parameter change allows the system to adapt deduplication operations based on data object characteristics, balancing storage savings with performance requirements.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If post-processing deduplication is used, then system performance is maintained, but storage space savings are delayed until after data is stored

Engineering Contradiction:
Improvesystem performanceVSAvoidtime to realize storage savings
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing deduplication operations before data objects are permanently stored to persistent storage. By selecting data objects for deduplication prior to storage and executing the deduplication in advance, the system realizes storage space savings sooner while still maintaining performance through selective application.

Inventive Principle:
Principle #10Preliminary action

3Loss of energy

If selective deduplication is applied, then resource efficiency is improved, but the complexity of determining selection criteria increases

Engineering Contradiction:
Improvecomputational resource efficiencyVSAvoidselection criterion complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent applies local quality by creating different selection criteria for different data objects based on their individual characteristics. Rather than using a single complex criterion for all objects, the system evaluates specific attributes of each data object to determine suitability for deduplication, localizing the decision-making process.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11169967B2Selective deduplication
Publication Date: 2021.11.09 NETAPP INC
  • US11169967B2 patent drawing
  • US11169967B2 patent drawing
  • US11169967B2 patent drawing

AI summary

Methods and apparatuses for performing selective deduplication in a storage system are introduced here. Techniques are provided for determining a probability of deduplication for a data object based on a characteristic of the data object and performing a deduplication operation on the data object in the storage system prior to the data object being stored in persistent storage of the storage system if the probability of deduplication for the data object has a specified relationship to a specified threshold.