Selective Deduplication in Storage Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data storage systems face challenges in efficiently using resources during inline deduplication, as they must perform deduplication operations in real-time, which can impact performance and require significant computational resources, while post-processing deduplication delays storage space savings until after data is initially stored.
Innovation Solution
Implementing selective deduplication by determining the probability of deduplication for each data object based on its characteristics and only performing deduplication on objects with a high probability of benefit, allowing resources to be focused on those that will yield storage space savings, thereby reducing the need for additional storage space and maintaining system performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If inline deduplication is performed on all data objects, then storage space savings are achieved, but system performance deteriorates due to computational overhead
Solution Approach 1:
The patent applies partial action by performing deduplication only on selected data objects rather than all data objects. A selection criterion is used to identify which data objects should undergo deduplication, thereby reducing computational overhead while still achieving storage space savings on the most beneficial objects.
Solution Approach 2:
The patent changes the parameter of deduplication application from universal to selective by introducing a selection criterion. This parameter change allows the system to adapt deduplication operations based on data object characteristics, balancing storage savings with performance requirements.
2Productivity
If post-processing deduplication is used, then system performance is maintained, but storage space savings are delayed until after data is stored
Solution Approach 1:
The patent applies preliminary action by performing deduplication operations before data objects are permanently stored to persistent storage. By selecting data objects for deduplication prior to storage and executing the deduplication in advance, the system realizes storage space savings sooner while still maintaining performance through selective application.
3Loss of energy
If selective deduplication is applied, then resource efficiency is improved, but the complexity of determining selection criteria increases
Solution Approach 1:
The patent applies local quality by creating different selection criteria for different data objects based on their individual characteristics. Rather than using a single complex criterion for all objects, the system evaluates specific attributes of each data object to determine suitability for deduplication, localizing the decision-making process.
Data Source
AI summary
Methods and apparatuses for performing selective deduplication in a storage system are introduced here. Techniques are provided for determining a probability of deduplication for a data object based on a characteristic of the data object and performing a deduplication operation on the data object in the storage system prior to the data object being stored in persistent storage of the storage system if the probability of deduplication for the data object has a specified relationship to a specified threshold.


