Polynomial-Based Signature Subspace for Deduplication Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deduplication methods in information processing systems require substantial computational and memory resources, leading to inefficiencies and performance issues in generating deduplication estimates for storage systems.
Innovation Solution
The use of polynomial-based signature subspaces to efficiently generate deduplication estimates by computing polynomial-based signatures for dataset pages and updating a deduplication estimate table, reducing the need for extensive computational and memory resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional deduplication methods are used to generate deduplication estimates, then accurate deduplication decisions can be made, but substantial computational and memory resources are consumed
Solution Approach 1:
The patent segments the deduplication estimation process into two stages: first computing lightweight polynomial-based signatures for all pages to identify candidate duplicates, then computing full content-based signatures only for pages that satisfy the subset inclusion characteristic. This segmentation reduces the number of expensive full signature computations while maintaining estimation accuracy.
Solution Approach 2:
The patent introduces polynomial-based signatures as an intermediary mechanism between page data and full content-based signatures. These intermediate signatures serve as a filtering layer that identifies potential duplicates without requiring the full computational overhead of content-based signature computation, thereby reducing overall resource consumption.
2Measurement precision
If conventional deduplication methods are used to generate deduplication estimates, then comprehensive deduplication analysis is achieved, but system performance is significantly undermined
Solution Approach 1:
The patent applies partial action by computing polynomial-based signatures for all pages (excessive action) to ensure comprehensive coverage, but only computes full content-based signatures for a subset of pages that satisfy the inclusion characteristic (partial action). This approach maintains comprehensiveness while improving performance by avoiding unnecessary full signature computations.
Solution Approach 2:
The patent performs preliminary computation of polynomial-based signatures for all pages before conducting the actual deduplication estimation. This preliminary action creates a filtered subset of candidate pages that are more likely to be duplicates, allowing the main deduplication estimation process to focus only on these candidates and thereby improving overall system performance.
3Use of energy by moving object
If polynomial-based signature subspaces are used to scan dataset pages, then computational resources are reduced, but the scanning process requires sophisticated signature computation and filtering
Solution Approach 1:
The patent changes the parameter of signature computation from full content-based signatures to polynomial-based signatures for the scanning process. This parameter change reduces computational complexity while maintaining the ability to identify duplicate pages. The polynomial-based signatures use simpler mathematical operations that are less resource-intensive than full hash computations.
Solution Approach 2:
The patent extracts only the essential characteristics needed for duplicate identification by using polynomial-based signatures that capture key features of page content without processing the entire page data. This extraction approach reduces computational resources by focusing only on the most relevant features for deduplication while simplifying the overall processing complexity.
Data Source
AI summary
An apparatus in one embodiment comprises at least one processing device comprising a processor coupled to a memory. The processing device is configured to identify a dataset to be scanned to generate a deduplication estimate for that dataset, to designate a subset inclusion characteristic to be utilized in the scan, and for each of a plurality of pages of the dataset, to scan the page, where scanning the page includes computing a polynomial-based signature for the page, determining whether or not the polynomial-based signature satisfies the designated subset inclusion characteristic, and responsive to the polynomial-based signature satisfying the designated subset inclusion characteristic, computing a content-based signature for the page and updating a corresponding entry of a deduplication estimate table for the dataset based at least in part on the content-based signature. The processing device generates the deduplication estimate for the dataset based at least in part on contents of the deduplication estimate table.


