Polynomial-Based Signature Subspace for Deduplication Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deduplication methods in information processing systems require substantial computational and memory resources, leading to inefficiencies and performance issues in generating deduplication estimates for storage systems.

Innovation Solution

The use of polynomial-based signature subspaces to efficiently generate deduplication estimates by computing polynomial-based signatures for dataset pages and updating a deduplication estimate table, reducing the need for extensive computational and memory resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional deduplication methods are used to generate deduplication estimates, then accurate deduplication decisions can be made, but substantial computational and memory resources are consumed

Engineering Contradiction:
Improvededuplication estimate accuracyVSAvoidcomputational and memory resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the deduplication estimation process into two stages: first computing lightweight polynomial-based signatures for all pages to identify candidate duplicates, then computing full content-based signatures only for pages that satisfy the subset inclusion characteristic. This segmentation reduces the number of expensive full signature computations while maintaining estimation accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces polynomial-based signatures as an intermediary mechanism between page data and full content-based signatures. These intermediate signatures serve as a filtering layer that identifies potential duplicates without requiring the full computational overhead of content-based signature computation, thereby reducing overall resource consumption.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If conventional deduplication methods are used to generate deduplication estimates, then comprehensive deduplication analysis is achieved, but system performance is significantly undermined

Engineering Contradiction:
Improvededuplication estimate comprehensivenessVSAvoidsystem performance
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies partial action by computing polynomial-based signatures for all pages (excessive action) to ensure comprehensive coverage, but only computes full content-based signatures for a subset of pages that satisfy the inclusion characteristic (partial action). This approach maintains comprehensiveness while improving performance by avoiding unnecessary full signature computations.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary computation of polynomial-based signatures for all pages before conducting the actual deduplication estimation. This preliminary action creates a filtered subset of candidate pages that are more likely to be duplicates, allowing the main deduplication estimation process to focus only on these candidates and thereby improving overall system performance.

Inventive Principle:
Principle #10Preliminary action

3Use of energy by moving object

If polynomial-based signature subspaces are used to scan dataset pages, then computational resources are reduced, but the scanning process requires sophisticated signature computation and filtering

Engineering Contradiction:
Improvecomputational resourcesVSAvoidsignature computation and filtering process
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent changes the parameter of signature computation from full content-based signatures to polynomial-based signatures for the scanning process. This parameter change reduces computational complexity while maintaining the ability to identify duplicate pages. The polynomial-based signatures use simpler mathematical operations that are less resource-intensive than full hash computations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent extracts only the essential characteristics needed for duplicate identification by using polynomial-based signatures that capture key features of page content without processing the entire page data. This extraction approach reduces computational resources by focusing only on the most relevant features for deduplication while simplifying the overall processing complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10983962B2Processing device utilizing polynomial-based signature subspace for efficient generation of deduplication estimate
Publication Date: 2021.04.20 EMC IP HLDG CO LLC
  • US10983962B2 patent drawing
  • US10983962B2 patent drawing
  • US10983962B2 patent drawing

AI summary

An apparatus in one embodiment comprises at least one processing device comprising a processor coupled to a memory. The processing device is configured to identify a dataset to be scanned to generate a deduplication estimate for that dataset, to designate a subset inclusion characteristic to be utilized in the scan, and for each of a plurality of pages of the dataset, to scan the page, where scanning the page includes computing a polynomial-based signature for the page, determining whether or not the polynomial-based signature satisfies the designated subset inclusion characteristic, and responsive to the polynomial-based signature satisfying the designated subset inclusion characteristic, computing a content-based signature for the page and updating a corresponding entry of a deduplication estimate table for the dataset based at least in part on the content-based signature. The processing device generates the deduplication estimate for the dataset based at least in part on contents of the deduplication estimate table.