Binary Trie Deduplication Hash Tracking Memory Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deduplication algorithms using 32-bit integer hashes require excessive memory to store all possible hashes, making it impractical for many modern devices to process and identify duplicates efficiently.

Innovation Solution

The implementation of a binary trie data structure with bit arrays that uses approximately 1 GB of memory to track all possible 32-bit integer hashes, eliminating the need for pointer references and optimizing memory usage by allowing constant insertion and retrieval times.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional data structures are used to store 32-bit integer hashes, then all possible hashes can be tracked, but memory consumption becomes excessive (tens or hundreds of GBs)

Engineering Contradiction:
Improvenumber of hashes that can be trackedVSAvoidmemory consumption
Core Design Contradiction:
Quantity of substanceVSWeight of stationary object

Solution Approach 1:

The patent divides the 32-bit hash space into multiple 16-bit segments. Instead of storing complete 32-bit hashes in a single large data structure, the hash is split into two 16-bit parts, and each part is tracked separately using smaller, more memory-efficient data structures. This segmentation reduces the memory footprint from tens or hundreds of GBs to a manageable size while still enabling detection of duplicate hashes.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameter representation from storing complete 32-bit hash values to storing segmented 16-bit hash parts. By transforming the data structure from a single large hash set to multiple smaller segmented structures, the memory consumption is reduced while maintaining the ability to identify duplicates through comparison of the segmented parts.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If memory is increased to store more hashes, then more duplicate hashes can be identified, but it becomes impractical for modern devices with memory limitations

Engineering Contradiction:
Improvenumber of hashes that can be trackedVSAvoidcompatibility with modern devices
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

By segmenting the hash storage into smaller 16-bit parts, the patent enables deployment on modern devices with limited memory. The segmented approach reduces the memory requirement from impractical sizes to fit within contemporary device constraints, while still maintaining comprehensive hash tracking capability across the full 32-bit hash space.

Inventive Principle:
Principle #1Segmentation

3Productivity

If a hash set is used to store 100 million hashes, then deduplication can be performed, but memory consumption reaches approximately 3 GB which may exceed internal memory limitations

Engineering Contradiction:
Improvededuplication processing capabilityVSAvoidmemory consumption
Core Design Contradiction:
ProductivityVSWeight of stationary object

Solution Approach 1:

The patent applies segmentation to reduce memory consumption from 3 GB to a smaller footprint by dividing the hash storage into segmented 16-bit parts. This allows the system to maintain deduplication processing capability for 100 million hashes while fitting within internal memory limitations of modern devices.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses lighter-weight segmented hash representations instead of complete 32-bit hash storage. By using smaller segmented structures that consume less memory, the system achieves the same deduplication functionality with reduced memory allocation, effectively replacing the memory-intensive approach with a more efficient alternative.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

Data Source

PatentUS11892980B2Memory optimized algorithm for evaluating deduplication hashes for large data sets
Publication Date: 2024.02.06 EMC IP HLDG CO LLC
  • US11892980B2 patent drawing
  • US11892980B2 patent drawing
  • US11892980B2 patent drawing

AI summary

One example method includes performing a hash of data to generate a hash value, checking a binary trie to determine if the hash value has previously been entered into the binary trie, if the hash value has previously been entered in the binary trie, declaring the data as a duplicate of other data, and if the hash value has not been previously entered in the binary trie, updating the binary trie to include the hash value.