Binary Trie Deduplication Hash Tracking Memory Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional deduplication algorithms using 32-bit integer hashes require excessive memory to store all possible hashes, making it impractical for many modern devices to process and identify duplicates efficiently.
Innovation Solution
The implementation of a binary trie data structure with bit arrays that uses approximately 1 GB of memory to track all possible 32-bit integer hashes, eliminating the need for pointer references and optimizing memory usage by allowing constant insertion and retrieval times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional data structures are used to store 32-bit integer hashes, then all possible hashes can be tracked, but memory consumption becomes excessive (tens or hundreds of GBs)
Solution Approach 1:
The patent divides the 32-bit hash space into multiple 16-bit segments. Instead of storing complete 32-bit hashes in a single large data structure, the hash is split into two 16-bit parts, and each part is tracked separately using smaller, more memory-efficient data structures. This segmentation reduces the memory footprint from tens or hundreds of GBs to a manageable size while still enabling detection of duplicate hashes.
Solution Approach 2:
The patent changes the parameter representation from storing complete 32-bit hash values to storing segmented 16-bit hash parts. By transforming the data structure from a single large hash set to multiple smaller segmented structures, the memory consumption is reduced while maintaining the ability to identify duplicates through comparison of the segmented parts.
2Quantity of substance
If memory is increased to store more hashes, then more duplicate hashes can be identified, but it becomes impractical for modern devices with memory limitations
Solution Approach 1:
By segmenting the hash storage into smaller 16-bit parts, the patent enables deployment on modern devices with limited memory. The segmented approach reduces the memory requirement from impractical sizes to fit within contemporary device constraints, while still maintaining comprehensive hash tracking capability across the full 32-bit hash space.
3Productivity
If a hash set is used to store 100 million hashes, then deduplication can be performed, but memory consumption reaches approximately 3 GB which may exceed internal memory limitations
Solution Approach 1:
The patent applies segmentation to reduce memory consumption from 3 GB to a smaller footprint by dividing the hash storage into segmented 16-bit parts. This allows the system to maintain deduplication processing capability for 100 million hashes while fitting within internal memory limitations of modern devices.
Solution Approach 2:
The patent uses lighter-weight segmented hash representations instead of complete 32-bit hash storage. By using smaller segmented structures that consume less memory, the system achieves the same deduplication functionality with reduced memory allocation, effectively replacing the memory-intensive approach with a more efficient alternative.
Data Source
AI summary
One example method includes performing a hash of data to generate a hash value, checking a binary trie to determine if the hash value has previously been entered into the binary trie, if the hash value has previously been entered in the binary trie, declaring the data as a duplicate of other data, and if the hash value has not been previously entered in the binary trie, updating the binary trie to include the hash value.


