Cuckoo Filter Resizing via Fingerprint Hashing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Approximate set membership data structures, such as cuckoo filters, face challenges in efficiently supporting deletions and resizing operations while maintaining low memory usage and fast lookup performance, especially when filled to high load factors or when the original data is inaccessible.
Innovation Solution
A cuckoo filter is implemented with a compressed format that stores only occupied slots, using a fullness counter array to encode empty slots, and a resizing mechanism that doubles or halves the filter's capacity without accessing original data, allowing for efficient insertions, deletions, and lookups by maintaining multiple views of the filter and using a resize counter to scale bucket indexes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the filter is filled to high load factors to reduce space usage, then storage efficiency is improved, but insertion throughput decreases substantially
Solution Approach 1:
The patent implements dynamic resizing capability that allows the filter to adapt its capacity based on load conditions. When the filter approaches high load factors, it can automatically resize to a larger capacity, maintaining both high storage efficiency during low-load periods and high insertion throughput during growth phases without manual intervention.
Solution Approach 2:
The patent changes the parameter of filter capacity dynamically through resizing operations. By adjusting the number of buckets and slots based on actual usage patterns and load factors, the system optimizes the balance between storage efficiency and insertion performance, allowing the filter to operate at optimal points rather than being fixed at a single capacity.
2Adaptability or versatility
If deletion operation is supported in Bloom filter variants, then functionality is improved, but storage cost increases 2× to 4×
Solution Approach 1:
The patent extracts only the necessary information for deletion support by using lazy deletion markers and reference counting mechanisms rather than storing complete deletion metadata. This approach enables deletion functionality while minimizing the additional storage overhead compared to full Bloom filter variants.
Solution Approach 2:
The patent uses fingerprint copying and reference counting to support deletions efficiently. Instead of storing multiple copies of deletion information, it maintains reference counts that track element occurrences, allowing deletion operations to be performed with minimal additional storage beyond the original filter structure.
3Speed
If resizing operation is performed without accessing original data, then operation speed is improved, but data availability requirements are relaxed
Solution Approach 1:
The patent uses fingerprint copying to enable resizing without accessing original data. By storing hashed fingerprints rather than original keys, the system can perform resizing operations on the fingerprint array independently, achieving fast resizing while the original data remains inaccessible without impacting the resizing process.
Solution Approach 2:
The patent performs preliminary hashing of original data into fingerprints before resizing operations are needed. This preliminary transformation allows subsequent resizing to be performed purely on the fingerprint structure without needing to access or reprocess the original data, significantly improving resizing speed.
Data Source
AI summary
A method of maintaining a probabilistic filter includes, in response to receiving a key K1 for adding to the probabilistic filter, generating a fingerprint F1 based on applying a fingerprint hash function HF to the key K1, identifying an initial bucket Bi1 by selecting between at least a first bucket B1 determined based on a first bucket hash function H1 of the key K1 and a second bucket B2 determined based on a second bucket hash function H2 of the key K1, and inserting the fingerprint F1 into the initial bucket Bi1; and resizing the probabilistic filter. Resizing the probabilistic filter includes incrementing a resize counter value, determining a bucket B′ for the fingerprint F1 based on a value of the fingerprint F1 and the resize counter value, and inserting the fingerprint F1 into the bucket B′ in the probabilistic filter.


