Clustered Hash Filters for Lower False Positives in Range Search
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Bloom filters suffer from high false positive rates while maintaining a small size, necessitating improvements in data storage systems for databases.
Innovation Solution
Implementing a hash-based filter that clusters data fields into smaller groups, using cluster identifiers and reduced hash functions to reduce the number of entries, thereby improving false positive rates and maintaining a compact filter size.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the number of entries m in the bloom filter is increased to improve false positive ratio, then the false positive rate decreases, but the filter size increases
Solution Approach 1:
The patent divides the data fields into multiple clusters, where each cluster is represented by a cluster identifier. Instead of creating a bloom filter for all fields individually, the system creates separate bloom filters for each cluster identifier. This segmentation reduces the number of entries needed in each individual bloom filter while maintaining the ability to accurately filter queries across all fields.
2Productivity
If a traditional bloom filter is used for all field values in a column group, then query acceleration is achieved, but the false positive rate increases
Solution Approach 1:
The patent segments the column group into multiple clusters based on field values, creating a distinct bloom filter for each cluster. This allows the system to maintain compact filter sizes for query acceleration while reducing false positives by comparing queries against multiple specialized filters rather than one large general filter.
Solution Approach 2:
Each bloom filter is tailored to a specific cluster of field values, optimizing the filter characteristics for that particular data subset. This local optimization ensures that each filter is highly effective for its designated cluster, improving overall reliability while maintaining the productivity benefits of bloom filter-based query acceleration.
Data Source
AI summary
A method for cluster based searching for a value range stored in a storage system, the method may include receiving a request to find a certain value range within a set of information elements that are stored in a storage system; wherein the set of information elements comprises subsets of information elements associated with subset hash based filters; wherein different subsets of information elements are associated with different subset hash based filters; determining a certain cluster value of a certain cluster that comprises the certain value range; applying one or more hush functions on the certain cluster value to provide one or more hash results; and determining whether one or more members of the certain cluster are possibly in a subset of information elements, based on the one or more hash results and on a subset hash based filter of the subset of information elements; and when determining that the one or more members of the certain cluster are possibly in the subset then searching, within the subset, a matching information element that matches the certain value range.


