Big Data Bloom Filter Segmentation for I/O Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing Bloom filters are inefficient for big data sets due to poor memory locality and high I/O and access times, especially in large key-value stores and database tables.
Innovation Solution
A novel Big Data Bloom Filter (BDBF) algorithm that segments a Bloom filter into equal-sized chunks, packing segments with the same index into extent data structures to improve filtering efficiency and reduce I/O costs by allowing multiple chunks to be filtered using a single extent.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If a standard Bloom filter is used for big data sets, then space efficiency is maintained, but I/O time and access time increase significantly
Solution Approach 1:
The Bloom filter is divided into multiple segments, where each segment can be independently stored and accessed. This segmentation allows the system to load only necessary segments into memory based on query patterns, reducing I/O operations while maintaining the probabilistic filtering capability. The patent applies this by partitioning the bit array into segments that can be selectively accessed.
Solution Approach 2:
The patent introduces a hierarchical dimension to the Bloom filter structure by organizing segments into different levels of caching. This multi-level hierarchy (L1 cache, L2 cache, disk storage) adds a spatial dimension to data access, allowing frequently accessed segments to reside in faster memory while less frequently accessed segments remain on disk, thereby reducing average access time without increasing overall space requirements.
2Adaptability or versatility
If Bloom filter size is increased to handle bigger data sets, then filtering capability improves, but memory locality deteriorates
Solution Approach 1:
By segmenting the large Bloom filter into smaller independent units, each segment maintains better memory locality characteristics. The system can load individual segments or groups of segments into cache based on access patterns, ensuring that the working set fits within available cache memory while the overall filter structure can handle billions of elements.
Solution Approach 2:
The patent implements dynamic segment loading and caching strategies where the system adaptively loads segments into memory based on query workloads. This dynamic behavior allows the filter to maintain high filtering capability for large data sets while optimizing memory locality by keeping only necessary segments in fast memory at any given time.
3Loss of time
If multiple Bloom filters are created to improve query performance, then access time reduces, but the number of files and storage complexity increases
Solution Approach 1:
The patent merges multiple Bloom filter segments into a unified hierarchical structure managed by a single system. Instead of treating each segment as a separate file requiring individual management, the system combines them into a cohesive multi-level cache structure with centralized control, reducing the effective number of files from the user perspective while maintaining parallel access capabilities.
Solution Approach 2:
The hierarchical Bloom filter structure serves multiple functions simultaneously: it provides fast filtering for hot data in L1/L2 cache, handles cold data from disk, and manages the entire filter structure as a single logical unit. This multi-functionality allows the system to achieve reduced access times through parallel segment access while presenting a simplified interface that hides the complexity of multiple underlying files.
Data Source
Figure 1
Figure 2A
Figure 3A
AI summary
The present invention provides a method for applying bloom filter on a large data set consisting of key-value pairs, using at least one processor. The method comprising the step of: • partitioning large data-set of key-value pairs into data chunks; • determining Bloom filter Vector and number of segments in the vector for each data chunk; • Encoding all keys of a given Chunk into a Bloom filter vector; • Determining the segment-id of a given key using H (0) hash function; • Encoding Key into a Bloom filter segment with said determined segment-id, using a K-bit array produced by H1,..Hk functions; and • Packing of segments into extent data structures where each extent includes segments of different chunks, but with the same segment-id wherein a single extent filters multiple chunks, depending on a packing factor (the number of segments packed into a single extent).