Big Data Bloom Filter Segmentation for I/O Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing Bloom filters are inefficient for big data sets due to poor memory locality and high I/O and access times, especially in large key-value stores and database tables.

Innovation Solution

A novel Big Data Bloom Filter (BDBF) algorithm that segments a Bloom filter into equal-sized chunks, packing segments with the same index into extent data structures to improve filtering efficiency and reduce I/O costs by allowing multiple chunks to be filtered using a single extent.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a standard Bloom filter is used for big data sets, then space efficiency is maintained, but I/O time and access time increase significantly

Engineering Contradiction:
Improvespace efficiencyVSAvoidI/O time and access time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The Bloom filter is divided into multiple segments, where each segment can be independently stored and accessed. This segmentation allows the system to load only necessary segments into memory based on query patterns, reducing I/O operations while maintaining the probabilistic filtering capability. The patent applies this by partitioning the bit array into segments that can be selectively accessed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to the Bloom filter structure by organizing segments into different levels of caching. This multi-level hierarchy (L1 cache, L2 cache, disk storage) adds a spatial dimension to data access, allowing frequently accessed segments to reside in faster memory while less frequently accessed segments remain on disk, thereby reducing average access time without increasing overall space requirements.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If Bloom filter size is increased to handle bigger data sets, then filtering capability improves, but memory locality deteriorates

Engineering Contradiction:
Improvefiltering capabilityVSAvoidmemory locality
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

By segmenting the large Bloom filter into smaller independent units, each segment maintains better memory locality characteristics. The system can load individual segments or groups of segments into cache based on access patterns, ensuring that the working set fits within available cache memory while the overall filter structure can handle billions of elements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic segment loading and caching strategies where the system adaptively loads segments into memory based on query workloads. This dynamic behavior allows the filter to maintain high filtering capability for large data sets while optimizing memory locality by keeping only necessary segments in fast memory at any given time.

Inventive Principle:
Principle #15Dynamics

3Loss of time

If multiple Bloom filters are created to improve query performance, then access time reduces, but the number of files and storage complexity increases

Engineering Contradiction:
Improveaccess timeVSAvoidnumber of files
Core Design Contradiction:
Loss of timeVSDevice complexity

Solution Approach 1:

The patent merges multiple Bloom filter segments into a unified hierarchical structure managed by a single system. Instead of treating each segment as a separate file requiring individual management, the system combines them into a cohesive multi-level cache structure with centralized control, reducing the effective number of files from the user perspective while maintaining parallel access capabilities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The hierarchical Bloom filter structure serves multiple functions simultaneously: it provides fast filtering for hot data in L1/L2 cache, handles cold data from disk, and manages the entire filter structure as a single logical unit. This multi-functionality allows the system to achieve reduced access times through parallel segment access while presenting a simplified interface that hides the complexity of multiple underlying files.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3683696B1System and method of bloom filter for big data
Publication Date: 2025.01.08 SQREAM TECH
  • EP3683696B1 patent drawingFigure 1
  • EP3683696B1 patent drawingFigure 2A
  • EP3683696B1 patent drawingFigure 3A

AI summary

The present invention provides a method for applying bloom filter on a large data set consisting of key-value pairs, using at least one processor. The method comprising the step of: • partitioning large data-set of key-value pairs into data chunks; • determining Bloom filter Vector and number of segments in the vector for each data chunk; • Encoding all keys of a given Chunk into a Bloom filter vector; • Determining the segment-id of a given key using H (0) hash function; • Encoding Key into a Bloom filter segment with said determined segment-id, using a K-bit array produced by H1,..Hk functions; and • Packing of segments into extent data structures where each extent includes segments of different chunks, but with the same segment-id wherein a single extent filters multiple chunks, depending on a packing factor (the number of segments packed into a single extent).