Vector Architecture Query Processing with Bloom Filter Bitmasks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for processing complex data queries in 'Big Data' stored in relational databases are inefficient, as they do not effectively leverage GPU capabilities for parallelized processing, limiting the ability to handle large datasets efficiently.

Innovation Solution

A method and system that utilize enhanced data structure representations, including logical data chunk boundaries and Bloom filter bitmasks, to enable simultaneous search across multiple data chunks using HWA processors, allowing for efficient query processing by identifying data item addresses and executing queries on respective data chunks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data queries are processed in parallel using known parallel computational processing methods, then processing capability is improved, but the methods only solve specific tailored scenarios and do not effectively leverage GPU capabilities for general Big Data queries

Engineering Contradiction:
Improvequery processing capabilityVSAvoidapplicability to different query scenarios
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent segments Big Data into multiple data chunks and organizes them in a tree-like hierarchical structure with chunk groups and chunks. This segmentation enables parallel processing of different chunks while maintaining the ability to handle various query types. Each chunk is independently processable by GPU threads, achieving both parallel productivity and versatility across different query scenarios.

Inventive Principle:
Principle #1Segmentation

2Productivity

If enhanced data structure representation with logical data chunk boundaries and Bloom filter bitmask is used, then search capability and compression ratio are improved, but data structure complexity increases

Engineering Contradiction:
Improvesearch efficiencyVSAvoiddata structure complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent pre-computes and stores Bloom filter bitmasks and logical data chunk boundaries during data loading and preprocessing stages. This preliminary action enables fast search operations during query execution without computing these structures in real-time. The enhanced data structure is built once and reused, improving search efficiency while the complexity is paid for only during the initial setup phase.

Inventive Principle:
Principle #10Preliminary action

3Speed

If simultaneous search across multiple data chunks is performed using HWA processors, then processing speed is improved, but memory requirements and computational overhead increase

Engineering Contradiction:
Improvesearch speedVSAvoidmemory consumption
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent uses Bloom filter bitmasks to perform partial filtering before full data chunk processing. The Bloom filter provides a probabilistic first layer of filtering that quickly eliminates chunks that definitely don't match the query, allowing HWA processors to focus only on relevant chunks. This partial action approach speeds up search while reducing the effective memory and computational load on the expensive HWA processing stage.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP2880566B1A method for pre-processing and processing query operation on multiple data chunk on vector enabled architecture
Publication Date: 2019.07.17 SQREAM TECH
  • EP2880566B1 patent drawingFigure 1
  • EP2880566B1 patent drawingFigure 2
  • EP2880566B1 patent drawingFigure 3

AI summary

The present invention provides a method for pre-processing and processing query operation on multiple data chunk on vector enabled architecture. The method comprising the step of : receiving user query having at least one a data item, accessing data chunk blocks having enhanced data structure representation, wherein the enhanced data structure representation includes data recursive presentation of data chunk boundaries and bloom filter bitmask of data chunks, search simultaneously at multiple data chunk blocks utilizing the recursive presentation of data chunk boundaries using HWA, identifying data item address by comparing calculated Bloom filter bitmask of the requested data item to calculated bitmask of the respective data chunks simultaneously by using multiple HWA, executing query on respective data chunk.