Hadoop OLAP Engine Distributed Cube Lattice

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current OLAP systems face limitations such as single file storage with less than two billion rows, lack of distributed architecture, and prolonged query execution times, which hinder efficient business intelligence data analysis on large datasets.

Innovation Solution

A Hadoop OLAP system is implemented, utilizing a distributed architecture with HDFS for storage and MapReduce for processing, coupled with a cube store like HBase for fast random access, enabling multi-dimensional analysis and scalable data storage and processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional OLAP systems use single file storage, then implementation is simple, but storage capacity is limited to less than two billion rows

Engineering Contradiction:
Improvestorage capacityVSAvoidarchitecture complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent divides the storage system into multiple segments including HDFS for distributed file storage, cube store for pre-computed data cubes, and relational database for metadata. This segmentation allows the system to handle terabyte-sized datasets by distributing storage across multiple nodes rather than relying on a single file, thereby resolving the contradiction between storage capacity and architecture complexity.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If traditional OLAP systems lack distributed architecture, then system management is simple, but scalability is limited

Engineering Contradiction:
ImprovescalabilityVSAvoidsystem architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a distributed architecture dimension by implementing Hadoop OLAP that leverages HDFS distributed file storage and MapReduce distributed processing. This adds a horizontal scaling dimension to the system, allowing it to expand across multiple nodes and handle terabyte-sized data cubes, thereby resolving the contradiction between scalability and system architecture complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If traditional OLAP systems process queries without distributed computing, then processing logic is simple, but query execution time exceeds 24 hours

Engineering Contradiction:
Improvequery execution speedVSAvoidprocessing architecture
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the query processing workload into multiple independent MapReduce jobs that can execute in parallel across distributed nodes. Each job handles a portion of the data cube construction or query processing, enabling terabyte-sized data to be processed efficiently within acceptable timeframes rather than exceeding 24 hours, thereby resolving the contradiction between query execution speed and processing architecture complexity.

Inventive Principle:
Principle #1Segmentation

4Quantity of substance

If OLAP systems support terabyte-sized data cubes, then data capacity increases, but memory requirements exceed available RAM

Engineering Contradiction:
Improvedata capacityVSAvoidmemory usage
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent extracts the data storage and processing functions from main memory by implementing a distributed file storage system (HDFS) and distributed processing framework (MapReduce). This allows terabyte-sized data cubes to be stored and processed across the distributed file system rather than requiring them to fit into limited RAM, thereby resolving the contradiction between data capacity and memory usage.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11537635B2Hadoop OLAP engine
Publication Date: 2022.12.27 EBAY INC
  • US11537635B2 patent drawing
  • US11537635B2 patent drawing
  • US11537635B2 patent drawing

AI summary

In various example embodiments, systems and methods for building data cubes to be stored in a cube store are presented. In some embodiments, a metadata engine generates the cube metadata. In further embodiments, cube data is generated by a cube build engine based on the cube metadata and source data. The cube build engine performs a multi-stage MapReduce job on the source data to produce a multi-dimensional cube lattice having multiple cuboids. In further embodiments, the cube data is provided to the cube store.