Cloud Data Parallel Clusters for Massive Analytics Latency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data warehouses struggle to provide reasonable response times due to the increasing volume of data, especially in industries like telecommunications, where massive data volumes hinder efficient data management and analysis, limiting the use of advanced analytics.

Innovation Solution

Implementing data parallel and compute parallel techniques over a cloud configuration, enabling efficient access and processing of massive structured data using mapping APIs for existing business intelligence tools, and utilizing query processors to generate query results, while supporting high-level query languages like SQL.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If data is stored in existing data warehouses, then data storage capacity is maintained, but response time to data management requests deteriorates due to expanding data volume

Engineering Contradiction:
Improveresponse timeVSAvoiddata volume
Core Design Contradiction:
Loss of timeVSQuantity of substance

Solution Approach 1:

The patent segments the monolithic data warehouse into multiple data parallel clusters, each capable of independent processing. This segmentation allows the system to handle expanding data volumes by distributing data across multiple nodes while maintaining responsive query performance through parallel processing capabilities.

Inventive Principle:
Principle #1Segmentation

2Productivity

If data is partitioned across multiple machines, then parallel I/O scan capability is improved, but system complexity increases

Engineering Contradiction:
Improveparallel I/O scan capabilityVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a universal query processor that can operate across multiple data parallel clusters with identical functionality. Each cluster node performs the same query processing operations, providing multi-functionality that simplifies the overall system architecture while maintaining high parallel I/O scan capability through coordinated operation across nodes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If advanced analytics applications are added to data warehouse, then analytics capability is improved, but performance on existing reporting functions deteriorates

Engineering Contradiction:
Improveanalytics capabilityVSAvoidreporting function performance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent segments the system into distinct data parallel clusters that can be dedicated to specific workloads. This allows advanced analytics applications to run on dedicated clusters without interfering with the performance of existing reporting functions on other clusters, thereby maintaining both analytics capability and reporting performance simultaneously.

Inventive Principle:
Principle #1Segmentation

4Loss of time

If data is processed in real-time, then responsiveness is improved, but processing cost increases

Engineering Contradiction:
Improveprocessing latencyVSAvoidprocessing cost
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The patent implements partial real-time processing where only the necessary portion of data is processed at high speed for time-sensitive operations, while other data can be processed at lower priority. This approach achieves real-time responsiveness for critical functions without incurring the full cost of real-time processing across all data operations.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS8874600B2System and method for building a cloud aware massive data analytics solution background
Publication Date: 2014.10.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US8874600B2 patent drawing
  • US8874600B2 patent drawing
  • US8874600B2 patent drawing

AI summary

Embodiments of the invention provide data management solutions that go beyond the traditional warehousing system to support advanced analytics. Furthermore, embodiments of the invention relate to systems and methods for extracting data from an existing data warehouse, storing the extracted data in a reusable (intermediate) form using data parallel and compute parallel techniques over cloud, query processing over the data with/without compute parallel techniques, and providing querying using high level querying languages.