Cloud Data Parallel Clusters for Massive Analytics Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data warehouses struggle to provide reasonable response times due to the increasing volume of data, especially in industries like telecommunications, where massive data volumes hinder efficient data management and analysis, limiting the use of advanced analytics.
Innovation Solution
Implementing data parallel and compute parallel techniques over a cloud configuration, enabling efficient access and processing of massive structured data using mapping APIs for existing business intelligence tools, and utilizing query processors to generate query results, while supporting high-level query languages like SQL.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of time
If data is stored in existing data warehouses, then data storage capacity is maintained, but response time to data management requests deteriorates due to expanding data volume
Solution Approach 1:
The patent segments the monolithic data warehouse into multiple data parallel clusters, each capable of independent processing. This segmentation allows the system to handle expanding data volumes by distributing data across multiple nodes while maintaining responsive query performance through parallel processing capabilities.
2Productivity
If data is partitioned across multiple machines, then parallel I/O scan capability is improved, but system complexity increases
Solution Approach 1:
The patent implements a universal query processor that can operate across multiple data parallel clusters with identical functionality. Each cluster node performs the same query processing operations, providing multi-functionality that simplifies the overall system architecture while maintaining high parallel I/O scan capability through coordinated operation across nodes.
3Adaptability or versatility
If advanced analytics applications are added to data warehouse, then analytics capability is improved, but performance on existing reporting functions deteriorates
Solution Approach 1:
The patent segments the system into distinct data parallel clusters that can be dedicated to specific workloads. This allows advanced analytics applications to run on dedicated clusters without interfering with the performance of existing reporting functions on other clusters, thereby maintaining both analytics capability and reporting performance simultaneously.
4Loss of time
If data is processed in real-time, then responsiveness is improved, but processing cost increases
Solution Approach 1:
The patent implements partial real-time processing where only the necessary portion of data is processed at high speed for time-sensitive operations, while other data can be processed at lower priority. This approach achieves real-time responsiveness for critical functions without incurring the full cost of real-time processing across all data operations.
Data Source
AI summary
Embodiments of the invention provide data management solutions that go beyond the traditional warehousing system to support advanced analytics. Furthermore, embodiments of the invention relate to systems and methods for extracting data from an existing data warehouse, storing the extracted data in a reusable (intermediate) form using data parallel and compute parallel techniques over cloud, query processing over the data with/without compute parallel techniques, and providing querying using high level querying languages.


