Data Warehouse Summarization via Relational Synopses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing and efficiently summarizing large and dynamically growing data collections in a way that facilitates quick derivation of pertinent information and trend discovery in complex processes is challenging due to rapid data growth and complexity.

Innovation Solution

A system and method utilizing a summarization component that employs a declarative framework with probabilistic states and observations to generate relational synopses, optimizing data analysis through pre-computation and materialization, and adaptively tuning performance using an optimizer and adaptor enhancer module.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional data warehousing methods are used to manage large data collections, then data storage capacity is maintained, but data summarization efficiency and query performance deteriorate as data volume grows

Engineering Contradiction:
Improvedata summarization efficiencyVSAvoiddata volume
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent segments the large data collection into multiple partitions or chunks that can be processed independently. The summarization component divides the data stream into manageable segments, allowing parallel processing and efficient summarization of each segment before aggregation, thus maintaining summarization efficiency despite growing data volume.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary summarization actions on incoming data streams before full analysis is required. By pre-computing summary statistics and maintaining relational synopses as data arrives, the system prepares processed information in advance, enabling rapid querying and trend discovery without reprocessing entire data sets.

Inventive Principle:
Principle #10Preliminary action

2Loss of information

If detailed analysis of large data sets is performed to discover trends, then information completeness is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveinformation completenessVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent creates simplified copies or synopses of the original data that capture essential patterns and trends. These relational synopses are maintained as separate data structures that mirror key characteristics of the full data set, allowing rapid analysis of trend information without processing the complete original data volume.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The summarization component acts as an intermediary between raw data and analysis queries. It transforms detailed data into intermediate summary representations that retain sufficient information for trend discovery while being computationally efficient to process, thus mediating between information completeness and processing speed requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If real-time summarization is performed on high data loads, then data currency is improved, but system performance and accuracy may deteriorate

Engineering Contradiction:
Improvedata accuracyVSAvoidsystem performance under load
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system dynamically adjusts its summarization parameters and processing intensity based on incoming data load conditions. When data volume increases, the summarization component adapts by adjusting window sizes, aggregation frequencies, or sampling rates to maintain accurate summaries while preserving system performance and preventing overload.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS7933861B2Process data warehouse
Publication Date: 2011.04.26 UNIV OF PITTSBURGH OF THE COMMONWEALTH SYST OF HIGHER EDUCATION
  • US7933861B2 patent drawing
  • US7933861B2 patent drawing
  • US7933861B2 patent drawing

AI summary

Systems and/or methods are presented that can efficiently analyze and summarize large collections of data. A summarization component can employ mapping rules to map received data into specified states and observations of interest, which can be utilized to facilitate creating relational tables that can be utilized to facilitate summarizing a collection of data based in part on predefined summarization criteria. An optimizer component can employ pre-computing and materialization of the process behavior to facilitate optimizing data analysis. An adaptor enhancer component can monitor and evaluate system performance and can generate mapping rules that can facilitate improving system performance.