Configurable Data Aggregation in Data Warehouses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data warehouse systems face challenges in managing large volumes of data over time, with issues related to storage and access efficiency, as well as the complexity of collecting data from multiple sources, leading to the need for improved data management and aggregation techniques.
Innovation Solution
A computer-implemented method and apparatus that identifies a policy for managing data in a storage system, aggregates raw data into summarized forms based on configurable policies, and provides mechanisms for pruning data to optimize storage usage, using intelligent remote agents to collect and process data from various sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If data is collected in fine granular format from different sources, then data completeness and detail are improved, but storage requirements and system complexity increase
Solution Approach 1:
The patent segments data into different aggregation levels (fine granular, intermediate, and summary levels), allowing the system to store detailed data when needed while maintaining compressed summary versions for historical periods, thus reducing overall storage requirements while preserving data completeness
Solution Approach 2:
The patent merges multiple data sources and multiple granularities into a unified data warehouse structure, combining fine-grained detailed data with coarser aggregated data in a single system, eliminating the need for separate storage systems and reducing overall complexity
2Loss of information
If data is collected in fine granular format from different sources, then data detail is improved, but device complexity increases
Solution Approach 1:
The system segments data management into distinct aggregation levels and uses automated ETL processes for each level, simplifying the overall system architecture by organizing complexity into manageable, standardized segments rather than handling all fine-grained data uniformly
Solution Approach 2:
The patent implements a universal data warehouse structure that can handle multiple data sources, multiple granularities, and multiple query types through a single standardized system, eliminating the need for separate specialized systems for each data type or source
3Duration of action of stationary object
If large amounts of data are accumulated over long periods, then historical data availability is improved, but access performance deteriorates
Solution Approach 1:
The patent merges detailed fine-grained data with pre-computed aggregated data in a unified query processing system, allowing queries to execute against compact aggregated representations when full detail is not required, thus maintaining fast access performance while preserving long-term historical data
Solution Approach 2:
The system dynamically adjusts query execution strategies based on data granularity and time ranges, automatically selecting optimized access paths for historical data queries versus recent detailed data queries, maintaining performance across different time periods and data volumes
4Quantity of substance
If older data is moved to secondary storage, then storage cost is reduced, but data access speed deteriorates
Solution Approach 1:
The patent merges primary and secondary storage systems into a unified data warehouse with automated data lifecycle management, allowing the system to present a single fast access interface while automatically managing data placement between storage tiers based on age and access patterns
Solution Approach 2:
The system performs preliminary aggregation and indexing of historical data during the data loading and ETL processes, preparing optimized query structures in advance so that historical data can be accessed efficiently without requiring real-time computation or complex data movement during query execution
Data Source
AI summary
A computer implemented method, apparatus, and computer usable program code to identify a policy for managing data in a data storage system. Raw data is located in the data storage system for processing to form located data. The located data is aggregated based on the policy to form aggregated data. The aggregated data is stored in the data storage system.


