Data Cube Granularity Selection Automation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in data visualization is selecting the appropriate granularity for data aggregation in data cubes, as existing methods require manual input of rules, which is cumbersome and often results in suboptimal granularities, making it difficult to clearly show inherent rules in transaction data.
Innovation Solution
A computer-implemented method selects a candidate granularity from a plurality of options based on predetermined conditions, such as periodicity, distinction degree, or correlation, to ensure the data distribution satisfies specific criteria, thereby automatically generating a data cube that reflects the underlying data rules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual rule-based method is used to select granularities, then user control over data cube generation is maintained, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The system automatically evaluates candidate granularities using predefined evaluation functions that assess data distribution characteristics, eliminating the need for manual rule input. The system serves itself by autonomously selecting optimal granularities based on statistical analysis of the data patterns.
Solution Approach 2:
The system changes the approach from manual parameter specification to automated parameter optimization by evaluating multiple candidate granularities using computational metrics. The evaluation functions compute statistical parameters such as data distribution entropy and pattern recognition scores to automatically determine optimal granularity settings.
2Extent of automation
If user sets rules for granularity selection, then some level of automation is achieved, but the results may be suboptimal if the user lacks experience
Solution Approach 1:
The system employs evaluation functions that provide feedback on the quality of each candidate granularity by analyzing data distribution patterns. The feedback mechanism computes metrics such as information entropy, pattern distinctiveness, and statistical significance to guide the automatic selection process toward optimal granularities.
Solution Approach 2:
The patent replaces the mechanical system of manual rule-setting with an automated computational system that uses statistical analysis and evaluation functions. This substitution eliminates human subjectivity and experience-dependent variations, providing consistent and objectively optimal granularity selection.
3Productivity
If improper granularity is selected, then data aggregation is performed, but the inherent rules in transaction data are not clearly shown
Solution Approach 1:
The system performs preliminary evaluation of candidate granularities before final data aggregation by computing evaluation metrics on sample data distributions. This preliminary action identifies the optimal granularity that best preserves inherent data patterns, ensuring that subsequent aggregation does not lose important information.
Solution Approach 2:
The system generates and evaluates multiple candidate granularity configurations, using each only for evaluation purposes before discarding them. This disposable approach allows extensive exploration of granularity options without committing to suboptimal choices, ensuring the final selection maximizes information preservation.
Data Source
AI summary
Disclosed are a computer-implemented method for generating a data cube from data, a system and a computer program product. The method comprises selecting a candidate granularity from a plurality of candidate granularities determined for a dimension of the data cube, where a data distribution obtained in the selected candidate granularity satisfies a predetermined condition; and generating the data cube based on the selected candidate granularity for the dimension.


