Dimension Candidate Scoring for Hierarchical Data Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users of data analytics software face overwhelming complexity when analyzing large and hierarchical datasets, making it difficult to determine which information is most important for decision-making.
Innovation Solution
A computer-implemented method that identifies and presents only the most influential 'dimension candidates' within each dimension of a dataset, focusing on statistical analysis to provide insights by determining a limited set of top contributors, reducing data overload and enhancing decision-making capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data analytics software provides comprehensive analysis of large datasets, then the volume of information available increases, but the complexity and overwhelming nature of the data increases
Solution Approach 1:
The patent extracts only the most influential dimension candidates from the complete set of dimensions. The system identifies and presents a limited subset of top contributors (e.g., top 5-10 dimensions) that have the greatest impact on the numerical measure, filtering out less relevant dimensions to reduce complexity while preserving essential information.
Solution Approach 2:
The system changes the parameter of dimension selection from providing all dimensions to providing only those above a certain influence threshold. By calculating influence scores based on statistical measures (variance, correlation, etc.) and filtering dimensions based on these scores, the system transforms the data presentation from comprehensive to optimized.
2Loss of information
If data analytics software provides comprehensive analysis of all dimensions, then complete information is available, but user ability to determine important information decreases
Solution Approach 1:
The patent applies local quality by providing different levels of information detail to different users or contexts. The system identifies which dimensions have local (specific) high influence on the measure and presents those prominently, while less influential dimensions are either omitted or presented with lower priority, allowing users to focus on locally important information.
Solution Approach 2:
The system changes the presentation parameter from equal weighting of all dimensions to differentiated weighting based on influence scores. Dimensions are ranked and presented according to their calculated influence, transforming the information structure from uniform to prioritized, thereby improving ease of identifying important information.
3Quantity of substance
If all dimension candidates are presented to users, then complete data coverage is achieved, but user workload increases
Solution Approach 1:
The system extracts and presents only the most relevant dimension candidates based on influence scoring. By automatically filtering out dimensions with low influence scores, the system reduces the number of dimensions users must review while maintaining coverage of the most important data aspects, thereby reducing user workload time.
Solution Approach 2:
The system performs preliminary analysis to calculate influence scores for all dimensions before presenting them to users. This pre-processing step automatically ranks dimensions by importance, so users receive pre-organized information that requires less manual analysis and sorting, reducing their workload time.
4Loss of information
If comprehensive statistical analysis is performed on all dimensions, then complete insights are generated, but processing time increases
Solution Approach 1:
The system extracts and performs detailed statistical analysis only on the most influential dimension candidates identified through initial screening. By applying comprehensive analysis selectively to top-ranked dimensions rather than all dimensions, the system maintains insight completeness for the most important factors while reducing overall processing time.
Solution Approach 2:
The system performs partial analysis by focusing computational resources on the subset of dimensions with highest influence scores. Rather than performing equal-depth analysis on all dimensions, the system applies more rigorous statistical methods to the most critical dimensions, achieving sufficient insight with reduced processing time.
Data Source
AI summary
The present disclosure pertains to processing a data model having hierarchical data. A front-end computer sends a request to a back-end computer for dimension candidates for the data model, which is visualized by the front-end computer. The front-end computer is configured score and rank such dimension candidates in order to determine statistics from the data model. The back-end computer determines dimension candidates based on their cardinality and hierarchical information and sends the dimension candidates to the front-end computer. The front-end computer performs the scoring and ranking of dimension candidates and determines statistics for a set of the dimension candidates. The statistics may be presented to a user along with charts and graphs representing the data model.


