Autogrouping Data Partitioning for Graph Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing large multidimensional datasets are inefficient and fail to identify important relationships, requiring sophisticated experts and being computationally intensive, with previous methods like clustering, linear regression, and principal component analysis being too sensitive to large scale distances and losing detail.
Innovation Solution
The system employs autogrouping techniques using a scoring function to partition data sets, generating a report that identifies exclusive subsets and visualizes relationships through an interactive visualization tool, allowing for exploratory data analysis and revealing structural patterns in data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If clustering methods are used to analyze large multidimensional datasets, then data can be grouped into categories, but the method is too blunt an instrument to identify important relationships and loses detail
Solution Approach 1:
The patent segments the data analysis process into multiple hierarchical levels, where data is progressively grouped from fine-grained to coarse-grained clusters. This segmentation allows identification of important relationships at multiple scales, preventing loss of detail while maintaining analytical efficiency.
Solution Approach 2:
The patent introduces a hierarchical dimension to traditional clustering by organizing clusters into a tree structure with multiple levels. This additional dimensional organization allows preservation of fine-grained relationships while enabling efficient high-level analysis, resolving the contradiction between detail preservation and analysis efficiency.
2Productivity
If traditional data analysis methods are used, then analysis can be performed on large datasets, but sophisticated experts are necessary to interpret and understand the output
Solution Approach 1:
The patent implements self-organizing maps that automatically structure and label clusters based on the data itself, without requiring expert intervention for interpretation. The system serves itself by generating human-readable descriptions and hierarchical organization of results, making complex data analysis accessible to non-experts.
Solution Approach 2:
The patent introduces an intermediary layer between raw data and human interpretation in the form of automatically generated cluster labels, descriptions, and hierarchical structures. This intermediary translates complex multidimensional relationships into intuitive visual and textual representations that non-experts can understand.
3Loss of information
If previous analysis methods are used, then graphs depicting relationships can be generated, but the graphs are not interactive and require considerable time for experts to understand
Solution Approach 1:
The patent transforms static relationship graphs into dynamic, interactive visualizations where users can navigate hierarchical levels, drill down into specific clusters, and explore relationships interactively. This dynamic presentation allows rapid understanding of relationships without requiring considerable analysis time.
Solution Approach 2:
The patent creates a multi-functional visualization system that serves multiple purposes: displaying hierarchical cluster structures, enabling interactive exploration, providing automated interpretations, and supporting both expert and non-expert users. This universal system replaces multiple separate analysis tools with a single integrated platform.
4Productivity
If existing analysis methods are used, then data can be processed, but the methods are too sensitive to large scale distances and lose detail
Solution Approach 1:
The patent segments the data space into hierarchical clusters that progressively group data points from local to global scales. This segmentation allows preservation of fine-grained details within local clusters while simultaneously capturing large-scale relationships through the hierarchical structure, preventing sensitivity to scale distances.
Solution Approach 2:
The patent applies local quality by allowing different clusters at different hierarchical levels to have different properties and resolutions. Local clusters preserve fine-grained detail with high resolution, while higher-level clusters provide coarse-grained overview, with each level optimized for its specific scale and purpose.
Data Source
AI summary
Autogrouping is described. An example method includes receiving a data set, building a first partition of subsets of the data set, computing a first subset score for each subset using a scoring function, generating a next partition including at least one subset that includes the elements of two or more subsets of the first partition, computing a second subset score for each subset of the next partition using the scoring function, defining a max score for each particular subset using a max score function, each max score being based on maximal subset scores of that particular subset and at least the subsets of the first partition related to that particular subset, selecting output subsets, selection of each of the output subsets being made using a maximum score of previously computed subset scores, and generating a report indicating an output partition, the output subsets being associated with the received data set.


