Autogrouping Data Partitioning for Graph Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for analyzing large multidimensional datasets are inefficient and fail to identify important relationships, requiring sophisticated experts and being computationally intensive, with previous methods like clustering, linear regression, and principal component analysis being too sensitive to large scale distances and losing detail.

Innovation Solution

The system employs autogrouping techniques using a scoring function to partition data sets, generating a report that identifies exclusive subsets and visualizes relationships through an interactive visualization tool, allowing for exploratory data analysis and revealing structural patterns in data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If clustering methods are used to analyze large multidimensional datasets, then data can be grouped into categories, but the method is too blunt an instrument to identify important relationships and loses detail

Engineering Contradiction:
Improvedata analysis efficiencyVSAvoidrelationship identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the data analysis process into multiple hierarchical levels, where data is progressively grouped from fine-grained to coarse-grained clusters. This segmentation allows identification of important relationships at multiple scales, preventing loss of detail while maintaining analytical efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a hierarchical dimension to traditional clustering by organizing clusters into a tree structure with multiple levels. This additional dimensional organization allows preservation of fine-grained relationships while enabling efficient high-level analysis, resolving the contradiction between detail preservation and analysis efficiency.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If traditional data analysis methods are used, then analysis can be performed on large datasets, but sophisticated experts are necessary to interpret and understand the output

Engineering Contradiction:
Improvedata analysis capabilityVSAvoidinterpretation complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements self-organizing maps that automatically structure and label clusters based on the data itself, without requiring expert intervention for interpretation. The system serves itself by generating human-readable descriptions and hierarchical organization of results, making complex data analysis accessible to non-experts.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces an intermediary layer between raw data and human interpretation in the form of automatically generated cluster labels, descriptions, and hierarchical structures. This intermediary translates complex multidimensional relationships into intuitive visual and textual representations that non-experts can understand.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If previous analysis methods are used, then graphs depicting relationships can be generated, but the graphs are not interactive and require considerable time for experts to understand

Engineering Contradiction:
Improverelationship visualizationVSAvoidanalysis time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent transforms static relationship graphs into dynamic, interactive visualizations where users can navigate hierarchical levels, drill down into specific clusters, and explore relationships interactively. This dynamic presentation allows rapid understanding of relationships without requiring considerable analysis time.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a multi-functional visualization system that serves multiple purposes: displaying hierarchical cluster structures, enabling interactive exploration, providing automated interpretations, and supporting both expert and non-expert users. This universal system replaces multiple separate analysis tools with a single integrated platform.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If existing analysis methods are used, then data can be processed, but the methods are too sensitive to large scale distances and lose detail

Engineering Contradiction:
Improvedata processing capabilityVSAvoiddetail preservation
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the data space into hierarchical clusters that progressively group data points from local to global scales. This segmentation allows preservation of fine-grained details within local clusters while simultaneously capturing large-scale relationships through the hierarchical structure, preventing sensitivity to scale distances.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by allowing different clusters at different hierarchical levels to have different properties and resolutions. Local clusters preserve fine-grained detail with high resolution, while higher-level clusters provide coarse-grained overview, with each level optimized for its specific scale and purpose.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10599669B2Grouping of data points in data analysis for graph generation
Publication Date: 2020.03.24 SYMPHONYAI SENSA LLC
  • US10599669B2 patent drawing
  • US10599669B2 patent drawing
  • US10599669B2 patent drawing

AI summary

Autogrouping is described. An example method includes receiving a data set, building a first partition of subsets of the data set, computing a first subset score for each subset using a scoring function, generating a next partition including at least one subset that includes the elements of two or more subsets of the first partition, computing a second subset score for each subset of the next partition using the scoring function, defining a max score for each particular subset using a max score function, each max score being based on maximal subset scores of that particular subset and at least the subsets of the first partition related to that particular subset, selecting output subsets, selection of each of the output subsets being made using a maximum score of previously computed subset scores, and generating a report indicating an output partition, the output subsets being associated with the received data set.