Topological Data Analysis for Multidimensional Dataset Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for analyzing large multidimensional datasets are insufficient in identifying important relationships, often breaking relationships, being computationally inefficient, and requiring sophisticated experts for interpretation, with outputs that are not interactive or exploratory.
Innovation Solution
A method involving topological data analysis (TDA) that maps data points from fact and dimension tables to a reference space, clusters them, and generates a segment data structure to identify significant dimensions and values, enabling interactive visualization and further analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional clustering methods are used to analyze large multidimensional datasets, then computational processing is performed, but important relationships in the data are broken and lost
Solution Approach 1:
The patent applies segmentation by dividing the data analysis process into distinct topological operations (filtering, covering, clustering) that preserve local relationships. Instead of treating the entire dataset as a single unit for clustering, the method segments the reference space into covered regions that maintain the connectivity and relationships among data points, thereby preventing the loss of important relationships while still enabling computational processing.
Solution Approach 2:
The patent introduces an intermediary reference space and covering mechanism between the original high-dimensional data and the final clustering result. This intermediary structure (the covered reference space with overlapping regions) acts as a mediator that preserves relationships during the transformation process, allowing data points to maintain their contextual connections while being processed computationally.
2Productivity
If traditional linear algebraic and analytic methods are used, then data processing is performed, but the methods are too sensitive to large scale distances and lose detail
Solution Approach 1:
The patent transforms the data analysis from direct high-dimensional space processing to a reference space with a covering structure. This dimensional transformation allows the method to handle large-scale distances in the reference space while preserving fine-grained details through the overlapping cover regions, effectively decoupling the sensitivity to scale from the detail retention capability.
Solution Approach 2:
The patent applies local quality by using overlapping cover regions in the reference space where different regions can have different resolution properties. This allows the method to maintain high measurement precision in local regions while handling global large-scale distances, as each covered region can be analyzed with appropriate detail without being affected by the overall scale of the entire dataset.
3Loss of information
If previous analysis methods are used, then graphs depicting relationships are generated, but the graphs are not interactive and require considerable time for expert interpretation
Solution Approach 1:
The patent introduces dynamics by making the data analysis system interactive and adaptable. The method allows users to modify analysis parameters, explore different coverings, and interactively adjust the reference space transformation, enabling rapid exploration of relationships without requiring time-consuming expert interpretation of static graphs. The system responds dynamically to user inputs and allows exploratory analysis.
Solution Approach 2:
The patent enables self-service by allowing users to directly interact with and modify the analysis parameters and explore the data relationships themselves, rather than requiring sophisticated experts to interpret complex static graphs. The interactive nature of the method empowers users to perform their own exploratory data analysis by adjusting parameters and observing results in real-time.
4Reliability
If previous analysis methods are used, then hypothesis testing is required, but exploratory data analysis where analysis can be quickly modified to discover new relationships is not allowed
Solution Approach 1:
The patent applies dynamics by enabling the analysis method to adapt quickly to different exploratory scenarios. Users can modify analysis parameters, change covering resolutions, and adjust reference space transformations interactively, allowing the system to support both rigorous hypothesis testing and flexible exploratory analysis without requiring pre-defined hypotheses. The dynamic nature allows rapid modification of analysis approaches to discover new relationships.
Data Source
AI summary
A method comprises receiving a selection of data from a fact table and one or more dimension tables stored in a data warehouse, mapping data points from the selection of the data from the fact table and the one or more dimension tables to a reference space utilizing a lens function, generating a cover of the reference space using a resolution function, clustering the data points mapped to the reference space using the cover and a metric function to determine each node of a plurality of nodes of a graph, each node including at least one data point, determining a plurality of segments of the graph, each segment including at least one node, and generating a segment data structure identifying each segment as well as membership of each segment, the membership of each segment including at least one node from the plurality of nodes in the graph.


