Multivariate Data Analysis Using Bipartite Synthesis Matrix
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data analysis methods, particularly axis-based and graph theory approaches, face challenges with large, sparse multivariate datasets due to limitations in handling noise, missing data, non-linearity, and computational efficiency, especially in establishing distance metrics and visualizing high-dimensional data.
Innovation Solution
A computer-implemented method using a bipartite matrix with multi-granular data aggregation to store and analyze multivariate data, generating adjacency matrices, and rendering them as ordinary graphs to establish path-independent distance metrics, enabling efficient data processing and visualization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If axis-based virtual coordinate assignment protocol is used to store and analyze multivariate data, then data storage is compact and relationships can be visualized, but the method becomes computationally problematic for large datasets and loses or distorts information when compressing dimensions
Solution Approach 1:
The patent segments the data analysis process into multiple granular levels, organizing data hierarchically from fine-grained individual data points to coarser-grained aggregated structures. This segmentation allows preservation of detailed information at lower levels while enabling efficient analysis at higher levels, avoiding the information loss inherent in traditional dimension compression.
Solution Approach 2:
The patent introduces a new dimensional framework that transcends traditional axis-based coordinate systems. By creating a multi-granular hierarchical structure that adds a granularity dimension, the system can represent data relationships without compressing existing dimensions, thereby preserving information while enabling efficient analysis of large multivariate datasets.
2Measurement precision
If traditional axis-based systems are used for data analysis, then relative position and distance measurements can be made, but path-dependent calculations become computationally problematic and uncertainty in relating data must be accounted for
Solution Approach 1:
The patent performs preliminary aggregation of data at multiple granular levels before analysis, pre-computing relationships and structures that would otherwise require complex path-dependent calculations. This preliminary organization of data hierarchically eliminates the need for computationally intensive path-dependent operations while maintaining measurement precision.
3Reliability
If high-dimensional data sets are analyzed using smooth axis-based systems, then interpolation or prediction can be performed, but large amounts of data collection are required to achieve statistically valid results
Solution Approach 1:
The patent performs preliminary aggregation at multiple granular levels, pre-computing statistical relationships and patterns from the data hierarchy. This preliminary processing enables reliable interpolation and prediction with fewer data points, as the aggregated structures capture essential patterns that would otherwise require extensive data collection to establish statistically.
4Productivity
If conventional computers are used to find hidden structures in extremely large datasets, then analysis may be performed, but it takes excessive time or may not be possible at all due to limited hardware resources
Solution Approach 1:
The patent segments large datasets into hierarchical structures organized by granularity level, enabling analysis to proceed from aggregated higher-level patterns to detailed lower-level structures. This segmentation allows conventional computers to analyze extremely large datasets efficiently by working with compressed representations at each level, dramatically reducing computation time while preserving the ability to discover hidden structures.
Data Source
AI summary
This invention is a computerized method which unites a multivariate dataset and then performs various operations, including data analytics. The set is stored in a “bipartite synthesis matrix” (BSM), e.g., a rectangular matrix with rows of data objects and columns of variable attributes, defined by a plurality of partitions (each with a numerical range and a characteristic scale). Links within the matrix between data objects and attribute(s) are based on shared correspondences within partitions. The process exploits mode reduction in which shared correspondences of a BSM (or its graph) interrelate data objects by producing an adjacency matrix or its associated graph. The partition scale is repeatedly and incrementally altered, varying the density of shared correspondences within the data, based on partition number and size; therefore, a fully connected and weighted unipartite network may be established. Shared correspondences' given scale and variable attribute provide distance metrics for edges within the network.


