Multivariate Data Analysis Using Bipartite Synthesis Matrix

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data analysis methods, particularly axis-based and graph theory approaches, face challenges with large, sparse multivariate datasets due to limitations in handling noise, missing data, non-linearity, and computational efficiency, especially in establishing distance metrics and visualizing high-dimensional data.

Innovation Solution

A computer-implemented method using a bipartite matrix with multi-granular data aggregation to store and analyze multivariate data, generating adjacency matrices, and rendering them as ordinary graphs to establish path-independent distance metrics, enabling efficient data processing and visualization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If axis-based virtual coordinate assignment protocol is used to store and analyze multivariate data, then data storage is compact and relationships can be visualized, but the method becomes computationally problematic for large datasets and loses or distorts information when compressing dimensions

Engineering Contradiction:
Improvedata storage efficiencyVSAvoidinformation loss during dimension compression
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent segments the data analysis process into multiple granular levels, organizing data hierarchically from fine-grained individual data points to coarser-grained aggregated structures. This segmentation allows preservation of detailed information at lower levels while enabling efficient analysis at higher levels, avoiding the information loss inherent in traditional dimension compression.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimensional framework that transcends traditional axis-based coordinate systems. By creating a multi-granular hierarchical structure that adds a granularity dimension, the system can represent data relationships without compressing existing dimensions, thereby preserving information while enabling efficient analysis of large multivariate datasets.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If traditional axis-based systems are used for data analysis, then relative position and distance measurements can be made, but path-dependent calculations become computationally problematic and uncertainty in relating data must be accounted for

Engineering Contradiction:
Improvedistance measurement capabilityVSAvoidcomputational complexity of path-dependent calculations
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary aggregation of data at multiple granular levels before analysis, pre-computing relationships and structures that would otherwise require complex path-dependent calculations. This preliminary organization of data hierarchically eliminates the need for computationally intensive path-dependent operations while maintaining measurement precision.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If high-dimensional data sets are analyzed using smooth axis-based systems, then interpolation or prediction can be performed, but large amounts of data collection are required to achieve statistically valid results

Engineering Contradiction:
Improvestatistical validity of interpolationVSAvoiddata collection requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent performs preliminary aggregation at multiple granular levels, pre-computing statistical relationships and patterns from the data hierarchy. This preliminary processing enables reliable interpolation and prediction with fewer data points, as the aggregated structures capture essential patterns that would otherwise require extensive data collection to establish statistically.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If conventional computers are used to find hidden structures in extremely large datasets, then analysis may be performed, but it takes excessive time or may not be possible at all due to limited hardware resources

Engineering Contradiction:
Improvedata analysis throughputVSAvoidexcessive computation time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments large datasets into hierarchical structures organized by granularity level, enabling analysis to proceed from aggregated higher-level patterns to detailed lower-level structures. This segmentation allows conventional computers to analyze extremely large datasets efficiently by working with compressed representations at each level, dramatically reducing computation time while preserving the ability to discover hidden structures.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9424307B2Multivariate data analysis method
Publication Date: 2016.08.23 LILIENTHAL SCOTT E
  • US9424307B2 patent drawing
  • US9424307B2 patent drawing
  • US9424307B2 patent drawing

AI summary

This invention is a computerized method which unites a multivariate dataset and then performs various operations, including data analytics. The set is stored in a “bipartite synthesis matrix” (BSM), e.g., a rectangular matrix with rows of data objects and columns of variable attributes, defined by a plurality of partitions (each with a numerical range and a characteristic scale). Links within the matrix between data objects and attribute(s) are based on shared correspondences within partitions. The process exploits mode reduction in which shared correspondences of a BSM (or its graph) interrelate data objects by producing an adjacency matrix or its associated graph. The partition scale is repeatedly and incrementally altered, varying the density of shared correspondences within the data, based on partition number and size; therefore, a fully connected and weighted unipartite network may be established. Shared correspondences' given scale and variable attribute provide distance metrics for edges within the network.