Multidimensional Correlated Data Extraction via Vector Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to find correlations among multiple dimensions in large datasets, such as purchase data, due to the complexity of setting and verifying hypotheses, making it difficult to identify correlations between elements like beef meat and detergent, or among three or more elements, especially in big data analysis where correlations are hidden and vast.
Innovation Solution
A multidimensional correlated data extracting device and method that uses a similarity indexing function to acquire a similarity index, a subset extracting function to identify correlated data, and a dimension extracting function to find featured dimensions with strong correlations, along with a map generating function to visualize these correlations in a multidimensional space.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If hypothesis verification methods are used to find correlations, then two-element correlations can be found, but multidimensional correlations among three or more elements cannot be effectively identified
Solution Approach 1:
The patent transforms the correlation analysis from traditional two-dimensional pairwise comparison to multidimensional space analysis. By representing data as vectors in multidimensional space and calculating similarity indices based on angular relationships, the system can simultaneously analyze correlations among three or more elements, overcoming the limitation of conventional hypothesis verification methods that only detect pairwise correlations.
Solution Approach 2:
The patent introduces a similarity index parameter that measures the angular relationship between data vectors in multidimensional space. This parameter transformation allows the system to detect correlations among multiple elements by evaluating the alignment of vectors rather than relying on traditional correlation coefficients that are limited to pairwise comparisons.
2Quantity of substance
If traditional analysis methods are applied to big data, then data processing can be performed, but correlations hidden in vast data amounts cannot be found
Solution Approach 1:
The patent extracts the essential correlation information from vast amounts of big data by projecting data vectors into a similarity space. Instead of analyzing all possible element combinations, the system extracts the angular relationships between data vectors, which reveals hidden correlations even in large datasets without requiring exhaustive analysis of every data point.
Solution Approach 2:
By transforming data into vector representations in multidimensional space and analyzing angular relationships, the patent enables effective correlation detection in big data. This dimensional transformation allows the system to process large volumes of data while maintaining the ability to detect subtle correlations that would be invisible in traditional analysis approaches.
3Loss of information
If all possible element combinations are analyzed, then complete correlation information can be obtained, but the complexity and time required increase rapidly
Solution Approach 1:
The patent merges multiple pairwise correlation analyses into a single multidimensional similarity assessment. By representing multiple elements as vectors in the same space and evaluating their mutual angular relationships, the system obtains complete correlation information among three or more elements simultaneously, rather than analyzing each combination separately, thus dramatically reducing analysis time while maintaining information completeness.
Data Source
AI summary
A method of extracting subsets from the whole population of data configured by values of many elements in a case where the subsets have a correlation for a plurality of elements and finding out correlated elements. More specifically, the method comprises modeling data as a vector based on values of all the elements configuring individual data, and, in a multidimensional space in which all the data included in the population is plotted, extracting subsets each having a multidimensional correlation based on the densities of plots, and finding out featured elements having a correlation in the subsets.


