Bulky Data Analysis via Cluster Center Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for analyzing bulky data from experiments, such as crash-test simulations, are computationally intensive and unable to handle the vast amounts of data produced, leading to time-consuming and often inaccurate results, failing to efficiently perform tasks like scatter analysis, causal analysis, and sensitivity analysis simultaneously.
Innovation Solution
An apparatus and method utilizing a cluster analyzer, correlation determiner, and data analyzer to reduce data complexity by determining cluster centers and calculating correlation values, allowing for the analysis of only strongly correlated data items, thereby accelerating and enhancing the precision of data analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional data analysis methods are used on bulky experimental data, then comprehensive analysis can be performed, but the computational complexity and processing time become excessively high
Solution Approach 1:
The patent extracts only the most relevant information from bulky data by identifying and analyzing cluster centers that represent groups of similar data items. Instead of processing all millions of data items, the method extracts a small subset of representative cluster centers, significantly reducing computational complexity while maintaining analysis precision.
Solution Approach 2:
The patent creates simplified representations (copies) of the original data through cluster centers. Each cluster center serves as a representative copy that captures the essential characteristics of multiple similar data items, allowing analysis to be performed on these simplified copies rather than the full dataset.
2Measurement precision
If traditional data analysis methods are used on bulky experimental data, then complete data coverage is achieved, but the processing time becomes unacceptably long
Solution Approach 1:
The patent segments the bulky data into multiple clusters based on similarity, where each cluster represents a distinct group of data items with comparable characteristics. By analyzing one representative from each cluster rather than all items individually, the method maintains comprehensive coverage of data patterns while dramatically reducing processing time.
Solution Approach 2:
The patent performs partial action by analyzing only a subset of data items (cluster centers) rather than the complete dataset. This selective approach is sufficient to capture all essential patterns and trends in the data, achieving analysis completeness without the need to process every single data item.
3Loss of information
If all data items are analyzed individually, then detailed insights can be obtained, but the scalability to handle massive datasets is lost
Solution Approach 1:
The patent merges multiple similar data items into clusters, where each cluster is represented by a single cluster center. This merging process preserves the essential information and patterns from all individual items while enabling efficient processing at scale, as the same analytical operations can be applied to cluster centers representing potentially millions of original data items.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus for analyzing bulky data comprises a cluster analyzer, a correlation determiner and a data analyzer. The bulky data comprises a plurality of data items, each data item comprises a data value for each experiment of a plurality of experiments by which the bulky data was obtained. The cluster analyzer determines a cluster center of a cluster of data items based on a cluster analysis of a plurality of data items. Further, the correlation determiner calculates a correlation value between the determined cluster center and each data item of the plurality of data items. A correlation value indicates a correlation strength of the cluster center and a data item. Additionally, the data analyzer analyzes the bulky data based on calculated correlation values.