Clustering Technique for Data Insight Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data clustering techniques often separate predictive analysis from input variables, failing to effectively detect strong relationships between input and output variables, which limits insight discovery in data analysis.
Innovation Solution
The approach involves compressing data into sub-clusters based on proximity values in a predictor space and merging these sub-clusters using a tightness factor in a target space to form sub-groups, allowing for the selection of sub-groups that provide insights into relationships between inputs and outputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is compressed into sub-clusters based on proximity values in predictor space, then data processing efficiency is improved, but the ability to detect strong relationships between input and output variables deteriorates
Solution Approach 1:
The patent segments data into sub-clusters based on proximity values in predictor space, then further segments these into sub-groups using tightness factors in target space. This hierarchical segmentation allows efficient processing of large datasets while maintaining the ability to detect relationships at multiple levels of granularity, resolving the contradiction between processing efficiency and relationship detection accuracy.
Solution Approach 2:
The patent transitions from one-dimensional clustering in predictor space to multi-dimensional analysis by incorporating target space tightness factors. This dimensional expansion allows the system to maintain processing efficiency through initial compression while detecting relationships through additional dimensional analysis, effectively resolving the contradiction.
2Loss of information
If hierarchical clustering techniques are used to generate subgroups, then insight extraction is improved, but computational complexity increases
Solution Approach 1:
The patent applies preliminary compression of data into sub-clusters based on proximity values before applying hierarchical clustering. This preliminary action reduces the data volume that requires complex hierarchical processing, thereby maintaining high-quality insight extraction while reducing computational complexity.
Solution Approach 2:
The patent applies different clustering strategies to different portions of the data - using proximity-based compression for overall structure and tightness-based hierarchical clustering for local relationship detection. This localized application of different methods optimizes both insight extraction quality and computational efficiency.
3Measurement precision
If sub-clusters are merged based on tightness values, then statistical inference accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary compression into sub-clusters using proximity values, which reduces the number of data points that require tightness-based merging. This preliminary action maintains statistical inference accuracy by preserving meaningful groupings while reducing the processing time required for subsequent tightness-based merging operations.
Data Source
AI summary
Disclosed aspects relate to data insight discovery using a clustering technique. A set of data may be compressed based on a set of proximity values with respect to a set of predictors to assemble a set of sub-clusters. A set of subgroups may be established by merging a plurality of individual sub-clusters of the set of sub-clusters using a tightness factor. A subset of the subgroups may be selected based on a selection criterion. A set of insight data which indicates a profile of the subset of the set of subgroups with respect to the set of data may be compiled for the subset of the set of subgroups.


