Clustering Technique for Data Insight Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data clustering techniques often separate predictive analysis from input variables, failing to effectively detect strong relationships between input and output variables, which limits insight discovery in data analysis.

Innovation Solution

The approach involves compressing data into sub-clusters based on proximity values in a predictor space and merging these sub-clusters using a tightness factor in a target space to form sub-groups, allowing for the selection of sub-groups that provide insights into relationships between inputs and outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is compressed into sub-clusters based on proximity values in predictor space, then data processing efficiency is improved, but the ability to detect strong relationships between input and output variables deteriorates

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidrelationship detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments data into sub-clusters based on proximity values in predictor space, then further segments these into sub-groups using tightness factors in target space. This hierarchical segmentation allows efficient processing of large datasets while maintaining the ability to detect relationships at multiple levels of granularity, resolving the contradiction between processing efficiency and relationship detection accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from one-dimensional clustering in predictor space to multi-dimensional analysis by incorporating target space tightness factors. This dimensional expansion allows the system to maintain processing efficiency through initial compression while detecting relationships through additional dimensional analysis, effectively resolving the contradiction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If hierarchical clustering techniques are used to generate subgroups, then insight extraction is improved, but computational complexity increases

Engineering Contradiction:
Improveinsight extraction qualityVSAvoidcomputational complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies preliminary compression of data into sub-clusters based on proximity values before applying hierarchical clustering. This preliminary action reduces the data volume that requires complex hierarchical processing, thereby maintaining high-quality insight extraction while reducing computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different clustering strategies to different portions of the data - using proximity-based compression for overall structure and tightness-based hierarchical clustering for local relationship detection. This localized application of different methods optimizes both insight extraction quality and computational efficiency.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If sub-clusters are merged based on tightness values, then statistical inference accuracy is improved, but processing time increases

Engineering Contradiction:
Improvestatistical inference accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary compression into sub-clusters using proximity values, which reduces the number of data points that require tightness-based merging. This preliminary action maintains statistical inference accuracy by preserving meaningful groupings while reducing the processing time required for subsequent tightness-based merging operations.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10579663B2Data insight discovery using a clustering technique
Publication Date: 2020.03.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10579663B2 patent drawing
  • US10579663B2 patent drawing
  • US10579663B2 patent drawing

AI summary

Disclosed aspects relate to data insight discovery using a clustering technique. A set of data may be compressed based on a set of proximity values with respect to a set of predictors to assemble a set of sub-clusters. A set of subgroups may be established by merging a plurality of individual sub-clusters of the set of sub-clusters using a tightness factor. A subset of the subgroups may be selected based on a selection criterion. A set of insight data which indicates a profile of the subset of the set of subgroups with respect to the set of data may be compiled for the subset of the set of subgroups.