Two-Dimensional Facet Cube Clustering for Text Mining

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for analyzing unstructured data, such as text in documents, lack effective methods for automatically deriving facet values and discovering relationships between them, leading to inefficiencies in data analysis.

Innovation Solution

A computer-implemented method and system that generates a two-dimensional facet cube as a correlation matrix for text mining, allowing for clustering of facets and identifying representative facets through correlation analysis, thereby automatically deriving facet values and discovering relationships between them.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual facet value addition is used, then facet hierarchy can be built, but the process is time-consuming and labor-intensive

Engineering Contradiction:
Improvefacet value derivation efficiencyVSAvoidtime for manual facet value addition
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system automatically derives facet values by performing text mining on documents themselves, eliminating the need for manual facet value addition. The documents serve their own analysis needs by providing the raw text that is mined to generate facet values automatically.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of adding facet values with an automated text mining system that uses correlation analysis and clustering algorithms to derive facet values programmatically from document content.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If traditional text analysis methods are used, then basic text processing can be performed, but relationship discovery between facet values is limited

Engineering Contradiction:
Improverelationship information between facet valuesVSAvoidcomplexity of correlation matrix and clustering system
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent combines multiple analysis functions into a unified correlation matrix that simultaneously captures relationships between all facet values. By merging the correlation analysis and clustering operations, the system discovers relationships that would be missed by separate analysis methods.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the analysis from traditional single-dimensional text processing to multi-dimensional correlation analysis by creating a facet cube with multiple dimensions representing different facet values and their interrelationships, enabling comprehensive relationship discovery.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If comprehensive text mining is performed, then accurate facet values can be derived, but the computational workload increases

Engineering Contradiction:
Improveaccuracy of facet value derivationVSAvoidcomputational complexity of text mining and clustering
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the comprehensive text mining process into distinct phases: correlation calculation between facet values, clustering based on correlation thresholds, and representative facet selection. This segmentation allows each component to be optimized independently while maintaining overall accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses correlation thresholds to perform partial analysis - only computing and clustering correlations that exceed certain thresholds, rather than analyzing all possible facet value pairs. This reduces computational complexity while maintaining the precision needed for accurate facet value derivation.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10657145B2Clustering facets on a two-dimensional facet cube for text mining
Publication Date: 2020.05.19 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10657145B2 patent drawing
  • US10657145B2 patent drawing
  • US10657145B2 patent drawing

AI summary

A computer-implemented method and system for clustering facets on a two-dimensional facet cube for text mining. The method and system performs text mining based on facets to analyze unstructured data in one or more documents by generating a two-dimensional facet cube that is a correlation matrix for one or more facets associated with a set of one or more of the documents; grouping one or more of the facets in the correlation matrix into at least one cluster; calculating a center for the cluster; and identifying facets that are located near the calculated center of the cluster as being representative of the cluster.