Two-Dimensional Facet Cube Clustering for Text Mining
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current techniques for analyzing unstructured data, such as text in documents, lack effective methods for automatically deriving facet values and discovering relationships between them, leading to inefficiencies in data analysis.
Innovation Solution
A computer-implemented method and system that generates a two-dimensional facet cube as a correlation matrix for text mining, allowing for clustering of facets and identifying representative facets through correlation analysis, thereby automatically deriving facet values and discovering relationships between them.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual facet value addition is used, then facet hierarchy can be built, but the process is time-consuming and labor-intensive
Solution Approach 1:
The system automatically derives facet values by performing text mining on documents themselves, eliminating the need for manual facet value addition. The documents serve their own analysis needs by providing the raw text that is mined to generate facet values automatically.
Solution Approach 2:
The patent replaces the manual mechanical process of adding facet values with an automated text mining system that uses correlation analysis and clustering algorithms to derive facet values programmatically from document content.
2Loss of information
If traditional text analysis methods are used, then basic text processing can be performed, but relationship discovery between facet values is limited
Solution Approach 1:
The patent combines multiple analysis functions into a unified correlation matrix that simultaneously captures relationships between all facet values. By merging the correlation analysis and clustering operations, the system discovers relationships that would be missed by separate analysis methods.
Solution Approach 2:
The patent transforms the analysis from traditional single-dimensional text processing to multi-dimensional correlation analysis by creating a facet cube with multiple dimensions representing different facet values and their interrelationships, enabling comprehensive relationship discovery.
3Measurement precision
If comprehensive text mining is performed, then accurate facet values can be derived, but the computational workload increases
Solution Approach 1:
The patent segments the comprehensive text mining process into distinct phases: correlation calculation between facet values, clustering based on correlation thresholds, and representative facet selection. This segmentation allows each component to be optimized independently while maintaining overall accuracy.
Solution Approach 2:
The patent uses correlation thresholds to perform partial analysis - only computing and clustering correlations that exceed certain thresholds, rather than analyzing all possible facet value pairs. This reduces computational complexity while maintaining the precision needed for accurate facet value derivation.
Data Source
AI summary
A computer-implemented method and system for clustering facets on a two-dimensional facet cube for text mining. The method and system performs text mining based on facets to analyze unstructured data in one or more documents by generating a two-dimensional facet cube that is a correlation matrix for one or more facets associated with a set of one or more of the documents; grouping one or more of the facets in the correlation matrix into at least one cluster; calculating a center for the cluster; and identifying facets that are located near the calculated center of the cluster as being representative of the cluster.


