Document Clustering System Using Reduced Term Space for Theme Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual classification of information is inefficient for unstructured data, and automatic classification techniques often misclassify documents due to ambiguous terms, leading to irrelevant results.

Innovation Solution

A clustering system processes documents using preprocessing and clustering engines to create a term-by-document matrix, applying Latent Semantic Indexing and singular value decomposition to identify clusters and themes, with a reduced term space and re-clustering to improve accuracy and disambiguate terms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification is performed, then classification accuracy is improved, but productivity deteriorates

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system enables automatic self-classification of documents through clustering algorithms that analyze term frequencies and document similarities without human intervention. The clustering engine autonomously organizes documents into thematic groups based on computational analysis of their content characteristics.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical classification processes with computational algorithms. The system uses automated text analysis, term frequency calculation, and clustering computations to substitute human classifiers, thereby maintaining accuracy while dramatically improving processing speed and productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If automatic classification using most frequent term is applied, then productivity is improved, but measurement precision deteriorates

Engineering Contradiction:
Improveclassification efficiencyVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system introduces clustering analysis as an intermediary between simple term frequency counting and final classification. Instead of directly using the most frequent term, the system first groups documents into clusters based on multiple term frequencies and similarities, then determines classification from these clusters, thereby improving accuracy while maintaining automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent moves from one-dimensional single-term classification to multi-dimensional clustering analysis. By considering multiple terms simultaneously and analyzing document relationships in a higher-dimensional space of term frequencies and similarities, the system achieves more accurate classification without sacrificing productivity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of operation

If term frequency-based classification is used, then ease of operation is improved, but reliability deteriorates due to ambiguous terms

Engineering Contradiction:
Improvesimplicity of classificationVSAvoidclassification consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system segments the classification process into distinct stages: term frequency extraction, document clustering, and theme determination. This segmentation allows the system to handle ambiguous terms by analyzing them in the context of cluster memberships rather than relying on single-term decisions, thereby improving reliability while maintaining operational simplicity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary clustering analysis before final classification determination. By pre-grouping documents into clusters based on overall term frequency patterns and similarities, the system resolves ambiguities in advance, ensuring that subsequent classification decisions are based on contextualized cluster information rather than isolated ambiguous terms.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS8886651B1Thematic clustering
Publication Date: 2014.11.11 REPUTATION COM
  • US8886651B1 patent drawing
  • US8886651B1 patent drawing
  • US8886651B1 patent drawing

AI summary

A data set is clustered into one or more initial clusters using a first term space. Initial themes for the initial clusters are determined. The first term space is reduced to create a reduced term space. At least a portion of the data set is reclustered into one or more baby clusters using the reduced term space. One or more singletons are reassigned to form one or more renovated clusters from the baby clusters. A renovated theme is determined for at least some of the renovated clusters. One or more of the renovated clusters and their respective themes are provided as output.