Clustering Space Embedding for Document Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Classifying and clustering large volumes of unstructured text is a computationally challenging task due to the vast number of possible partitions and similarities between documents, which exceeds human capabilities and existing algorithms often require significant human intervention or sacrifice optimization for simplicity.

Innovation Solution

A computer-assisted clustering method that generates a two-dimensional clustering space by isometrically embedding the space of all possible clusterings, sampling, and creating partitions that tessellate the space, allowing users to explore and visualize clusterings, with features like landmark multidimensional scaling and animated visualization to facilitate the selection of informative clusterings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If fully automatic clustering algorithms are used to partition documents, then computational efficiency is improved, but the ability to produce insightful and useful clusterings deteriorates

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidinsightful information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent introduces computer-assisted clustering as an intermediary between fully automatic clustering and manual classification. The system presents multiple clustering options to users, who can then select and refine clusterings that provide the most insight. This intermediary approach combines the efficiency of automated algorithms with the discernment of human judgment, resolving the contradiction between computational speed and information quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If computer-assisted clustering methods are used to explore clustering space, then insightfulness of clusterings is improved, but user time investment increases

Engineering Contradiction:
Improveinsightful informationVSAvoiduser time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating multiple candidate clusterings using different algorithms and parameters before presenting them to the user. This pre-computation reduces the user's workload by eliminating the need to manually create and evaluate numerous clusterings, thereby improving insightfulness while minimizing the time users must invest in the exploration process.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the number of possible clusterings is increased to provide more options, then user choice is improved, but computational complexity increases

Engineering Contradiction:
Improveuser choiceVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the vast clustering space into manageable subsets by applying multiple clustering algorithms with different parameters to generate a diverse but有限 set of candidate clusterings. This segmentation provides users with meaningful choices across different clustering perspectives while avoiding the computational intractability of evaluating all possible clusterings, thus balancing user choice with computational feasibility.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9519705B2Method and apparatus for selecting clusterings to classify a data set
Publication Date: 2016.12.13 PRESIDENT & FELLOWS OF HARVARD COLLEGE
  • US9519705B2 patent drawing
  • US9519705B2 patent drawing
  • US9519705B2 patent drawing

AI summary

In a computer assisted clustering method, a clustering space is generated from fixed basis partitions that embed the entire space of all possible clusterings. A lower dimensional clustering space is fu-reated from the space of all possible clusterings by isometrically embedding the space of all possible clusterings in a lower dimensional Euclidean space. This lower dimensional space is then sampled based on the number of documents in the corpus. Partitions are then developed based on the samples that tessellate the space. Finally, using clusterings representative of these tessellations, a two-dimensional representation for users to explore is created.