Clustering Space Visualization for Text Data Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Classifying large volumes of unstructured text into optimal clusters is a computationally challenging task due to the vast number of possible partitions and similarities that need to be assessed, exceeding human capabilities and existing algorithm limitations, with no universally applicable clustering method across diverse data sets.

Innovation Solution

A method involving multiple known clustering algorithms applied sequentially to generate a metric space, projected to a lower dimension for visualization, using a local cluster ensemble approach to create new clusterings and animated visualization to aid in selecting the most informative clusterings, incorporating techniques like Sammon multidimensional scaling and weighted averaging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple clustering methods are applied to generate comprehensive clusterings, then the coverage of clustering space is improved, but the computational complexity increases

Engineering Contradiction:
Improvecoverage of clustering spaceVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the vast clustering space into multiple regions using a small set of representative clustering methods. Each method explores a specific region of the clustering space, and the results are combined to achieve comprehensive coverage without exhaustively searching all possible clusterings. This segmentation approach reduces computational complexity while maintaining versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal framework that combines results from multiple clustering methods into a unified clustering space representation. This multi-functional system allows different clustering methods to work together, with each contributing to different aspects of the overall clustering analysis, achieving comprehensive coverage through coordinated effort rather than individual exhaustive search.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If the number of clusters is increased to capture more granular patterns, then the classification precision is improved, but the difficulty of selecting meaningful clusterings increases

Engineering Contradiction:
Improveclassification precisionVSAvoiddifficulty of selecting meaningful clusterings
Core Design Contradiction:
Measurement precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The patent adds a new dimension to the clustering analysis by introducing a visual mapping space where clusterings are represented as points. This dimensional transformation allows researchers to navigate and select meaningful clusterings based on their position in the visual space rather than manually evaluating numerous clusterings, making the selection process more intuitive despite increased granularity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an intermediary visual mapping system that mediates between the complex clustering results and human interpretation. This intermediary layer transforms detailed clustering data into a manageable visual representation, allowing researchers to select meaningful clusterings through visual inspection rather than direct analysis of complex clustering parameters.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If traditional clustering methods are used without domain knowledge, then the applicability across different data sets is improved, but the classification accuracy deteriorates

Engineering Contradiction:
Improveapplicability across data setsVSAvoidclassification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by applying multiple clustering methods before the actual analysis. This preliminary step generates a diverse set of clusterings that capture different aspects of the data, creating a foundation that can be adapted to various domains. The preliminary clusterings serve as starting points that can be refined based on domain-specific requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a composite approach by combining results from multiple clustering methods into a unified framework. Similar to composite materials that combine different substances to achieve superior properties, this composite clustering framework combines strengths of different methods to achieve both broad applicability and high classification accuracy across diverse data sets.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS8438162B2Method and apparatus for selecting clusterings to classify a predetermined data set
Publication Date: 2013.05.07 PRESIDENT & FELLOWS OF HARVARD COLLEGE
  • US8438162B2 patent drawing
  • US8438162B2 patent drawing
  • US8438162B2 patent drawing

AI summary

A method for selecting clusterings to classify a predetermined data set of numerical data comprises five steps. First, a plurality of known clustering methods are applied, one at a time, to the data set to generate clusterings for each method. Second, a metric space of clusterings is generated using a metric that measures the similarity between two clusterings. Third, the metric space is projected to a lower dimensional representation useful for visualization. Fourth, a “local cluster ensemble” method generates a clustering for each point in the lower dimensional space. Fifth, an animated visualization method uses the output of the local cluster ensemble method to display the lower dimensional space and to allow a user to move around and explore the space of clustering.