Document Classification Using Thesaurus Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional document classification methods face challenges in presenting document analysis results intelligibly due to the high number of document clusters and biased document distribution, particularly when using thesauruses for classification, which leads to a lack of versatility in viewpoint dictionaries and increased complexity in interpreting classification labels.

Innovation Solution

A document classification apparatus that extracts feature words from documents, clusters them into subtrees of a thesaurus to equalize document distribution, and assigns classification labels to clusters, allowing for intelligible presentation of results without relying on field-specific viewpoint dictionaries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If document clusters are classified using registered words on the same layer of thesaurus, then classification granularity is standardized, but the number of document clusters increases excessively

Engineering Contradiction:
Improveclassification granularityVSAvoidnumber of document clusters
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent merges multiple document clusters that share the same registered word on the thesaurus into a single integrated cluster. Instead of creating separate clusters for each registered word occurrence, the system combines them under one unified cluster, thereby reducing the total number of clusters while preserving the standardized granularity provided by the thesaurus structure.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If registered words from lower level concepts in thesaurus are used as classification labels, then classification precision is improved, but interpretability of results deteriorates

Engineering Contradiction:
Improveclassification precisionVSAvoidinterpretability
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements a nested labeling structure where classification labels are organized in hierarchical levels. The display label comes from a higher-level concept in the thesaurus hierarchy, providing broad interpretability, while the actual classification is performed using more specific lower-level registered words, maintaining precision. This nested approach allows users to see both the general category and the specific classification criteria.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Measurement precision

If field-specific viewpoint dictionaries are used for reputation analysis, then analysis accuracy is improved, but versatility of the system deteriorates

Engineering Contradiction:
Improveanalysis accuracyVSAvoidsystem versatility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent makes the viewpoint dictionary field-independent by using registered words from a general thesaurus that can apply across multiple domains. The system extracts and utilizes viewpoint expressions that are not tied to specific fields, allowing the same reputation analysis mechanism to work effectively across different domains without requiring field-specific customization of the viewpoint dictionary.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS9507857B2Apparatus and method for classifying document, and computer program product
Publication Date: 2016.11.29 KK TOSHIBA
  • US9507857B2 patent drawing
  • US9507857B2 patent drawing
  • US9507857B2 patent drawing

AI summary

According to an embodiment, a document classification apparatus includes an extraction unit, a clustering unit, a classification unit, and a label assignment unit. The extraction unit is configured to extract feature words from documents. The clustering unit is configured to cluster the feature words into clusters so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value. The classification unit is configured to classify the documents into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs. The label assignment unit is configured to assign a classification label to each cluster as a word representative of the corresponding feature words.