Document Classification Using Thesaurus Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document classification methods face challenges in presenting document analysis results intelligibly due to the high number of document clusters and biased document distribution, particularly when using thesauruses for classification, which leads to a lack of versatility in viewpoint dictionaries and increased complexity in interpreting classification labels.
Innovation Solution
A document classification apparatus that extracts feature words from documents, clusters them into subtrees of a thesaurus to equalize document distribution, and assigns classification labels to clusters, allowing for intelligible presentation of results without relying on field-specific viewpoint dictionaries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If document clusters are classified using registered words on the same layer of thesaurus, then classification granularity is standardized, but the number of document clusters increases excessively
Solution Approach 1:
The patent merges multiple document clusters that share the same registered word on the thesaurus into a single integrated cluster. Instead of creating separate clusters for each registered word occurrence, the system combines them under one unified cluster, thereby reducing the total number of clusters while preserving the standardized granularity provided by the thesaurus structure.
2Measurement precision
If registered words from lower level concepts in thesaurus are used as classification labels, then classification precision is improved, but interpretability of results deteriorates
Solution Approach 1:
The patent implements a nested labeling structure where classification labels are organized in hierarchical levels. The display label comes from a higher-level concept in the thesaurus hierarchy, providing broad interpretability, while the actual classification is performed using more specific lower-level registered words, maintaining precision. This nested approach allows users to see both the general category and the specific classification criteria.
3Measurement precision
If field-specific viewpoint dictionaries are used for reputation analysis, then analysis accuracy is improved, but versatility of the system deteriorates
Solution Approach 1:
The patent makes the viewpoint dictionary field-independent by using registered words from a general thesaurus that can apply across multiple domains. The system extracts and utilizes viewpoint expressions that are not tied to specific fields, allowing the same reputation analysis mechanism to work effectively across different domains without requiring field-specific customization of the viewpoint dictionary.
Data Source
AI summary
According to an embodiment, a document classification apparatus includes an extraction unit, a clustering unit, a classification unit, and a label assignment unit. The extraction unit is configured to extract feature words from documents. The clustering unit is configured to cluster the feature words into clusters so that a difference between the number of documents each including any one of the feature words belonging to one cluster and the number of documents each including any one of the feature words belonging to another cluster is equal to or less than a predetermined reference value. The classification unit is configured to classify the documents into the clusters so that each document belongs to the cluster to which the feature word included in the each document belongs. The label assignment unit is configured to assign a classification label to each cluster as a word representative of the corresponding feature words.


