Document Classification Using Clustering and Category Updates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification techniques face challenges in accurately classifying documents due to ambiguous category boundaries, variability in classification methods among individuals, and difficulties in automatically assigning categories, leading to inconsistent and inaccurate results.

Innovation Solution

A document classification device comprising a characteristic extraction unit, a clustering unit, and a category update unit that extracts features from document data, clusters similar data, and updates categories based on appearance ratios to improve classification accuracy and independence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If categories are finely defined to reduce classification fluctuations, then classification precision improves, but category setting cost increases

Engineering Contradiction:
Improveclassification precisionVSAvoidcategory setting cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically learning category boundaries and characteristics from training data without requiring manual fine-tuning of category definitions. The machine learning model autonomously adapts to the data distribution, eliminating the need for expensive and time-consuming manual category setting while maintaining high classification precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system changes parameters dynamically by adjusting classification thresholds and category boundaries based on learned data characteristics rather than using fixed, pre-defined categories. This allows the system to achieve high precision without the overhead of manually setting and maintaining detailed category parameters.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If categories are determined automatically, then productivity improves, but measurement precision deteriorates

Engineering Contradiction:
Improveclassification speedVSAvoidclassification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system performs preliminary action by pre-training on labeled training data to learn optimal category boundaries and characteristics before actual classification. This pre-learning phase enables the model to achieve both high speed and high accuracy during automated classification, as the model has already internalized the nuanced patterns of different categories.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system incorporates feedback mechanisms where classification results are continuously evaluated and used to refine the model's understanding of category boundaries. This feedback loop allows the automated system to improve its precision over time while maintaining high productivity, as the model learns from its own performance and adjusts accordingly.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If manual classification is used, then measurement precision improves, but productivity deteriorates

Engineering Contradiction:
Improveclassification accuracyVSAvoidclassification speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system replaces the mechanical process of manual classification with an automated machine learning system. The ML model replicates and enhances human classification capabilities by learning from training data, achieving both the precision of expert human classifiers and the high productivity of automated processing, thereby eliminating the trade-off between manual accuracy and automated speed.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Ease of operation

If categories are ambiguously defined, then ease of operation improves, but measurement precision deteriorates

Engineering Contradiction:
Improvecategory definition simplicityVSAvoidclassification accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system performs self-service by automatically learning precise category boundaries and characteristics from training data without requiring manual definition of detailed category criteria. The model autonomously identifies the nuanced differences between categories, achieving high precision while maintaining ease of operation, as the system self-adapts to the data rather than requiring explicit parameter specification.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10353925B2Document classification device, document classification method, and computer readable medium
Publication Date: 2019.07.16 FUJIFILM BUSINESS INNOVATION CORP
  • US10353925B2 patent drawing
  • US10353925B2 patent drawing
  • US10353925B2 patent drawing

AI summary

A document classification device includes a characteristic extraction unit, a clustering unit, and a category update unit. The characteristic extraction unit extracts characteristic information from each of plural document data which are classified in advance into specific categories. The clustering unit classifies the document data with similar appearance frequency of the characteristic information into a same cluster. The category update unit assigns the document data which is classified into the same cluster with a category of different document data which is classified into the same cluster as a category of the document data.