Document Classification Using Multi-Dimensional Feature Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification methods are inefficient in accurately identifying desired documents within large sets, often requiring manual review to filter out unnecessary documents, and struggle to provide a general overview of document sets, leading to increased time and effort in searching and analysis.

Innovation Solution

A document classifying device that generates multi-dimensional feature vectors using classification codes and employs cluster analysis or latent topic analysis to classify documents, allowing for efficient identification of relevant documents and capturing the general picture of a document set by grouping similar content together.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing document classification methods are used, then documents can be classified using basic classification codes, but the accuracy of identifying desired documents is insufficient and manual review is required

Engineering Contradiction:
Improveclassification accuracyVSAvoidtime for manual review
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent transforms classification codes into multi-dimensional feature vectors, adding dimensional depth to the classification representation. This allows documents to be classified not just by single codes but by complex patterns across multiple dimensions, significantly improving classification accuracy and enabling automated identification of desired documents without manual review

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines multiple classification codes into composite feature vectors that capture nuanced document characteristics. By integrating information from multiple classification dimensions into a unified vector representation, the system achieves higher classification precision and can automatically distinguish desired documents from irrelevant ones

Inventive Principle:
Principle #40Composite materials

2Loss of information

If existing document classification methods are used, then basic classification can be performed, but the ability to provide a general overview of document sets is limited

Engineering Contradiction:
Improveoverview informationVSAvoidclassification system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts essential classification information from multiple codes and consolidates it into representative feature vectors. This extraction process captures the general overview of document sets by identifying dominant classification patterns and characteristics, providing comprehensive insights without requiring complex manual analysis

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a multi-functional classification system where the same feature vector representation serves multiple purposes: detailed document classification, general overview generation, and pattern recognition across document sets. This universal approach enables both specific and general analysis functions within a unified framework

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If manual review is used to filter documents, then accurate identification is possible, but the process requires increased time and effort

Engineering Contradiction:
Improvedocument identification reliabilityVSAvoidsearching efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent enables the classification system to automatically perform the filtering and identification functions that previously required manual review. By using multi-dimensional feature vectors and advanced classification algorithms, the system achieves reliable automated document identification, maintaining high accuracy while dramatically improving searching efficiency and reducing human effort

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10984344B2Document classifying device
Publication Date: 2021.04.20 KAO CORP
  • US10984344B2 patent drawing
  • US10984344B2 patent drawing
  • US10984344B2 patent drawing

AI summary

A document classifying device (10) includes: a unit (22) configured to acquire information regarding a to-be-classified document set in which classification codes based on multi-viewpoint classification are assigned to each document in advance; a unit (23) configured to generate a multi-dimensional feature vector for each document in the to-be-classified document set, the multi-dimensional feature vector having, as elements, all or part of the classification codes assigned to the to-be-classified document set; classifying unit (24) configured to classify the to-be-classified document set using the feature vector of each document; and a generating unit (25) that generates document classification information indicating a result of the classification.