Automated Document Classification Using Keyword and Related Term Scoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current forensic systems require manual effort and cost to classify large amounts of digitized document information for legal actions, as they lack automation in determining the validity of documents as evidentiary materials.

Innovation Solution

A document classification system that utilizes keyword and related term databases to automatically attach classification marks to documents based on their association with legal actions, using keyword-corresponding and related term-corresponding information, and a scoring system to determine the degree of association, thereby reducing reviewer burden and improving classification efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual classification of document information is performed by reviewers, then classification precision can be maintained, but the effort and cost increase significantly

Engineering Contradiction:
Improveclassification precisionVSAvoidclassification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical classification process with an automated information processing system that uses databases, extraction units, and classification units to automatically attach classification marks to documents, eliminating the need for manual reviewer intervention while maintaining classification accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service classification by automatically processing document information through extraction units that identify keywords and classification units that assign appropriate classification marks without requiring external reviewer input, allowing the system to classify documents independently

Inventive Principle:
Principle #25Self-service

2Reliability

If manual classification of document information is performed by reviewers, then accurate classification can be achieved, but the cost and burden increase

Engineering Contradiction:
Improveclassification reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the classification system into distinct functional modules including extraction units for identifying keywords and classification units for assigning classification marks, allowing each component to perform its specific function reliably while keeping the overall system manageable through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces databases as intermediary components that store and manage classification information, serving as a bridge between document input and classification output, which simplifies the overall system architecture while ensuring reliable and consistent classification results

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automatic classification is implemented, then reviewer burden is reduced, but classification precision may deteriorate

Engineering Contradiction:
Improveclassification efficiencyVSAvoidclassification precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent replaces manual classification with an automated system that uses extraction units to identify keywords and classification units to assign classification marks based on predefined criteria, achieving both high productivity and maintained precision through systematic automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9495445B2Document sorting system, document sorting method, and document sorting program
Publication Date: 2016.11.15 FRONTEO INC
  • US9495445B2 patent drawing
  • US9495445B2 patent drawing
  • US9495445B2 patent drawing

AI summary

It is possible to analyze digitized document information gathered to be provided as evidence in a legal action and to classify the document information to be easily accessible in the legal action. A document classification system includes a keyword database, a related term database, a first classification unit which extracts a document including a keyword recorded in the keyword database from document information and attaches a specific classification mark to the extracted document based on keyword-corresponding information, and a second classification unit which extracts a document including a related term recorded in the related term database from document information, to which the specific classification mark is not attached in the first classification unit, calculates a score based on an evaluated value of the related term included in the extracted document and the number of related terms, and attaches a predetermined classification mark to a document, for which the score exceeds a given value, among the documents including the related term based on the score and the related term-corresponding information.