Automated Document Classification Using Keyword and Related Term Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current forensic systems require manual effort and cost to classify large amounts of digitized document information for legal actions, as they lack automation in determining the validity of documents as evidentiary materials.
Innovation Solution
A document classification system that utilizes keyword and related term databases to automatically attach classification marks to documents based on their association with legal actions, using keyword-corresponding and related term-corresponding information, and a scoring system to determine the degree of association, thereby reducing reviewer burden and improving classification efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual classification of document information is performed by reviewers, then classification precision can be maintained, but the effort and cost increase significantly
Solution Approach 1:
The patent replaces the manual mechanical classification process with an automated information processing system that uses databases, extraction units, and classification units to automatically attach classification marks to documents, eliminating the need for manual reviewer intervention while maintaining classification accuracy
Solution Approach 2:
The system enables self-service classification by automatically processing document information through extraction units that identify keywords and classification units that assign appropriate classification marks without requiring external reviewer input, allowing the system to classify documents independently
2Reliability
If manual classification of document information is performed by reviewers, then accurate classification can be achieved, but the cost and burden increase
Solution Approach 1:
The patent divides the classification system into distinct functional modules including extraction units for identifying keywords and classification units for assigning classification marks, allowing each component to perform its specific function reliably while keeping the overall system manageable through modular design
Solution Approach 2:
The patent introduces databases as intermediary components that store and manage classification information, serving as a bridge between document input and classification output, which simplifies the overall system architecture while ensuring reliable and consistent classification results
3Productivity
If automatic classification is implemented, then reviewer burden is reduced, but classification precision may deteriorate
Solution Approach 1:
The patent replaces manual classification with an automated system that uses extraction units to identify keywords and classification units to assign classification marks based on predefined criteria, achieving both high productivity and maintained precision through systematic automated processing
Data Source
AI summary
It is possible to analyze digitized document information gathered to be provided as evidence in a legal action and to classify the document information to be easily accessible in the legal action. A document classification system includes a keyword database, a related term database, a first classification unit which extracts a document including a keyword recorded in the keyword database from document information and attaches a specific classification mark to the extracted document based on keyword-corresponding information, and a second classification unit which extracts a document including a related term recorded in the related term database from document information, to which the specific classification mark is not attached in the first classification unit, calculates a score based on an evaluated value of the related term included in the extracted document and the number of related terms, and attaches a predetermined classification mark to a document, for which the score exceeds a given value, among the documents including the related term based on the score and the related term-corresponding information.


