Document Classification System Iterative Score Recalculation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document classification systems for legal cases, such as those disclosed in PTL 1 to PTL 3, face challenges in enhancing precision and recall of document classification results, leading to potential inclusion of irrelevant or confidential information in evidentiary materials.
Innovation Solution
A document classification system that extracts and classifies documents using a score calculation unit to evaluate linkage strength between documents and classification codes, with iterative recalculations based on user-defined criteria, including keyword selection and weighting, to improve precision and recall.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing document classification systems are used, then documents can be classified, but the precision and recall of classification results are insufficient
Solution Approach 1:
The patent implements feedback by iteratively recalculating scores based on classification results. The system calculates initial scores for documents, performs classification, then uses the classification results to recalculate scores with improved accuracy. This feedback loop continues until convergence, progressively enhancing both precision and reliability of document classification.
Solution Approach 2:
The patent applies preliminary action by pre-calculating scores for all documents before final classification. The system performs initial score calculations and preliminary classifications, then uses these results to guide subsequent iterative refinements. This preliminary processing establishes a foundation that improves the efficiency and accuracy of the overall classification process.
2Productivity
If existing document classification systems are used, then classification can be performed, but irrelevant or confidential information may be included in evidentiary materials
Solution Approach 1:
The system uses feedback mechanisms to iteratively improve classification accuracy while maintaining efficiency. By recalculating scores based on classification results and refining the classification process through multiple iterations, the system progressively reduces inclusion of irrelevant information while preserving productive classification of relevant documents.
Solution Approach 2:
The patent applies partial action by focusing computational resources on documents that require refinement. Rather than uniformly processing all documents with equal intensity, the system identifies documents with lower confidence classifications and applies additional iterative processing specifically to those cases, improving accuracy without excessive overall processing time.
Data Source
AI summary
The present invention includes: an extraction unit that extracts a specified quantity of documents, as targets to be classified by a user, from document information; a classification code accepting unit that accepts a classification code which is an identifier used when categorizing the documents, and is assigned by the user to each of the extracted documents; a database that records keywords selected from the extracted documents on the basis of the classification code; a score calculation unit that calculates a score which evaluates linkage strength between documents included in the document information, and the classification code on the basis of the keywords wherein the score calculation unit recalculates the score on the basis of a result of further extraction, by the extraction unit, of a specified quantity of documents, as targets to be classified by the user, from the document information according to the score.


