Expert-Guided Text Classifier for Document Relevance Scoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computerized systems for electronic document analysis are highly dependent on human intuition and creativity, leading to uncontrolled variance in errors, risks, and costs, with manual keyword methods being inefficient and lacking graduated relevance scores, resulting in clumsy and rigid culling and review processes.
Innovation Solution
An expert-guided system that reviews sample documents, learns to score relevance, and iteratively improves accuracy, providing graduated relevance scores, automated prioritization, and keyword generation, enabling more efficient culling and review processes through machine learning and text classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If manual keyword methods are used for document culling, then the process is simple to implement, but the precision of relevance identification deteriorates
Solution Approach 1:
The patent replaces manual keyword-based mechanical classification with machine learning-based automated text classification. The system uses trained classifiers to automatically analyze document content and assign relevance scores, substituting human expert intuition with algorithmic processing that provides both automation and improved precision through learned patterns from training data.
Solution Approach 2:
The patent transforms the binary relevant/not-relevant classification into a graduated relevance scoring system. By changing the output parameter from discrete categories to continuous scores, the system provides more nuanced relevance identification while maintaining automated processing. This parameter change enables more precise differentiation between documents of varying relevance levels.
2Device complexity
If manual keyword methods are used for document culling, then the process is rigid and binary, but the adaptability to budget changes and case evolution deteriorates
Solution Approach 1:
The patent implements dynamic adaptability by allowing the system to adjust culling thresholds and re-run classification based on changing budget constraints or case priorities. The graduated scoring system enables flexible threshold adjustment without requiring process redesign, and the system can be retrained on new training data to adapt to evolving case requirements, providing both structural simplicity and operational flexibility.
3Extent of automation
If expert-based manual review is used, then creativity and intuition can be applied, but the variance in errors and costs increases
Solution Approach 1:
The patent implements self-service through automated machine learning classification that performs document review without requiring human expert intervention for each document. The system trains on initial data and then autonomously classifies subsequent documents, eliminating the need for continuous human expert involvement while maintaining consistent application of classification criteria, thereby reducing variance in errors and costs.
Solution Approach 2:
The patent incorporates feedback mechanisms where classification results can be reviewed and used to retrain and improve the classifier. This feedback loop allows the system to learn from corrections and improve its reliability over time while maintaining automated operation, reducing error variance through continuous improvement based on performance feedback.
4Measurement precision
If graduated relevance scores are implemented, then document prioritization precision improves, but the system complexity increases
Solution Approach 1:
The patent segments the document review process into distinct phases: training phase where classifiers are trained on labeled data, and deployment phase where trained classifiers automatically score documents. This segmentation allows the complex machine learning operations to be confined to the initial training phase, while the operational phase remains relatively simple, requiring only threshold application and score generation without continuous complex processing.
Data Source
AI summary
An electronic document analysis method receiving N electronic documents pertaining to a case encompassing a set of issues including at least one issue and establishing relevance of at least the N documents to at least one individual issue in the set of issues, the method comprising, for at least one individual issue from among the set of issues, receiving an output of a categorization process applied to each document in training and control subsets of the at least N documents, the output including, for each document in the subsets, one of a relevant-to-the-individual issue indication and a non-relevant-to-the-individual issue indication; building a text classifier simulating the categorization process using the output for all documents in the training subset of documents; and running the text classifier on the at least N documents thereby to obtain a ranking of the extent of relevance of each of the at least N documents to the individual issue. The method may also comprise evaluating the text classifier's quality using the output for all documents in the control subset.


