Active Learning Document Classification via Predictive Forking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current document classification systems in e-discovery are inefficient and require significant human intervention, computational resources, and infrastructure, often failing to achieve acceptable levels of precision and recall, especially when dealing with large document collections.
Innovation Solution
An efficient active learning platform that uses a single-phase approach to classify documents with minimal human input, leveraging predicted classifiers and forking to process documents interactively, reducing the need for seed sets and infrastructure, and implementing a user-friendly, portable solution for document classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of each document is performed, then precision and recall are improved, but time consumption and cost increase significantly
Solution Approach 1:
The system enables documents to classify themselves through automated machine learning algorithms. The computerized system performs document review autonomously by applying trained classifiers to categorize documents, eliminating the need for manual human review of each document while maintaining high precision and recall rates.
Solution Approach 2:
The patent replaces the mechanical human review process with an automated computerized classification system. Machine learning algorithms and predictive coding techniques substitute for human reviewers, using computational methods to analyze and categorize documents at scale without the time and resource constraints of manual review.
2Loss of time
If keyword-based culling is used to filter documents, then time consumption is reduced, but precision and recall deteriorate
Solution Approach 1:
The system transitions from simple keyword matching to sophisticated machine learning parameters. Instead of relying on basic text search keywords, the patent employs trained classifiers that analyze multiple document attributes and features simultaneously, dynamically adjusting classification thresholds to optimize both speed and accuracy.
Solution Approach 2:
The system implements feedback loops where classification results are continuously refined. Human reviewers provide feedback on initial automated classifications, and this feedback is used to retrain and improve the machine learning models, progressively enhancing precision while maintaining efficient processing speeds.
3Productivity
If technology-assisted review tools are deployed, then productivity increases, but device complexity and infrastructure requirements increase
Solution Approach 1:
The patent creates a multi-functional platform that combines document classification, predictive coding, active learning, and review management in a single integrated system. This universal tool performs multiple e-discovery functions simultaneously, reducing the need for separate specialized systems and complex infrastructure.
Solution Approach 2:
The system introduces an intelligent intermediary layer between raw documents and final classification. The machine learning models act as mediators that translate unstructured document data into structured classification outcomes, simplifying the overall system architecture while maintaining high productivity through automated intelligent processing.
4Measurement precision
If active learning with seed sets is implemented, then classification accuracy improves, but the need for manual input and computational resources increases
Solution Approach 1:
The system applies partial active learning by processing only the most uncertain or borderline documents that require human review. Instead of manually reviewing all documents or using extensive seed sets, the patent identifies and focuses computational resources on the subset of documents where machine classification confidence is lowest, achieving high accuracy with reduced resource consumption.
Data Source
AI summary
Systems and methods for classifying electronic information or documents into a number of classes and subclasses are provided through an active learning algorithm. In certain embodiments, seed sets may be eliminated by merging relevance feedback and machine learning phases. In certain embodiments, the active learning algorithm forks a number of classification paths corresponding to predicted user coding decisions for a selected document. The active learning algorithm determines an order in which the documents of the collection may be processed and scored by the forked classification paths. Such document classification systems are easily scalable for large document collections, require less manpower and can be employed on a single computer, thus requiring fewer resources. Furthermore, the classification systems and methods described can be used for any pattern recognition or classification effort in a wide variety of fields.


