Scalable Continuous Active Learning for Document Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technology-assisted review (TAR) methods are inefficient in minimizing human review effort and achieving high recall for large datasets, particularly in electronic discovery, where they often require burdensome protocols with little assurance of success.
Innovation Solution
The implementation of a scalable continuous active learning (S-CAL) approach that iteratively selects and labels documents using a sub-sample, updates classifiers, and estimates prevalence to reduce human review effort and provide calibrated recall and precision, allowing for accurate classification of large document collections.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional TAR methods are used to classify large document collections, then recall can be achieved, but human review effort becomes excessively burdensome
Solution Approach 1:
The patent segments the document collection into batches that are further divided into subsamples for reviewer assessment. This segmentation allows the system to process large document collections efficiently by evaluating only small portions (e.g., 10-50 documents per batch) while using machine learning classifiers to handle the remaining documents, thereby reducing human review effort while maintaining high recall
Solution Approach 2:
The patent implements continuous feedback loops where reviewer assessments of subsamples are used to update and retrain classifiers, which then refine their predictions for the entire batch. This feedback mechanism allows the system to learn from limited human input and improve its classification accuracy iteratively, achieving high recall without requiring extensive human review
2Reliability
If more documents are reviewed by humans to improve recall, then classification accuracy increases, but the process becomes less scalable
Solution Approach 1:
The patent enables the classification system to serve itself by using automatically trained classifiers to pre-screen and rank documents within batches. The system only requires minimal human input for subsample assessment, and the classifiers independently handle the majority of classification tasks, making the process scalable to large document collections without proportionally increasing human review effort
Solution Approach 2:
The patent dynamically adjusts parameters such as batch size and subsample size based on the evolving performance of classifiers and the characteristics of the document collection. This allows the system to optimize the balance between human review effort and classification accuracy for different scales of document collections, maintaining scalability while adapting to specific task requirements
3Measurement precision
If traditional active learning is used, then some classification accuracy is achieved, but it cannot provide calibrated estimates of recall and precision
Solution Approach 1:
The patent uses reviewer assessments of subsamples as feedback to not only retrain classifiers but also to calculate calibrated estimates of recall and precision. By comparing classifier predictions against actual reviewer judgments in the subsamples, the system can estimate the performance characteristics of the entire batch, providing valuable information about the completeness and accuracy of the classification results
Solution Approach 2:
The patent replaces traditional mechanical sampling methods with a systematic approach that uses classifier predictions to identify and prioritize documents for reviewer assessment. This substitution allows for more efficient estimation of recall and precision by focusing human review on documents where the classifier is least certain, thereby obtaining better calibrated estimates with fewer reviewed documents
Data Source
AI summary
Systems and methods for classifying electronic information are provided by way of a Technology-Assisted Review (“TAR”) process. In certain embodiments, the TAR process is a Scalable Continuous Active Learning (“S-CAL”) approach. In certain embodiments, S-CAL selects an initial sample from a document collection, trains a classifier by using a default classification for a portion of the initial sample, scores the initial sample, selects a sub-sample from the initial sample for review, removes the reviewed sub-sample from the initial sample, and repeats the process by re-training the classifier until the initial sample is exhausted. In certain embodiments, a classification threshold is determined using a calculated estimate of the prevalence of relevant information such that the threshold classifies the information in accordance with a determined target criteria. In certain embodiments, the estimate of prevalence is determined from the results of iterations of a TAR process such as S-CAL.


