Batch-Mode Active Learning for Technology-Assisted Review
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technology-assisted review (TAR) methods face challenges in efficiently balancing the training of systems for document classification, particularly in large datasets, where they require significant expert time and struggle with model stabilization and recall optimization, especially in highly imbalanced document distributions.
Innovation Solution
The implementation of batch-mode active learning using Support Vector Machines (SVM) with techniques like Diversity Sampler and Biased Probabilistic Sampler to select unlabeled instances for training, along with a Kappa agreement-based stopping criterion, to optimize the training process and improve recall efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional linear review methods are used for large document datasets, then comprehensive document review is achieved, but time consumption and cost increase significantly
Solution Approach 1:
The system enables self-service through automated machine learning models that perform document classification independently. The TAR system trains classifiers on labeled documents and automatically applies them to unlabeled documents, reducing reliance on continuous expert intervention and enabling the system to serve itself in the document review process.
Solution Approach 2:
The patent replaces manual mechanical review processes with automated computational systems. Machine learning classifiers substitute for human reviewers in sorting and classifying documents, while active learning algorithms automate the selection of documents for labeling, transforming a labor-intensive mechanical process into an automated computational workflow.
2Measurement precision
If extensive expert labeling is performed to improve model accuracy, then classification performance increases, but training time and expert effort increase
Solution Approach 1:
The active learning system applies partial action by selecting only the most informative subset of documents for labeling rather than labeling all documents. The uncertainty sampling mechanism identifies documents that provide maximum information gain, allowing the system to achieve high classification accuracy with a fraction of the labeling effort required by traditional methods.
Solution Approach 2:
The system implements feedback loops where classification results are continuously evaluated and used to guide subsequent labeling decisions. The active learning process uses prediction uncertainty as feedback to determine which documents should be labeled next, creating an iterative improvement cycle that increases accuracy while minimizing labeling effort.
3Productivity
If batch-mode active learning with uncertainty sampling is used, then labeling efficiency improves, but model stabilization becomes difficult in imbalanced distributions
Solution Approach 1:
The system applies local quality by using different sampling strategies for different regions of the feature space. For imbalanced distributions, the patent employs stratified sampling or region-specific uncertainty thresholds that adapt to local data characteristics, ensuring stable model training by appropriately representing minority classes while maintaining overall labeling efficiency.
Solution Approach 2:
The system stabilizes model training in imbalanced distributions by dynamically adjusting parameters such as uncertainty thresholds, batch sizes, and sampling probabilities. The active learning process modifies these parameters based on observed class distributions and model performance, enabling stable convergence even when data is highly imbalanced.
4Measurement precision
If more documents are reviewed in the second pass to improve recall, then more relevant documents are found, but attorney time and effort increase
Solution Approach 1:
The system performs preliminary action by pre-classifying and ranking documents before attorney review. The trained classifier processes the entire document collection and ranks documents by predicted relevance, allowing attorneys to review only the top-ranked documents that are most likely to be relevant, thereby achieving high recall with minimal attorney time investment.
Solution Approach 2:
The system creates a copied and simplified version of the document review process through automated classification. Instead of attorneys reviewing all documents directly, the system creates a filtered subset of priority documents based on machine learning predictions, effectively copying the essential review task onto a manageable subset while maintaining high recall.
Data Source
AI summary
The present disclosure relates to the electronic document review field and, more particularly, to various apparatuses and methods of implementing batch-mode active learning for technology-assisted review (TAR) of documents (e.g., legal documents).


