Asynchronous Interactive ML for Legal Document Discovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of electronic discovery in legal proceedings is complex and costly due to the high volume and diversity of electronic documents, leading to time-consuming and expensive disputes over the preservation, review, and admissibility of evidence.
Innovation Solution
An Asynchronous and Interactive Machine Learning (AIML) system that uses machine learning techniques to predict the likelihood of document tagging, allowing for iterative training and retraining based on user interactions, enabling efficient sorting and annotation of documents within a data corpus.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional manual review methods are used for electronic document discovery, then quality assurance and thoroughness can be maintained, but the process becomes excessively time-consuming and costly
Solution Approach 1:
The patent introduces a machine learning model as an intermediary between the document corpus and human reviewers. The model processes documents asynchronously, generating predictions about relevance and sensitivity that guide human review efforts, thereby maintaining quality assurance while reducing overall discovery time and cost
Solution Approach 2:
The system performs preliminary processing of documents using machine learning models before human review. By pre-analyzing documents and generating predictions in advance, the system prepares data structures and identifies candidate documents, allowing human reviewers to focus only on uncertain cases and thereby reducing total review time
2Reliability
If comprehensive manual review of all electronic documents is conducted, then high quality standards can be maintained, but costs and resource requirements increase significantly
Solution Approach 1:
The patent implements feedback loops where human reviewer decisions are used to retrain and improve the machine learning model. The system continuously learns from corrected predictions and user interactions, allowing quality standards to be maintained while progressively improving efficiency as the model becomes more accurate
Solution Approach 2:
The machine learning model serves itself by automatically processing documents, generating predictions, and identifying cases requiring human review. The system autonomously manages the discovery process, only engaging human reviewers when necessary, thereby maintaining quality standards while significantly improving overall productivity
3Measurement precision
If synchronous machine learning training is used during document review, then model accuracy can be improved, but the review process is interrupted and becomes less efficient
Solution Approach 1:
The patent implements periodic retraining of the machine learning model at predetermined intervals rather than continuously during review. This allows the model to be improved periodically while maintaining steady review throughput between training cycles, balancing accuracy improvements with productivity
Solution Approach 2:
The system dynamically adjusts between synchronous and asynchronous operations based on model confidence levels. When the model is highly confident, documents are processed asynchronously without interruption. When uncertainty is detected, the system synchronously engages human reviewers and uses their input for immediate model updates, optimizing both accuracy and throughput
4Quantity of substance
If large volumes of diverse electronic documents are processed manually, then complete coverage can be achieved, but the complexity and cost of discovery increases
Solution Approach 1:
The patent segments the document corpus into manageable batches that can be processed asynchronously by the machine learning model. Documents are divided into smaller units for independent processing, allowing complete coverage of large volumes while simplifying the overall discovery process through modular, parallel execution
Data Source
AI summary
A machine learning system continuously receives tag signals indicating membership relations between data objects from a data corpus and tag targets. The machine learning system is asynchronously and iteratively trained with the received tag signals to identify further data objects from the data corpus predicted to have a membership relation with the single tag target. The machine learning system constantly improves its predictive accuracy in short time by the continuous training of a backend machine learning model based on implicit and explicit tag signals gathered from a non-intrusive monitoring of user interactions during a review process of the data corpus.


