Asynchronous Interactive ML for Legal Document Discovery

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The process of electronic discovery in legal proceedings is complex and costly due to the high volume and diversity of electronic documents, leading to time-consuming and expensive disputes over the preservation, review, and admissibility of evidence.

Innovation Solution

An Asynchronous and Interactive Machine Learning (AIML) system that uses machine learning techniques to predict the likelihood of document tagging, allowing for iterative training and retraining based on user interactions, enabling efficient sorting and annotation of documents within a data corpus.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional manual review methods are used for electronic document discovery, then quality assurance and thoroughness can be maintained, but the process becomes excessively time-consuming and costly

Engineering Contradiction:
Improvequality assuranceVSAvoiddiscovery time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a machine learning model as an intermediary between the document corpus and human reviewers. The model processes documents asynchronously, generating predictions about relevance and sensitivity that guide human review efforts, thereby maintaining quality assurance while reducing overall discovery time and cost

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing of documents using machine learning models before human review. By pre-analyzing documents and generating predictions in advance, the system prepares data structures and identifies candidate documents, allowing human reviewers to focus only on uncertain cases and thereby reducing total review time

Inventive Principle:
Principle #10Preliminary action

2Reliability

If comprehensive manual review of all electronic documents is conducted, then high quality standards can be maintained, but costs and resource requirements increase significantly

Engineering Contradiction:
Improvequality standardVSAvoiddiscovery efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent implements feedback loops where human reviewer decisions are used to retrain and improve the machine learning model. The system continuously learns from corrected predictions and user interactions, allowing quality standards to be maintained while progressively improving efficiency as the model becomes more accurate

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The machine learning model serves itself by automatically processing documents, generating predictions, and identifying cases requiring human review. The system autonomously manages the discovery process, only engaging human reviewers when necessary, thereby maintaining quality standards while significantly improving overall productivity

Inventive Principle:
Principle #25Self-service

3Measurement precision

If synchronous machine learning training is used during document review, then model accuracy can be improved, but the review process is interrupted and becomes less efficient

Engineering Contradiction:
Improvemodel accuracyVSAvoidreview throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements periodic retraining of the machine learning model at predetermined intervals rather than continuously during review. This allows the model to be improved periodically while maintaining steady review throughput between training cycles, balancing accuracy improvements with productivity

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system dynamically adjusts between synchronous and asynchronous operations based on model confidence levels. When the model is highly confident, documents are processed asynchronously without interruption. When uncertainty is detected, the system synchronously engages human reviewers and uses their input for immediate model updates, optimizing both accuracy and throughput

Inventive Principle:
Principle #15Dynamics

4Quantity of substance

If large volumes of diverse electronic documents are processed manually, then complete coverage can be achieved, but the complexity and cost of discovery increases

Engineering Contradiction:
Improvedocument coverageVSAvoiddiscovery process complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the document corpus into manageable batches that can be processed asynchronously by the machine learning model. Documents are divided into smaller units for independent processing, allowing complete coverage of large volumes while simplifying the overall discovery process through modular, parallel execution

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11270225B1Methods and apparatus for asynchronous and interactive machine learning using word embedding within text-based documents and multimodal documents
Publication Date: 2022.03.08 CS DISCO INC
  • US11270225B1 patent drawing
  • US11270225B1 patent drawing
  • US11270225B1 patent drawing

AI summary

A machine learning system continuously receives tag signals indicating membership relations between data objects from a data corpus and tag targets. The machine learning system is asynchronously and iteratively trained with the received tag signals to identify further data objects from the data corpus predicted to have a membership relation with the single tag target. The machine learning system constantly improves its predictive accuracy in short time by the continuous training of a backend machine learning model based on implicit and explicit tag signals gathered from a non-intrusive monitoring of user interactions during a review process of the data corpus.