E-Discovery Effectiveness Estimation via Statistical Sampling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for evaluating the effectiveness of information retrieval systems in electronic discovery are costly and time-consuming, requiring extensive human review to determine if the systems accurately classify documents.

Innovation Solution

A server computing system estimates the effectiveness of an information retrieval system by calculating statistics from a subset of test documents, including false negatives, true positives, and false positives, using a classification model and user classifications to determine recall and F-measure, thereby reducing the need for extensive human review.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human reviewers manually evaluate a large number of classified documents to determine classification accuracy, then measurement precision of system effectiveness is improved, but loss of time and loss of substance (human review cost) worsen

Engineering Contradiction:
Improveeffectiveness evaluation accuracyVSAvoidevaluation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses a small set of test documents as a representative copy/sample of the entire corpus to estimate system effectiveness. Instead of evaluating all documents, a manageable subset is selected and evaluated, with statistics from this sample used to infer overall system performance, thereby reducing time and resource requirements while maintaining measurement precision

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies partial action by evaluating only a portion (small set) of documents rather than the entire corpus. This partial evaluation is sufficient to generate reliable effectiveness estimates through statistical inference, avoiding the excessive time and cost of complete evaluation

Inventive Principle:
Principle #16Partial or excessive action

2Measurement precision

If human reviewers manually evaluate a large number of classified documents to determine classification accuracy, then measurement precision of system effectiveness is improved, but loss of substance (human review cost) worsens

Engineering Contradiction:
Improveeffectiveness evaluation accuracyVSAvoidhuman review cost
Core Design Contradiction:
Measurement precisionVSLoss of substance

Solution Approach 1:

The patent uses a small set of test documents as a representative copy/sample of the entire corpus to estimate system effectiveness. Instead of evaluating all documents, a manageable subset is selected and evaluated, with statistics from this sample used to infer overall system performance, thereby reducing time and resource requirements while maintaining measurement precision

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies partial action by evaluating only a portion (small set) of documents rather than the entire corpus. This partial evaluation is sufficient to generate reliable effectiveness estimates through statistical inference, avoiding the excessive time and cost of complete evaluation

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If automated review using predictive coding is implemented, then productivity is improved, but device complexity worsens

Engineering Contradiction:
Improvedocument review efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an effectiveness estimation module as an intermediary between the predictive coding system and final deployment. This module trains classification models on training documents, evaluates them on test documents, and provides effectiveness metrics, thereby managing system complexity while enabling automated high-productivity review

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the document corpus into training documents and test documents, and segments the evaluation process into model training, effectiveness estimation, and deployment phases. This segmentation allows automated processing to handle large volumes efficiently while keeping the complexity of each segment manageable

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9613319B1Method and system for information retrieval effectiveness estimation in e-discovery
Publication Date: 2017.04.04 ARCTERA US LLC
  • US9613319B1 patent drawing
  • US9613319B1 patent drawing
  • US9613319B1 patent drawing

AI summary

A server computing system determines a plurality of statistics for a plurality of test documents, determines a number of false negatives for a corpus of documents based on one or more of the plurality of statistics for the plurality of test documents. The classification of a document of the corpus of documents is a false negative if classification of the document by a classification model is negative and classification of the document by a user is positive. The server computing system calculates an effectiveness of an information retrieval system on a corpus of documents based on the number of false negatives for the corpus of documents.