E-Discovery Effectiveness Estimation via Statistical Sampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for evaluating the effectiveness of information retrieval systems in electronic discovery are costly and time-consuming, requiring extensive human review to determine if the systems accurately classify documents.
Innovation Solution
A server computing system estimates the effectiveness of an information retrieval system by calculating statistics from a subset of test documents, including false negatives, true positives, and false positives, using a classification model and user classifications to determine recall and F-measure, thereby reducing the need for extensive human review.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human reviewers manually evaluate a large number of classified documents to determine classification accuracy, then measurement precision of system effectiveness is improved, but loss of time and loss of substance (human review cost) worsen
Solution Approach 1:
The patent uses a small set of test documents as a representative copy/sample of the entire corpus to estimate system effectiveness. Instead of evaluating all documents, a manageable subset is selected and evaluated, with statistics from this sample used to infer overall system performance, thereby reducing time and resource requirements while maintaining measurement precision
Solution Approach 2:
The patent applies partial action by evaluating only a portion (small set) of documents rather than the entire corpus. This partial evaluation is sufficient to generate reliable effectiveness estimates through statistical inference, avoiding the excessive time and cost of complete evaluation
2Measurement precision
If human reviewers manually evaluate a large number of classified documents to determine classification accuracy, then measurement precision of system effectiveness is improved, but loss of substance (human review cost) worsens
Solution Approach 1:
The patent uses a small set of test documents as a representative copy/sample of the entire corpus to estimate system effectiveness. Instead of evaluating all documents, a manageable subset is selected and evaluated, with statistics from this sample used to infer overall system performance, thereby reducing time and resource requirements while maintaining measurement precision
Solution Approach 2:
The patent applies partial action by evaluating only a portion (small set) of documents rather than the entire corpus. This partial evaluation is sufficient to generate reliable effectiveness estimates through statistical inference, avoiding the excessive time and cost of complete evaluation
3Productivity
If automated review using predictive coding is implemented, then productivity is improved, but device complexity worsens
Solution Approach 1:
The patent introduces an effectiveness estimation module as an intermediary between the predictive coding system and final deployment. This module trains classification models on training documents, evaluates them on test documents, and provides effectiveness metrics, thereby managing system complexity while enabling automated high-productivity review
Solution Approach 2:
The patent segments the document corpus into training documents and test documents, and segments the evaluation process into model training, effectiveness estimation, and deployment phases. This segmentation allows automated processing to handle large volumes efficiently while keeping the complexity of each segment manageable
Data Source
AI summary
A server computing system determines a plurality of statistics for a plurality of test documents, determines a number of false negatives for a corpus of documents based on one or more of the plurality of statistics for the plurality of test documents. The classification of a document of the corpus of documents is a false negative if classification of the document by a classification model is negative and classification of the document by a user is positive. The server computing system calculates an effectiveness of an information retrieval system on a corpus of documents based on the number of false negatives for the corpus of documents.


