Recall Estimation via Random Permutation and Confidence Sequences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current evaluation methods for one-phase Technology-Assisted Review (TAR) workflows in electronic discovery lack practicality and flexibility, failing to account for operational constraints and biases in human review processes, leading to inaccurate recall estimates and limited manager control over the review process.
Innovation Solution
A novel evaluation method using confidence sequences, which assigns unique permanent random numbers to documents, generates random permutations, and employs the Wauby-Smith and Ramdas algorithm to provide valid frequentist confidence intervals on recall, allowing for flexible review management and reduced bias through sequential sampling and updated estimates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional evaluation methods are used for one-phase TAR workflows, then the review process is simpler to operate, but the recall estimates become inaccurate and lack statistical validity
Solution Approach 1:
The patent applies preliminary action by pre-generating a random permutation of all documents in the database before the review process begins. This permanent random ordering is established in advance and remains fixed throughout the review, allowing statistical validity to be maintained while providing complete control over the review process through the permutation structure.
Solution Approach 2:
The patent implements feedback mechanisms through confidence sequences that provide updated recall estimates at every point during the review process. The system continuously monitors coding decisions and updates confidence intervals, allowing managers to see the impact of their review actions on recall in real-time while maintaining statistical validity.
2Measurement precision
If larger sample sizes are used for evaluation, then the statistical validity and precision of recall estimates improve, but the time and resources required for review increase
Solution Approach 1:
The patent applies dynamics by making the evaluation process adaptive and flexible. The random permutation structure allows the review to proceed dynamically with complete control over the process, and the confidence sequences provide updated estimates at any point, enabling efficient allocation of review time and resources while maintaining statistical validity.
Solution Approach 2:
The patent utilizes parameter changes by allowing the evaluation to adapt to changing conditions during the review process. The confidence sequences update their parameters based on incoming coding decisions, and the system can adjust the review strategy based on observed recall patterns, optimizing the balance between sample size requirements and review time.
3Adaptability or versatility
If human reviewers are used in one-phase TAR, then operational flexibility and adaptability are improved, but bias and operational constraints affect review accuracy
Solution Approach 1:
The patent introduces an intermediary structure in the form of a permanent random permutation that mediates between human reviewer flexibility and the need for consistent, unbiased evaluation. The random permutation serves as an objective framework that guides the review process, ensuring that human adaptability is exercised within a statistically valid structure that reduces bias.
Data Source
AI summary
A method for estimating recall in database workflows involves assigning unique keys to documents in a database, generating a random permutation of the database, accessing coding decisions for documents, determining the longest prefix of the permutation with corresponding coding decisions, calculating the number of relevant documents, applying a confidence sequence generating algorithm to the prefix to establish a confidence interval on the relevant documents, and computing a confidence interval estimate on recall for the database. This method provides a systematic approach to assess recall in database workflows, enabling accurate evaluation of the retrieval performance of relevant documents within the database.


