Automatic Sampling Evaluation for Electronic Discovery Recall
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge in electronic discovery is the difficulty in efficiently reviewing and producing large volumes of electronic messages due to diverse formats and the need to identify responsive documents, which is time-consuming and costly, especially in compliance with regulatory requirements.
Innovation Solution
A semantic space is created using document and term vectors to represent electronically stored information, enabling automatic sampling evaluation and improving recall by analyzing similarity between documents and generating new search criteria based on noun phrases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review of electronic messages is performed to identify responsive documents, then precision in identifying relevant documents is improved, but productivity and cost efficiency deteriorate due to the time-consuming nature of the process
Solution Approach 1:
The patent introduces an automatic sampling evaluation system as an intermediary between manual review processes. This system uses document feature vectors, similarity calculations, and convergence detection to automatically evaluate whether additional searches will yield new responsive documents, thereby reducing the need for extensive manual review while maintaining identification accuracy
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated computational system. The system uses document feature vectors, similarity metrics (cosine similarity), and algorithmic convergence detection to substitute human reviewers in evaluating search effectiveness and identifying responsive documents
2Reliability
If comprehensive search of all documents is performed to maximize recall of responsive documents, then completeness of identification is improved, but loss of time and resources worsens
Solution Approach 1:
The patent implements feedback mechanisms where the automatic sampling evaluation system continuously monitors search results, calculates similarity between newly found documents and previously identified responsive documents, and provides feedback on whether to continue or terminate the search process based on detected convergence
Solution Approach 2:
The patent applies partial action by performing sampling evaluation on a subset of documents rather than comprehensive review of all documents. The system determines whether sampling provides sufficient evidence of convergence to justify stopping the search, thereby avoiding excessive action while maintaining reliability
3Adaptability or versatility
If diverse electronic message formats are reviewed manually, then adaptability to different formats is improved, but device complexity and operational difficulty worsen
Solution Approach 1:
The patent creates a universal document representation system using feature vectors that can represent diverse electronic message formats (emails, attachments, different file types) in a unified mathematical space. This universal representation enables the same similarity calculation and convergence detection algorithms to handle all formats without requiring format-specific processing logic
Data Source
AI summary
Techniques are provided for automatic sampling evaluation. An automatic sampling evaluation system enables users to evaluate convergence of one or more search processes. For example, given a set of searches that were validated by human review, a system can implement a retrieval process that samples one or more non-retrieved collections. Each individual document's similarity in the one or more non-retrieved collections is automatically evaluated to other documents in any retrieved sets. Given a goal of achieving a high recall, documents with high similarity can then be analyzed for additional noun phrases that may be used for a next iteration of a search. Convergence can be expected if the information gain in the new feedback loop is less than previous iterations, and if the additional documents identified are below a certain threshold document count.


