Dynamic Random Document Sampling for Representative eDiscovery Review
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for maintaining a random sample of documents struggle to keep the sample representative as the document pool size and content change during eDiscovery processes, often requiring additional manual review or less accurate measurements of classifier precision and recall.
Innovation Solution
A method and system for maintaining a random sample by ingesting documents into a pool, sorting them randomly, providing them to a review platform, defining the initial sample based on applied labels, and interleaving new documents to maintain relative order, thereby adjusting the sample size and variety without manual review.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a random sample is taken at the onset of the eDiscovery process, then the initial assessment of classifier precision and recall can be performed, but the sample becomes unrepresentative as the document pool changes over time
Solution Approach 1:
The system dynamically updates the random sample as documents are reviewed and added to the pool. Instead of using a static initial sample, the methodology continuously regenerates the random sample to reflect the current state of the document pool, ensuring the sample remains representative throughout the eDiscovery process
Solution Approach 2:
The system uses feedback from the review process to update the random sample. As documents are reviewed and labeled, this information feeds back into the system to adjust and regenerate the random sample, ensuring it accurately reflects the evolving document pool composition
2Measurement precision
If a new random sample is generated each time the corpus changes, then the precision and recall measurements remain accurate, but additional manual review of documents is required
Solution Approach 1:
The system performs self-service by automatically regenerating the random sample using computational methods. Instead of requiring manual selection and review of documents to create a new sample, the system uses algorithms to automatically generate representative samples from the evolving document pool, reducing manual intervention
Solution Approach 2:
The methodology replaces manual mechanical processes with automated computational systems. Instead of manually reviewing and selecting documents for the random sample, computer algorithms automatically generate the sample based on the current document pool, substituting human effort with automated processing
3Quantity of substance
If the document pool size changes during the eDiscovery process, then more documents can be reviewed, but the original random sample no longer represents the updated pool
Solution Approach 1:
The random sample composition dynamically adapts to changes in the document pool. As documents are added or removed from the pool during the eDiscovery process, the system regenerates the random sample to maintain the correct proportional representation of different document types and categories
Solution Approach 2:
The system changes the parameters of the random sample based on the current state of the document pool. When the pool composition changes, the sampling parameters are adjusted to reflect the new distribution of document types, ensuring the sample accurately represents the updated pool
Data Source
AI summary
Systems and methods related to maintaining a random sample of documents that is representative of a pool of documents are provided. As documents are ingested into the pool of documents, a random number may be assigned to the documents. The documents may then be sorted into an ordered list. As the documents in the pool are provided to a review platform for manual review, the documents may be included in a review queue based at least in part on the ordered list. As additional documents are added to the pool of documents, the new documents are interleaved into the ordered list to maintain the random and representative nature of the random sample.


