Document Analytics System for E-Discovery Tagging Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The process of document searching and analysis in legal discovery is time-consuming due to the large volume of documents that need to be searched, reviewed, and identified, particularly in e-discovery where thousands of documents require manual tagging and analysis.
Innovation Solution
A method and system that includes generating a graphical user interface with an analytics panel for document analytics, allowing users to receive manual tags, perform iterations to identify improperly associated documents based on content and metadata, and update reports with reclassified tags, while also categorizing documents by sender domain, illustrating word prevalence, and identifying irrelevant documents by file type.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual tagging and review of thousands of documents is performed, then document classification accuracy can be improved, but the time required for discovery increases significantly
Solution Approach 1:
The system performs preliminary automated classification of documents using machine learning models before manual review. This preliminary action pre-sorts documents into categories, so that during manual review, attorneys only need to verify or adjust classifications rather than classify from scratch, significantly reducing the time required while maintaining accuracy
Solution Approach 2:
The system implements an iterative feedback loop where attorney corrections to automated classifications are fed back into the machine learning model to improve future classifications. This allows the system to learn from human expertise and continuously improve accuracy, reducing the need for extensive manual review over time
2Reliability
If thousands of documents are manually reviewed and tagged, then comprehensive document identification is achieved, but productivity decreases due to the large volume of work
Solution Approach 1:
The system segments the large volume of documents into manageable batches or clusters based on similarity metrics. Instead of reviewing thousands of documents individually, attorneys review smaller, organized groups, which maintains comprehensive identification while significantly improving productivity through reduced cognitive load and faster processing of segmented data
Solution Approach 2:
The system replaces the mechanical process of manual document-by-document review with an automated machine learning-based classification system. This substitution handles the bulk of classification work automatically, allowing attorneys to focus only on edge cases or verify results, thereby dramatically increasing productivity while maintaining reliability
3Measurement precision
If detailed manual analysis of document content is performed, then tagging accuracy improves, but the complexity of the review process increases
Solution Approach 1:
The machine learning system performs self-service classification by automatically analyzing document content, metadata, and patterns to assign classifications. This self-service capability handles the complex analysis work that would otherwise require detailed manual review, maintaining tagging accuracy while reducing the perceived complexity for users who only need to verify rather than perform the detailed analysis
Data Source
AI summary
Implementations generally relate to providing document analytics. In some implementations, a method includes receiving a plurality of documents related to e-discovery. The method further includes generating a graphical user interface that includes an analytics panel that provides analytics information about the plurality of documents. The method further includes receiving, from one or more users, manual tags for one or more documents of the plurality of documents. The method further includes performing a first iteration that determines a first group of documents that are improperly associated with one or more of the manual tags based on at least one of content and metadata of the plurality of documents. The method further includes performing a second iteration that determines a second group of documents that are improperly associated with one or more of the manual tags based on the reclassification.


