Generative AI Validation Sample Updates for Consistent Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing eDiscovery processes face inefficiencies and inconsistencies due to the need for manual review and varying interpretations of document responsiveness using machine learning models, leading to conflicting classifications and the requirement of extensive labeled training examples.
Innovation Solution
A generative AI model is used to classify documents based on prompt criteria, allowing for efficient classification with reduced manual review, and includes a prompt-based model that can handle multiple classifications with a single call, assisted by a generative AI model to reconcile different user inputs and update prompts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning models are deployed for document classification, then classification speed is improved, but classification consistency deteriorates due to conflicting interpretations by different attorneys
Solution Approach 1:
The system implements feedback mechanisms where classification results are reviewed and validated, allowing the model to learn from discrepancies and improve consistency. User feedback on misclassifications is incorporated to refine the model's performance across different document types and legal contexts.
Solution Approach 2:
The system dynamically adjusts classification parameters and thresholds based on the specific legal context and document characteristics. By adapting parameters to match different attorneys' interpretive frameworks, the system maintains both speed and consistency across varying classification scenarios.
2Measurement precision
If manual review of documents is performed to generate labeled training examples, then classification accuracy is improved, but time consumption increases significantly
Solution Approach 1:
The system performs preliminary actions by automatically generating initial classifications and training data annotations before human review. This pre-processing reduces the amount of manual work needed, as reviewers only need to validate or correct the automated preliminary results rather than starting from scratch.
Solution Approach 2:
The system enables self-service by allowing it to automatically generate, validate, and refine its own training data and classification models with minimal human intervention. The model can self-correct errors and improve its performance autonomously through iterative feedback loops.
3Adaptability or versatility
If multiple machine learning models are trained for different classification tasks, then classification comprehensiveness is improved, but system complexity increases
Solution Approach 1:
The system implements a universal classification model that can handle multiple classification tasks and document types through a single integrated framework. Rather than training separate models for different tasks, the system uses one versatile model that adapts to various legal classification requirements through prompt engineering and parameter adjustment.
Solution Approach 2:
The system merges multiple classification functions into a unified model architecture that can perform various classification tasks simultaneously. By combining different classification capabilities within a single model rather than maintaining separate models, the system reduces overall complexity while maintaining comprehensive functionality.
Data Source
AI summary
The following relates generally to using generative AI to: (i) classify documents; (ii) generate prompts (and/or criteria for prompts) to classify documents; (iii) explain document classifications; and/or (iv) explain updates to prompts (and/or prompt criteria). In some embodiments, one or more processors: obtain an initial set of documents from a corpus of documents; classify documents within the initial set of documents by inputting a prompt and the documents within the initial set of documents into a generative artificial intelligence (AI) model; and evaluate classification performance of the prompt to identify (i) that the initial set of documents does not include enough documents associated with a first issue of the one or more issues, or (ii) that the corpus of documents is associated with a new issue.


