AI Document Analysis With Reusable Predictive Coding Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document analysis systems face inefficiencies in iterative human coding processes for large volumes of electronic documents, particularly in contexts like litigation, leading to wasteful labor and time consumption due to the need for repeated coding of similar corpora.
Innovation Solution
The reuse of previously coded datasets and trained models across corpora through predictive coding, allowing for the formation of boosted codes and models that can be applied to new corpora without additional human input, leveraging machine learning to enhance predictive scoring and reduce labor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional iterative human coding processes are used for large volumes of electronic documents, then coding accuracy can be maintained through human review, but labor consumption and time requirements increase significantly
Solution Approach 1:
The system performs preliminary coding actions by training machine learning models on previously coded datasets before applying them to new corpora. This preliminary training phase captures coding patterns and decisions, enabling automated reuse of coding logic without requiring human reviewers to manually code each document from scratch, thereby reducing labor consumption while maintaining accuracy through the pretrained model's understanding of coding standards.
Solution Approach 2:
The system creates copies of previously coded datasets and trained models to apply to new corpora. By copying the learned coding patterns and model parameters from source corpora to target corpora, the system replicates successful coding approaches without requiring identical human review processes, thus reducing labor consumption while preserving coding accuracy through the transferred knowledge.
2Adaptability or versatility
If coding is performed manually for each new corpus, then coding decisions can be tailored to specific corpus characteristics, but time consumption increases due to repeated coding processes
Solution Approach 1:
The system creates universal trained models that can be applied across multiple different corpora. These models learn general coding patterns that are adaptable to various corpus characteristics while maintaining consistent coding standards. The same model can serve multiple functions across different document sets, reducing time consumption by eliminating the need to create separate coding frameworks for each corpus while still adapting to specific characteristics through the model's generalized understanding.
Solution Approach 2:
The system adjusts model parameters and training configurations based on the specific characteristics of different corpora. By changing parameters such as training data selection, model architecture adjustments, or hyperparameter optimization tailored to each corpus's unique features, the system maintains adaptability to specific corpus characteristics while using automated processes to reduce the time required compared to manual tailoring for each corpus.
3Productivity
If previously coded datasets are reused across corpora, then labor is reduced and efficiency improves, but data availability may be limited when data is ephemeral or segregated
Solution Approach 1:
The system extracts the essential coding knowledge and patterns from previously coded datasets by training machine learning models on the data. Once the model is trained, the actual source data can be removed or archived while the model retains the learned coding logic. This extraction process allows the system to reuse coding knowledge without requiring continuous access to the original ephemeral or segregated data sources, maintaining efficiency while overcoming data availability limitations.
Solution Approach 2:
The trained machine learning model serves as an intermediary between the original coded datasets and new corpora. The model captures and stores the coding knowledge from source data, then applies this knowledge to target corpora without requiring direct access to the original data. This intermediary approach enables efficient knowledge transfer while bypassing data availability constraints, as the model preserves the essential coding patterns independently of the source data's availability.
Data Source
AI summary
Artificial intelligence based document analysis systems and methods are disclosed. Embodiments of document analysis systems may allow the reuse of coded datasets defined in association with a particular code by allowing these datasets to be bundled to define a dataset for another code, where that code may be associated with a target corpus of documents. A model can then be trained based on that dataset and used to provide predictive scores for the documents of the target corpora with respect to the code. Furthermore, this code can be applied not just to the target corpus of documents, but additionally can be applied against any other corpora.


