Document Classification With Verified Evidence Text Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document classification systems, particularly in medical coding, are labor-intensive and lack transparency and credibility in their decision-making processes, leading to inefficiencies and potential biases.
Innovation Solution
A dual-machine learning approach is employed, using a generative model to assign categorical identifiers to text segments and a verifier model to refine predictions based on evidence text portions, enhancing explainability and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional document classification methods are used, then labor intensity is high and processing efficiency is low, but manual processing provides some level of control and oversight
Solution Approach 1:
The patent introduces attribution maps as an intermediary mechanism between the machine learning model and the end user. These attribution maps provide explainability about model decisions without requiring full manual review, thus improving productivity while maintaining appropriate levels of automation. The attribution maps act as a mediator that bridges automated processing with human oversight needs.
2Speed
If machine learning models are used for document classification, then processing speed increases, but explainability and transparency of decision-making decrease
Solution Approach 1:
The patent implements feedback mechanisms through attribution maps that provide information about why the model made certain classifications. This feedback loop allows users to understand model decisions while maintaining fast automated processing. The attribution maps feed back explanatory information without slowing down the core classification process.
3Reliability
If traditional attribution maps are used to explain model classifications, then some level of explainability is provided, but explanation robustness and credibility are limited
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing attribution information during the model training phase. This allows the system to provide robust and credible explanations without adding complexity to the inference process. The attribution maps are prepared in advance, making them readily available and reliable when needed.
4Reliability
If generative machine learning models are used, then recall is improved, but precision decreases due to excessive prediction of categorical identifiers
Solution Approach 1:
The patent applies partial action by using generative models to generate multiple possible classifications and then filtering these predictions through attribution analysis. Rather than making a single excessive prediction, the system generates multiple candidates and selectively validates them, improving precision while maintaining high recall through the generative approach.
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for identifying one or more evidence text portions comprising one or more bases relied on by a generative machine learning model for assigning a plurality of model-assigned categorical identifiers to a plurality of text segment data objects associated with a document data object, and verifying the one or more evidence text portions with a verifier machine learning model to generate one or more classifications of the document data object and provide the one or more verified evidence text portions along with the one or more classifications.


