Document Classification With Verified Evidence Text Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing document classification systems, particularly in medical coding, are labor-intensive and lack transparency and credibility in their decision-making processes, leading to inefficiencies and potential biases.

Innovation Solution

A dual-machine learning approach is employed, using a generative model to assign categorical identifiers to text segments and a verifier model to refine predictions based on evidence text portions, enhancing explainability and accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional document classification methods are used, then labor intensity is high and processing efficiency is low, but manual processing provides some level of control and oversight

Engineering Contradiction:
Improvedocument processing efficiencyVSAvoidautomation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The patent introduces attribution maps as an intermediary mechanism between the machine learning model and the end user. These attribution maps provide explainability about model decisions without requiring full manual review, thus improving productivity while maintaining appropriate levels of automation. The attribution maps act as a mediator that bridges automated processing with human oversight needs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If machine learning models are used for document classification, then processing speed increases, but explainability and transparency of decision-making decrease

Engineering Contradiction:
Improveclassification processing speedVSAvoiddecision-making transparency
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent implements feedback mechanisms through attribution maps that provide information about why the model made certain classifications. This feedback loop allows users to understand model decisions while maintaining fast automated processing. The attribution maps feed back explanatory information without slowing down the core classification process.

Inventive Principle:
Principle #23Feedback

3Reliability

If traditional attribution maps are used to explain model classifications, then some level of explainability is provided, but explanation robustness and credibility are limited

Engineering Contradiction:
Improveexplanation credibilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-computing and storing attribution information during the model training phase. This allows the system to provide robust and credible explanations without adding complexity to the inference process. The attribution maps are prepared in advance, making them readily available and reliable when needed.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If generative machine learning models are used, then recall is improved, but precision decreases due to excessive prediction of categorical identifiers

Engineering Contradiction:
Improveprediction recallVSAvoidclassification precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent applies partial action by using generative models to generate multiple possible classifications and then filtering these predictions through attribution analysis. Rather than making a single excessive prediction, the system generates multiple candidates and selectively validates them, improving precision while maintaining high recall through the generative approach.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12579179B2Machine learning techniques for classifying document data objects
Publication Date: 2026.03.17 OPTUM INC
  • US12579179B2 patent drawing
  • US12579179B2 patent drawing
  • US12579179B2 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for identifying one or more evidence text portions comprising one or more bases relied on by a generative machine learning model for assigning a plurality of model-assigned categorical identifiers to a plurality of text segment data objects associated with a document data object, and verifying the one or more evidence text portions with a verifier machine learning model to generate one or more classifications of the document data object and provide the one or more verified evidence text portions along with the one or more classifications.