Document Classifier-Explainer Synthetic QA Generation for Model Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models face challenges with low accuracy and low recall in generating predictions, particularly in question-answer (QA) systems.

Innovation Solution

The generation of synthetic QA training datasets using structured label-explanation datasets and prediction score thresholds to train QA machine learning models, improving prediction accuracy and reliability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning models are used for QA predictions, then the system is simple to implement, but the accuracy and recall of predictions are low

Engineering Contradiction:
Improveprediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by generating synthetic QA training datasets before actual model training. The system pre-processes unstructured text data into structured QA pairs with labels and explanations, creating a ready-to-use training corpus that improves subsequent model accuracy without increasing operational complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary synthetic data generation layer between raw text data and the QA model. This intermediary process creates structured training examples with predicted labels and explanations, acting as a bridge that transforms unstructured data into a format suitable for high-accuracy model training

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If more training data is collected manually, then the model accuracy improves, but the time and resources required increase significantly

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata preparation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service by enabling the system to automatically generate its own training data. The synthetic QA generation process uses existing unstructured text and automated label prediction to create training examples without human intervention, allowing the system to self-populate its training corpus efficiently

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses copying by generating synthetic QA pairs that replicate the structure and characteristics of real QA data. Instead of manually collecting diverse examples, the system copies the essential patterns from existing text data to create realistic training scenarios that improve model accuracy

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12443800B2Generation of synthetic question-answer pairs using a document classifier and classification explainer
Publication Date: 2025.10.14 UNITEDHEALTH GROUP INC
  • US12443800B2 patent drawing
  • US12443800B2 patent drawing
  • US12443800B2 patent drawing

AI summary

Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for improving question-answer (QA) machine learning model training based on generating predicted label indicators, generating prediction score indicators and prediction explanation indicators, generating structured label-explanation datasets, generating synthetic QA training datasets, generating a prediction output, and initiating the performance of one or more prediction-based operations based on the prediction output.