Document Classifier-Explainer Synthetic QA Generation for Model Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models face challenges with low accuracy and low recall in generating predictions, particularly in question-answer (QA) systems.
Innovation Solution
The generation of synthetic QA training datasets using structured label-explanation datasets and prediction score thresholds to train QA machine learning models, improving prediction accuracy and reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional machine learning models are used for QA predictions, then the system is simple to implement, but the accuracy and recall of predictions are low
Solution Approach 1:
The patent applies preliminary action by generating synthetic QA training datasets before actual model training. The system pre-processes unstructured text data into structured QA pairs with labels and explanations, creating a ready-to-use training corpus that improves subsequent model accuracy without increasing operational complexity
Solution Approach 2:
The patent introduces an intermediary synthetic data generation layer between raw text data and the QA model. This intermediary process creates structured training examples with predicted labels and explanations, acting as a bridge that transforms unstructured data into a format suitable for high-accuracy model training
2Measurement precision
If more training data is collected manually, then the model accuracy improves, but the time and resources required increase significantly
Solution Approach 1:
The patent implements self-service by enabling the system to automatically generate its own training data. The synthetic QA generation process uses existing unstructured text and automated label prediction to create training examples without human intervention, allowing the system to self-populate its training corpus efficiently
Solution Approach 2:
The patent uses copying by generating synthetic QA pairs that replicate the structure and characteristics of real QA data. Instead of manually collecting diverse examples, the system copies the essential patterns from existing text data to create realistic training scenarios that improve model accuracy
Data Source
AI summary
Various embodiments of the present disclosure provide methods, apparatus, systems, computing devices, computing entities, and/or the like for improving question-answer (QA) machine learning model training based on generating predicted label indicators, generating prediction score indicators and prediction explanation indicators, generating structured label-explanation datasets, generating synthetic QA training datasets, generating a prediction output, and initiating the performance of one or more prediction-based operations based on the prediction output.


