Synthetic Training Data Tool for Bias Rectification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
AI/ML models inherit biases and deficiencies from incomplete or flawed training datasets, leading to unreliable performance in applications.
Innovation Solution
A synthetic training data generation tool that identifies and rectifies deficiencies in training datasets by combining them with remediating data, using interpolation and extrapolation techniques, and continuously refining the dataset to remove identified biases and gaps.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a training dataset is used to train AI/ML models, then the models can be developed and deployed, but the models inherit biases and deficiencies from the training dataset, leading to unreliable performance
Solution Approach 1:
The patent identifies biases and deficiencies in the training dataset (harmful factors) and converts them into opportunities for improvement by generating targeted synthetic remediating data. The system analyzes the deficiencies and creates synthetic data that specifically addresses these issues, transforming the harmful effect of biased data into a beneficial refinement process that enhances model reliability.
Solution Approach 2:
The patent performs preliminary analysis of the training dataset to identify biases and deficiencies before the actual model training occurs. By conducting this assessment upfront and generating remediating data in advance, the system prevents harmful factors from negatively impacting model performance, rather than addressing issues after they have already affected the training process.
2Quantity of substance
If the training dataset is expanded to include more data sources, then the quantity of training data increases, but the complexity of data integration and processing increases
Solution Approach 1:
The patent employs an automated system that self-manages the entire process of identifying data deficiencies, generating synthetic remediating data, and integrating it with the original training dataset. The system automatically analyzes the training data, determines what deficiencies exist, generates appropriate synthetic data without human intervention, and integrates everything seamlessly, thereby reducing processing complexity despite increasing data quantity.
Solution Approach 2:
The patent transforms the training dataset by changing its composition parameters - adding synthetic remediating data that addresses specific deficiencies while maintaining compatibility with the original data structure. This parameter change approach allows the system to increase data quantity and quality without proportionally increasing processing complexity, as the synthetic data is generated to match existing data schemas and formats.
3Reliability
If synthetic data is generated to rectify deficiencies, then the quality of training data improves, but the time and computational resources required for data processing increase
Solution Approach 1:
The patent performs preliminary identification of data deficiencies and generates synthetic remediating data before the actual model training begins. By completing the data refinement process upfront, the system avoids time losses during the training phase itself, allowing the model training to proceed efficiently with already-refined data.
Solution Approach 2:
The patent implements a continuous process where synthetic remediating data is generated and integrated in an ongoing manner to address deficiencies as they are identified. This continuous refinement approach ensures that data quality improves progressively without requiring complete reprocessing of the entire dataset at once, thereby reducing overall processing time while maintaining high data quality standards.
Data Source
AI summary
A system for removing deficiencies from a dataset. The system may include a processor and memory that stores instructions that, when executed by the processor, cause the processor to perform operations. The operations may include removing deficiencies from a dataset that may have been obtained via an input of the synthetic training data generation tool. The removing of deficiencies from the dataset may comprise: determining that the dataset includes training deficiencies; retrieving, from one or more data sources, first remediating data that rectifies a first deficiency; rectifying the first deficiency by updating the dataset with the first remediating data; determining that the updated dataset still includes a training deficiency; and synthesizing second remediating data that rectifies the training deficiency.


