Synthetic Training Data Tool for Bias Rectification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

AI/ML models inherit biases and deficiencies from incomplete or flawed training datasets, leading to unreliable performance in applications.

Innovation Solution

A synthetic training data generation tool that identifies and rectifies deficiencies in training datasets by combining them with remediating data, using interpolation and extrapolation techniques, and continuously refining the dataset to remove identified biases and gaps.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a training dataset is used to train AI/ML models, then the models can be developed and deployed, but the models inherit biases and deficiencies from the training dataset, leading to unreliable performance

Engineering Contradiction:
Improvemodel performance reliabilityVSAvoidbiases and deficiencies in training data
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent identifies biases and deficiencies in the training dataset (harmful factors) and converts them into opportunities for improvement by generating targeted synthetic remediating data. The system analyzes the deficiencies and creates synthetic data that specifically addresses these issues, transforming the harmful effect of biased data into a beneficial refinement process that enhances model reliability.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent performs preliminary analysis of the training dataset to identify biases and deficiencies before the actual model training occurs. By conducting this assessment upfront and generating remediating data in advance, the system prevents harmful factors from negatively impacting model performance, rather than addressing issues after they have already affected the training process.

Inventive Principle:
Principle #10Preliminary action

2Quantity of substance

If the training dataset is expanded to include more data sources, then the quantity of training data increases, but the complexity of data integration and processing increases

Engineering Contradiction:
Improvequantity of training dataVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent employs an automated system that self-manages the entire process of identifying data deficiencies, generating synthetic remediating data, and integrating it with the original training dataset. The system automatically analyzes the training data, determines what deficiencies exist, generates appropriate synthetic data without human intervention, and integrates everything seamlessly, thereby reducing processing complexity despite increasing data quantity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the training dataset by changing its composition parameters - adding synthetic remediating data that addresses specific deficiencies while maintaining compatibility with the original data structure. This parameter change approach allows the system to increase data quantity and quality without proportionally increasing processing complexity, as the synthetic data is generated to match existing data schemas and formats.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If synthetic data is generated to rectify deficiencies, then the quality of training data improves, but the time and computational resources required for data processing increase

Engineering Contradiction:
Improvetraining data qualityVSAvoiddata processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary identification of data deficiencies and generates synthetic remediating data before the actual model training begins. By completing the data refinement process upfront, the system avoids time losses during the training phase itself, allowing the model training to proceed efficiently with already-refined data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a continuous process where synthetic remediating data is generated and integrated in an ongoing manner to address deficiencies as they are identified. This continuous refinement approach ensures that data quality improves progressively without requiring complete reprocessing of the entire dataset at once, thereby reducing overall processing time while maintaining high data quality standards.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20250124333A1Method and system for removing deficiencies from a dataset
Publication Date: 2025.04.17 JPMORGAN CHASE BANK NA
  • US20250124333A1 patent drawing
  • US20250124333A1 patent drawing
  • US20250124333A1 patent drawing

AI summary

A system for removing deficiencies from a dataset. The system may include a processor and memory that stores instructions that, when executed by the processor, cause the processor to perform operations. The operations may include removing deficiencies from a dataset that may have been obtained via an input of the synthetic training data generation tool. The removing of deficiencies from the dataset may comprise: determining that the dataset includes training deficiencies; retrieving, from one or more data sources, first remediating data that rectifies a first deficiency; rectifying the first deficiency by updating the dataset with the first remediating data; determining that the updated dataset still includes a training deficiency; and synthesizing second remediating data that rectifies the training deficiency.