Iterative Synthetic Data Generation for Machine Learning Adaptability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning systems are limited by their dependence on historical data for training, which reduces their reliability and adaptiveness, especially in scenarios with rapid data environment changes or adversarial conditions.

Innovation Solution

An iterative synthetic data generation system using multiple machine learning models and generative adversarial neural networks to identify emerging patterns, broaden their scope, and generate synthetic data sets for continuous training, enabling improved pattern detection and adaptability.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning systems rely on historical data for training, then model training is straightforward and reliable, but the system's adaptiveness to rapid data environment changes deteriorates

Engineering Contradiction:
Improvetraining reliabilityVSAvoidadaptiveness to data changes
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by generating synthetic data in advance that anticipates potential future data patterns and adversarial scenarios. This synthetic data is prepared beforehand and used to pre-train models, enabling them to adapt more quickly when actual data changes occur, thus resolving the contradiction between training reliability and adaptiveness.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of historical data in the form of synthetic data that mimics real data patterns, distributions, and characteristics. These synthetic copies are generated using generative models trained on historical data, allowing the system to expand training datasets without relying solely on additional historical data, thereby improving adaptiveness while maintaining training reliability.

Inventive Principle:
Principle #26Copying

2Measurement precision

If more historical data is collected for training, then model accuracy improves, but the system's responsiveness to real-time changes deteriorates

Engineering Contradiction:
Improvemodel accuracyVSAvoidresponsiveness to changes
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The system generates synthetic copies of historical training data that capture the essential patterns and distributions. These synthetic data copies serve as substitutes for collecting and processing additional real historical data, allowing models to achieve high accuracy through training on synthesized datasets while maintaining the ability to respond quickly to real-time changes without being burdened by large volumes of historical data processing.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameter of data representation by transforming real historical data into synthetic data with modified characteristics. This transformation allows the model to learn from diverse data patterns without being constrained by the volume or recency of actual historical data, thereby improving both accuracy and responsiveness simultaneously.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If synthetic data is generated to expand training data, then adaptiveness improves, but data quality and reliability may deteriorate

Engineering Contradiction:
ImproveadaptivenessVSAvoiddata quality
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system implements feedback mechanisms where generated synthetic data is evaluated against ground truth data and model performance metrics. This feedback loop allows the system to identify and correct quality issues in synthetic data, ensuring that only high-quality synthetic data that maintains reliability is used for training. The feedback process continuously refines the synthetic data generation to preserve data quality while enhancing adaptiveness.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system carefully controls and adjusts parameters of the synthetic data generation process to maintain data quality. By modifying generation parameters such as noise levels, transformation intensities, and distribution matching constraints, the system ensures that synthetic data retains the essential characteristics and reliability of real data while providing the adaptiveness needed for diverse scenarios.

Inventive Principle:
Principle #35Parameter changes

4Manufacturing precision

If iterative refinement of synthetic data is performed, then data quality improves, but processing time and complexity increase

Engineering Contradiction:
Improvesynthetic data qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing and pre-refining synthetic data before it is needed for training. This advance preparation includes generating multiple iterations of synthetic data and pre-evaluating their quality, so that when the data is needed, high-quality refined data is already available. This reduces the need for complex real-time refinement processes and lowers overall system complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial refinement to synthetic data by focusing computational resources on the most critical aspects of data quality improvement rather than exhaustive refinement of all data characteristics. This selective refinement approach achieves sufficient data quality for training purposes while avoiding the excessive processing time and complexity that would result from complete iterative refinement of all data attributes.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11531883B2System and methods for iterative synthetic data generation and refinement of machine learning models
Publication Date: 2022.12.20 BANK OF AMERICA CORP
  • US11531883B2 patent drawing
  • US11531883B2 patent drawing
  • US11531883B2 patent drawing

AI summary

Embodiments of the present invention provide an improvement to convention machine model training techniques by providing an innovative system, method and computer program product for the generation of synthetic data using an iterative process that incorporates multiple machine learning models and neural network approaches. A collaborative system for receiving data and continuously analyzing the data to determine emerging patterns is provided. Common characteristics of data from the identified emerging patterns are broadened in scope and used to generate a synthetic data set using a generative neural network approach. The resulting synthetic data set is narrowed based on analysis of the synthetic data as compared to the detected emerging patterns, and can then be used to further train one or more machine learning models for further pattern detection.