Class-Conditioned Synthetic Data Retraining for Statistical Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional synthetic data generators lack class-awareness, produce unrealistic artifacts, and do not support recursive simulation dynamics to mimic real-world data characteristics, leading to degraded model performance and bias propagation in AI deployments, especially in data-scarce or regulated domains.

Innovation Solution

A novel architecture integrating class-specific ensemble modeling, structured noise injection, and recursive feature integration, with fidelity monitoring using divergence metrics like Wasserstein distance and Kullback-Leibler divergence, triggering retraining cycles for affected classes via a per-class update mechanism.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional generative approaches are used for synthetic data generation, then the generation process is simple, but the class-conditional feature distributions are not preserved and statistical fidelity is poor

Engineering Contradiction:
Improvestatistical fidelityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the synthetic data generation process into distinct components: a class-conditioned sample generator that processes different class labels separately, and a fidelity evaluation module that assesses statistical properties. This segmentation allows each component to specialize in maintaining class-conditional feature distributions, thereby improving statistical fidelity without requiring complete system redesign

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a feedback mechanism where the fidelity evaluation module continuously assesses the statistical properties of generated synthetic data and provides guidance back to the class-conditioned sample generator. This closed-loop feedback ensures that class-conditional feature distributions are preserved while allowing the system to adapt and improve statistical fidelity over time

Inventive Principle:
Principle #23Feedback

2Reliability

If conventional synthetic data generators are used, then the system is easy to operate, but unrealistic artifacts are produced and model performance degrades

Engineering Contradiction:
Improvemodel performanceVSAvoidoperational simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces dynamic class-conditioning into the sample generator, allowing the system to adapt its generation process based on the specific class label being processed. This dynamic approach enables the generator to maintain appropriate feature distributions for each class, producing more realistic synthetic data that improves model performance without requiring manual intervention

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent replaces simple random sampling mechanisms with a sophisticated class-conditioned generation system that uses learned representations and statistical modeling. This substitution eliminates the production of unrealistic artifacts while maintaining automated operation, thereby improving reliability without significantly increasing operational complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Manufacturing precision

If traditional generative models are used, then the processing speed is fast, but class balance and statistical fidelity are not maintained

Engineering Contradiction:
Improveclass balanceVSAvoidgeneration speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent segments the generation process by class label, allowing parallel processing of different classes while maintaining their specific statistical properties. This segmentation enables the system to preserve class balance and statistical fidelity without creating a sequential bottleneck, thereby maintaining generation speed while improving precision

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent dynamically adjusts generation parameters based on the target class label and desired statistical properties. By changing parameters such as feature distribution constraints and sampling strategies according to class-specific requirements, the system maintains both class balance and generation efficiency without requiring slow iterative refinement

Inventive Principle:
Principle #35Parameter changes

4Stability of the object's composition

If synthetic data is generated without fidelity monitoring, then the generation process is quick, but bias propagation occurs and statistical consistency is lost

Engineering Contradiction:
Improvestatistical consistencyVSAvoidevaluation time
Core Design Contradiction:
Stability of the object's compositionVSLoss of time

Solution Approach 1:

The patent implements targeted fidelity evaluation that monitors key statistical properties of the generated synthetic data and provides feedback to adjust the generation process. This feedback mechanism detects distributional shifts and bias propagation early, allowing corrective actions to be taken before significant statistical consistency is lost, thereby maintaining stability without requiring exhaustive evaluation

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent extracts and monitors only the critical statistical properties and class-conditional feature distributions that are essential for maintaining statistical consistency. By focusing evaluation efforts on these key metrics rather than进行全面 analysis, the system achieves statistical stability with minimal evaluation time, preventing bias propagation without excessive time loss

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20250384289A1Generating class-balanced synthetic data with fidelity-guided retraining
Publication Date: 2025.12.18 THE TORONTO DOMINION BANK
  • US20250384289A1 patent drawing
  • US20250384289A1 patent drawing
  • US20250384289A1 patent drawing

AI summary

An example operation may include at least one of producing, by a class-conditioned sample generator executing on at least one processor communicatively coupled to a memory on a host platform, a synthetic feature set based on a label sequence and class information derived from received data, transmitting, by the host platform, a finalized synthetic sample to a computing device when the synthetic feature set satisfies a fidelity threshold, generating, by the computing device, a fidelity score based on a comparison of the finalized synthetic sample to the label sequence and the class information, retraining, by the computing device, the class-conditioned sample generator based on the fidelity score, and validating, by the computing device, the class-conditioned sample generator by transmitting a test prompt to the host platform, receiving a synthetic response generated by the class-conditioned sample generator, and comparing the synthetic response to previously stored synthetic data to validate the class-conditioned sample generator.