GAN Synthetic Training Data Fairness Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning training data sets often contain biases and inaccuracies, which can lead to unfair or undesirable outcomes in machine learning models.

Innovation Solution

A computer-implemented method is introduced to generate synthetic training data with increased fairness using a generative adversarial network (GAN) based on an initial set of training data and its fairness profile, aiming to improve the fairness score of independent variables.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing machine learning training data sets are used, then the training process can proceed with available data, but the data contains biases and inaccuracies that lead to unfair or undesirable outcomes

Engineering Contradiction:
Improvefairness of machine learning outcomesVSAvoidaccuracy and fairness of training data
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The patent creates synthetic copies of training data through GANs that replicate the statistical properties and distributions of original data while eliminating biases. The generated synthetic data mimics real data patterns but with corrected fairness characteristics, allowing the system to use copied data structures without inheriting original biases

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent modifies key parameters of the training data including fairness scores, demographic distributions, and outcome correlations. By adjusting these parameters through the GAN training process, the system transforms biased data into fair data while maintaining statistical validity and predictive power

Inventive Principle:
Principle #35Parameter changes

2Reliability

If synthetic training data is generated using GANs, then the fairness score of independent variables is improved, but the complexity of the data processing system increases

Engineering Contradiction:
Improvefairness score of training dataVSAvoidcomplexity of data generation system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces a fairness assessment module as an intermediary that evaluates training data quality and guides the GAN generation process. This mediator calculates fairness scores and provides feedback to the GAN, enabling automated control of the complex data generation process without requiring manual intervention

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a feedback loop where fairness scores are continuously calculated on generated synthetic data and fed back to the GAN training process. This feedback mechanism automatically adjusts the generation parameters to improve fairness, reducing the need for complex manual tuning and system management

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250139418A1Generation of statistically fair training data
Publication Date: 2025.05.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250139418A1 patent drawing
  • US20250139418A1 patent drawing
  • US20250139418A1 patent drawing

AI summary

A computer-implemented method to generate training data with increased fairness. The method includes identifying a set of training data comprising at least one independent variable and dependent variable. The method includes analyzing the set of training data to identify correlations between the independent and the dependent variables. The method further includes identifying at least one correlation between the one or more independent variable and the dependent variable. The method includes calculating a fairness score for each first independent variable against the dependent variable. The method further includes creating, based on the analyzing, a fairness profile for the set of training data. The method also includes generating, by a generative adversarial network (GAN) and based on the set of training data and the fairness profile, a set of synthetic training, where the GAN is configured to increases the fairness score for each variable with a disparate effect score above a threshold.