Fairness-Aware Training Data Selection for AI Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models often become biased when trained with data sets that favor certain groups or data points, leading to unpredictable behavior when encountering inputs outside the trained feature space, resulting in biased decision-making and inefficient training.

Innovation Solution

A method for generating fairness-aware training data sets by calculating diversity scores and model attribution scores to select and prioritize data points that are diverse and informative, thereby creating a balanced and unbiased training dataset for AI models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional training data sampling is used, then training efficiency is maintained, but model bias increases and fairness deteriorates

Engineering Contradiction:
Improvemodel fairnessVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent transforms the training data selection process by changing parameters from traditional random or stratified sampling to a fairness-aware sampling method that incorporates diversity scores and model attribution scores. This parameter change enables the system to select data points that maximize both diversity and informational value, resolving the contradiction between model fairness and training efficiency by redefining the selection criteria rather than increasing computational overhead.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical sampling methods (random sampling, stratified sampling) with an intelligent selection mechanism that uses diversity scores and model attribution scores. This substitution eliminates the need for manual bias correction or extensive retraining, as the new selection mechanism inherently produces fairer models by prioritizing diverse and informative data points, thus maintaining training efficiency while improving fairness.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If diverse training data is prioritized, then model fairness improves, but data selection complexity increases

Engineering Contradiction:
Improvemodel fairnessVSAvoiddata selection complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the data selection process into two independent scoring components: diversity scores (measuring representativeness across different groups) and model attribution scores (measuring informational value for model learning). This segmentation simplifies the overall complexity by breaking down the complex fairness-aware selection into manageable, independently calculable metrics that can be combined to guide data selection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces diversity scores and model attribution scores as intermediary metrics that mediate between the raw training data and the model training process. These intermediary scores serve as proxies for fairness and informational value, simplifying the selection process by providing clear, quantifiable criteria for data point selection without requiring complex fairness algorithms or manual intervention.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If representative data sampling is used, then generalization ability improves, but training data requirements increase

Engineering Contradiction:
Improvegeneralization abilityVSAvoidtraining data quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent enables the model training process to self-select representative training data by using model attribution scores that identify which data points provide the most informational value for learning. This self-service mechanism allows the system to automatically prioritize high-value data points that maximize generalization ability, eliminating the need to uniformly sample large quantities of data across all categories and thus reducing overall training data requirements.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the selection parameter from uniform or random sampling to a value-based selection using model attribution scores. This parameter change enables the system to achieve better generalization with fewer data points by selectively sampling high-information-value examples rather than requiring large quantities of uniformly distributed data, thus resolving the contradiction between generalization ability and data quantity requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240177051A1Adjustment of training data sets for fairness-aware artificial intelligence models
Publication Date: 2024.05.30 PAYPAL INC
  • US20240177051A1 patent drawing
  • US20240177051A1 patent drawing
  • US20240177051A1 patent drawing

AI summary

There are provided systems and methods for adjustment of training data sets for fairness-aware artificial intelligence models. A service provider, such as an electronic transaction processor for digital transactions, may utilize different decision services that implement rules and/or artificial intelligence models for decision-making of data including data in production computing environment. Decision services may be used for data processing and decision-making, where multiple decision services may be invoked during run-time in order to complete a data processing request. When processing data, machine learning and other artificial intelligence models may be utilized by such decision services. These may be trained using a sampled training data set that takes into account data records' diversity and model attribution scores as providing valuable data points or observations for training and/or retraining the ML model. The sampled training data may be analyzed to determine these scores and generated for training.