Fairness-Aware Training Data Selection for AI Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models often become biased when trained with data sets that favor certain groups or data points, leading to unpredictable behavior when encountering inputs outside the trained feature space, resulting in biased decision-making and inefficient training.
Innovation Solution
A method for generating fairness-aware training data sets by calculating diversity scores and model attribution scores to select and prioritize data points that are diverse and informative, thereby creating a balanced and unbiased training dataset for AI models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional training data sampling is used, then training efficiency is maintained, but model bias increases and fairness deteriorates
Solution Approach 1:
The patent transforms the training data selection process by changing parameters from traditional random or stratified sampling to a fairness-aware sampling method that incorporates diversity scores and model attribution scores. This parameter change enables the system to select data points that maximize both diversity and informational value, resolving the contradiction between model fairness and training efficiency by redefining the selection criteria rather than increasing computational overhead.
Solution Approach 2:
The patent replaces traditional mechanical sampling methods (random sampling, stratified sampling) with an intelligent selection mechanism that uses diversity scores and model attribution scores. This substitution eliminates the need for manual bias correction or extensive retraining, as the new selection mechanism inherently produces fairer models by prioritizing diverse and informative data points, thus maintaining training efficiency while improving fairness.
2Reliability
If diverse training data is prioritized, then model fairness improves, but data selection complexity increases
Solution Approach 1:
The patent segments the data selection process into two independent scoring components: diversity scores (measuring representativeness across different groups) and model attribution scores (measuring informational value for model learning). This segmentation simplifies the overall complexity by breaking down the complex fairness-aware selection into manageable, independently calculable metrics that can be combined to guide data selection.
Solution Approach 2:
The patent introduces diversity scores and model attribution scores as intermediary metrics that mediate between the raw training data and the model training process. These intermediary scores serve as proxies for fairness and informational value, simplifying the selection process by providing clear, quantifiable criteria for data point selection without requiring complex fairness algorithms or manual intervention.
3Adaptability or versatility
If representative data sampling is used, then generalization ability improves, but training data requirements increase
Solution Approach 1:
The patent enables the model training process to self-select representative training data by using model attribution scores that identify which data points provide the most informational value for learning. This self-service mechanism allows the system to automatically prioritize high-value data points that maximize generalization ability, eliminating the need to uniformly sample large quantities of data across all categories and thus reducing overall training data requirements.
Solution Approach 2:
The patent changes the selection parameter from uniform or random sampling to a value-based selection using model attribution scores. This parameter change enables the system to achieve better generalization with fewer data points by selectively sampling high-information-value examples rather than requiring large quantities of uniformly distributed data, thus resolving the contradiction between generalization ability and data quantity requirements.
Data Source
AI summary
There are provided systems and methods for adjustment of training data sets for fairness-aware artificial intelligence models. A service provider, such as an electronic transaction processor for digital transactions, may utilize different decision services that implement rules and/or artificial intelligence models for decision-making of data including data in production computing environment. Decision services may be used for data processing and decision-making, where multiple decision services may be invoked during run-time in order to complete a data processing request. When processing data, machine learning and other artificial intelligence models may be utilized by such decision services. These may be trained using a sampled training data set that takes into account data records' diversity and model attribution scores as providing valuable data points or observations for training and/or retraining the ML model. The sampled training data may be analyzed to determine these scores and generated for training.


