Machine Learning Model Fairness Assessment via Synthetic Data Resampling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automated or semi-automated decisioning systems that use machine learning algorithms often introduce or perpetuate undesired and unlawful disparity between different classes or categories of data, leading to unfair predictions.
Innovation Solution
A system and method that enables simultaneous prediction distribution matching with indiscernibility constraints to optimize the learning of a target machine learning model, reducing discrimination between classes, and includes a resampling algorithm to generate synthetic datasets for evaluating candidate models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning algorithms are used in automated decisioning systems, then prediction accuracy and decision-making efficiency are improved, but disparity and unfairness between different classes of data are introduced or perpetuated
Solution Approach 1:
The system performs preliminary actions by generating multiple candidate models before final selection. Each candidate model is trained with different fairness constraints and performance objectives in advance, allowing the system to pre-compute and store model performances on synthetic datasets. This preliminary generation of diverse models enables subsequent efficient selection without retraining, thus maintaining high productivity while addressing disparity issues through pre-prepared fair models.
Solution Approach 2:
The patent uses synthetic datasets that copy and resample from original training data to create multiple virtual versions. These synthetic datasets replicate the statistical properties of the original data while enabling diverse model training scenarios. By copying and transforming the original data into multiple synthetic versions, the system can evaluate model performance across different demographic groups without requiring additional real-world data collection, thus improving fairness while maintaining efficiency.
2Object-affected harmful factors
If multiple candidate models are generated to reduce disparity, then fairness is improved, but the complexity of model selection and assessment increases
Solution Approach 1:
The system segments the model assessment process into distinct phases: (1) generating synthetic datasets, (2) training candidate models on these datasets, (3) evaluating models on multiple metrics (performance, fairness, disparity mitigation), and (4) selecting the best model. This segmentation allows each phase to be optimized independently and enables automated decision-making at each stage, reducing the manual complexity of model selection while maintaining comprehensive fairness evaluation across multiple candidate models.
Solution Approach 2:
The system implements feedback mechanisms where model performances on synthetic datasets are continuously monitored and used to guide further model generation and selection. The assessment metrics (performance efficacy, fairness efficacy, disparity-mitigating viability score) provide feedback that informs which models should be selected or refined. This feedback loop automates the selection process and reduces complexity by using objective criteria rather than manual evaluation, while still thoroughly assessing multiple candidate models for fairness.
3Reliability
If synthetic datasets are generated through resampling, then model evaluation robustness is improved, but computational resources and time are consumed
Solution Approach 1:
The system generates a limited number of synthetic datasets through resampling rather than creating exhaustive datasets. By using resampling techniques, the system creates sufficient synthetic data to robustly evaluate model performance and fairness without requiring complete enumeration of all possible data scenarios. This partial action approach provides adequate robustness for model evaluation while significantly reducing computational resources compared to generating all possible datasets exhaustively.
Data Source
AI summary
A system and method includes obtaining an incumbent model and a candidate model, generating a plurality of synthetic model input datasets, computing, for each synthetic model input dataset, a model performance efficacy metric and a model fairness efficacy metric for the incumbent model based on assessing model output data of the incumbent model that corresponds to each respective synthetic model input dataset of the plurality of synthetic model input datasets, computing, for each synthetic model input dataset, a model performance efficacy metric and a model fairness efficacy metric for the candidate model based on assessing model output data of the candidate model that corresponds to each respective synthetic model input dataset of the plurality of synthetic model input datasets, computing, for the candidate model, a disparity-mitigating model viability score, and displaying, via a graphical user interface, a representation of the candidate model in association with the disparity-mitigating model viability score.


