Fairness-Aware Data Valuation for ML Bias Mitigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data valuation methods in machine learning are not performance and fairness-aware, leading to the propagation and exacerbation of biases in high-stakes decision-making settings, where biased models can discriminate against certain social groups.
Innovation Solution
A fairness-aware data valuation framework that uses an entropy-based notion of value and utility to measure the contribution of training instances to both performance and fairness, incorporating protected attributes and promoting subgroup fairness, allowing for the identification of data bias and its mitigation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data valuation is performed using traditional methods focused only on performance metrics, then model accuracy is improved, but bias propagation and discrimination against certain social groups worsens
Solution Approach 1:
The patent changes the valuation parameters from solely performance-based metrics to a dual metric system that incorporates both performance and fairness measures. This allows the data valuation framework to assess training instances based on their contribution to both model accuracy and fairness, thereby preventing bias propagation while maintaining reliability.
Solution Approach 2:
The patent introduces fairness-aware data valuation as an intermediary layer between data selection and model training. This intermediary assesses the fairness impact of training instances before they are used for training, acting as a mediator that prevents biased data from negatively affecting the model while still allowing accurate data to be utilized.
2Object-affected harmful factors
If fairness-aware data valuation is implemented to mitigate bias, then discrimination against social groups is reduced, but model accuracy may deteriorate
Solution Approach 1:
The patent employs parameter changes by adjusting the valuation metrics to include both fairness and performance dimensions. By modifying the assessment parameters to consider dual objectives, the system can identify training instances that maintain accuracy while reducing discrimination, thus resolving the trade-off between fairness and model performance.
3Productivity
If data valuation focuses exclusively on performance metrics, then computational efficiency is maintained, but fairness assessment capability is lost
Solution Approach 1:
The patent implements multi-functionality by creating a unified data valuation framework that simultaneously performs both performance assessment and fairness evaluation. This universal approach allows the same system to handle multiple objectives (accuracy and fairness) without requiring separate computational processes, thereby maintaining efficiency while gaining fairness assessment capability.
Data Source
AI summary
The present document discloses a method for machine-learning fairness-aware data valuation processing for supervised learning, from a training dataset comprising a plurality of data records each containing data for a training instance for said supervised learning, wherein the data for each training instance comprises one or more target variables, one or more input variables and one or more protected-attribute variables, wherein fairness is defined as minimizing a data bias present in the training set in respect of the one or more protected variables.


