Model Comparison via Tournament Ranking for Bias Evaluation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models face challenges in comparing and evaluating multiple evaluation criteria for demographic bias, as different criteria yield non-comparable scores, making it difficult to determine the best performance and identify the most important bias criterion, leading to unfair and discriminatory outcomes.
Innovation Solution
A computer-implemented method and system that generates raw scores for models based on multiple measures of demographic bias and performance, creates a raw score matrix, determines rank scores, and uses a tournament matrix to compare models pairwise, ultimately selecting and presenting the least biased model via a user interface.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple measures of demographic bias are used to evaluate models, then the comprehensiveness of bias evaluation is improved, but the comparability of scores across different measures deteriorates
Solution Approach 1:
The patent introduces an intermediary normalization process that transforms scores from different bias measures into a common scale. The normalization module standardizes scores across multiple demographic bias measures (e.g., demographic parity, equal opportunity) so they become comparable, while preserving the comprehensiveness of evaluating multiple bias types simultaneously.
Solution Approach 2:
The patent transforms the parameter space by converting raw bias scores into normalized scores through mathematical transformation. This parameter change allows scores from different bias measures to be expressed in a unified framework, enabling direct comparison while maintaining the ability to evaluate multiple bias dimensions.
2Reliability
If multiple bias criteria are evaluated simultaneously, then the thoroughness of model assessment is improved, but the difficulty of determining overall best performance increases
Solution Approach 1:
The patent segments the evaluation process into distinct modular components: a scoring module that evaluates individual bias criteria, a normalization module that standardizes scores, and an aggregation module that combines normalized scores. This segmentation makes the overall complex evaluation process manageable and systematic while maintaining thorough assessment across multiple criteria.
Solution Approach 2:
The patent creates a universal evaluation framework that handles multiple different bias criteria through a single integrated system. The aggregation module universally combines scores from various bias measures using a consistent method, providing a unified approach to determining overall best performance regardless of the specific number or type of bias criteria being evaluated.
3Adaptability or versatility
If different bias scoring methods are applied, then the coverage of bias detection is improved, but the agreement on identifying the most important bias criterion deteriorates
Solution Approach 1:
The patent incorporates feedback mechanisms where the aggregation module uses weighted combinations of normalized scores from multiple bias criteria. The weighting system provides feedback on the relative importance of different bias measures, allowing the system to identify and rank the most important bias criteria based on their contribution to overall model evaluation, thereby achieving agreement on importance ranking.
Data Source
AI summary
Systems and methods are disclosed for comparing a plurality of models. The method includes generating raw scores for the plurality of models based on multiple measures of demographic bias and performance. The raw scores for each of the plurality of models are stored in corresponding locations of a raw score matrix. The rank scores for the plurality of models are determined based on comparing the raw scores of the plurality models in each of the multiple measures of demographic bias and performance. The rank scores for each of the plurality of models are stored in corresponding locations of a rank matrix. Tournament scores for the plurality of models are determined based on performing a pairwise comparison of the rank scores. The tournament scores are stored in corresponding locations of a tournament matrix. The tournament scores are tallied to determine a rank for each of the plurality of models.


