AI Bias Removal via Population Segmentation and Subpopulation Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing artificial intelligence algorithms often exhibit bias due to training data, leading to unfair outcomes, as they may focus on proxy factors even when prohibited from considering specific traits, such as race, resulting in amplified and reinforced biases.
Innovation Solution
A platform that segments populations by specific traits into subpopulations and trains separate models for each, allowing for the specification of ratios or amounts to be selected from each subpopulation, thereby reducing bias by explicitly setting conditions on the results and enabling continuous learning to identify value within subgroups.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single AI model is trained on the entire population without segmentation, then the model achieves broad applicability and simplicity, but bias emerges due to systematic errors that privilege certain groups over others
Solution Approach 1:
The patent divides the population into distinct subpopulations based on protected traits (e.g., race, gender, age). Instead of training a single model on the entire population, separate models are trained for each subpopulation. This segmentation allows the system to account for systematic errors that affect different groups differently, thereby reducing algorithmic bias while maintaining model effectiveness.
2Reliability
If the AI model is trained to consider all population members equally, then fairness across groups is improved, but the model fails to account for systematic errors that differently affect various groups
Solution Approach 1:
By segmenting the population into subpopulations based on protected traits, the patent enables the system to identify and correct systematic errors that differently affect various groups. Each subpopulation model can be optimized for its specific characteristics, improving both fairness and prediction accuracy simultaneously.
Solution Approach 2:
The patent applies different model configurations and training approaches to different subpopulations based on their specific characteristics. Each subpopulation receives tailored modeling that accounts for its unique patterns and systematic errors, rather than applying a uniform approach to all population members.
3Object-affected harmful factors
If separate models are trained for each subpopulation, then bias is reduced and fairness is improved, but the system complexity increases
Solution Approach 1:
While segmentation into multiple models does increase complexity, the patent manages this by implementing a modular architecture where each subpopulation model is independent and can be developed, tested, and maintained separately. This modular approach makes the complexity manageable and allows for systematic deployment.
Solution Approach 2:
The patent introduces an intermediary layer that manages the multiple subpopulation models. This intermediary handles model selection, coordination, and aggregation of results, thereby managing the complexity of having multiple models while maintaining the bias-reduction benefits of segmentation.
4Adaptability or versatility
If proxy factors are used to make predictions when direct consideration of protected traits is prohibited, then legal compliance is maintained, but bias is amplified and reinforced
Solution Approach 1:
By segmenting the population based on protected traits before model training, the patent allows the system to directly account for these traits in a lawful manner. This approach complies with legal requirements while preventing the amplification of bias that occurs when proxy factors are used, as the segmented models can directly learn from subpopulation-specific patterns without relying on problematic proxies.
Data Source
AI summary
Data is received characterizing a population and a target trait characteristic for selecting candidates from the population. The population is segmented into at least a first subpopulation and a second subpopulation. A first number of candidates is selected from the first subpopulation and using a first model. The first number of candidates is selected according to the target trait characteristic. The first model having been trained using a first training population in which all members of the first training population are part of the first class of the two or more classes. A second number of candidates is selected from the second subpopulation and using a second model. The second model having been trained using a second training population in which all members of the second training population are part of the second class of the two or more classes. Related apparatus, systems, techniques and articles are also described.


