Cluster Targeting for Machine Learning Bias Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning models often suffer from biases and errors due to large, unclean datasets, making it difficult to accurately represent less common situations and leading to mislabeling or undesirable outcomes, which are costly and time-consuming to correct.
Innovation Solution
A system and method that uses unsupervised learning to inspect and improve supervised machine learning models by clustering data, identifying biases, and allowing for subjective interpretation to refine the training process, enabling more precise data curation and ethical AI behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If large datasets are used to train machine learning models, then model performance may improve, but data cleaning costs and time increase significantly
Solution Approach 1:
The patent applies preliminary action by using unsupervised learning models to pre-inspect and cluster training data before supervised learning. This preliminary clustering identifies potential biases and edge cases in advance, allowing data scientists to clean and curate data more efficiently before the main training process, thereby reducing the time and cost of data cleaning while maintaining model performance.
Solution Approach 2:
The patent introduces an intermediary unsupervised learning model that acts as a mediator between raw training data and the supervised learning model. This intermediary clusters the data and identifies biases, serving as a bridge that enables more efficient data cleaning and preparation processes, ultimately reducing the time required for data cleaning while improving model reliability.
2Quantity of substance
If large datasets are used to train machine learning models, then more training data is available, but visibility into data quality and patterns decreases
Solution Approach 1:
The patent applies segmentation by using unsupervised learning to divide large datasets into distinct clusters based on similarities and patterns. This segmentation provides visibility into data quality and characteristics by organizing vast amounts of data into manageable groups, allowing data scientists to inspect and understand data patterns that would be invisible in raw large-scale datasets.
Solution Approach 2:
The unsupervised learning model serves as an intermediary that processes large datasets and transforms them into clustered representations with visible patterns. This intermediary layer maintains the quantity of training data while recovering information about data quality and patterns through clustering, thereby preventing loss of information about data characteristics.
3Ease of manufacture
If traditional supervised learning is used without inspection, then training process is straightforward, but biases and errors in data are amplified
Solution Approach 1:
The patent applies preliminary action by inserting an unsupervised learning inspection step before supervised learning. This preliminary clustering and bias identification maintains relative simplicity of the training process while preventing the amplification of biases and errors, thereby improving model accuracy without significantly complicating the overall training workflow.
4Reliability
If data cleaning is performed manually to correct biases, then model accuracy can be improved, but the process is slow and expensive
Solution Approach 1:
The patent applies self-service by using unsupervised learning models to automatically inspect, cluster, and identify biases in training data without requiring extensive manual intervention. This automated inspection process significantly improves data cleaning efficiency and productivity while maintaining the ability to achieve high model accuracy through targeted corrections based on unsupervised insights.
Solution Approach 2:
The unsupervised learning model acts as an intermediary that automates the data inspection process, replacing slow and expensive manual data cleaning. This intermediary provides automated bias identification and clustering, thereby improving productivity while still enabling the achievement of high model accuracy through the insights it provides.
Data Source
AI summary
A system and method for training, using a supervised learning process, a first learning model with a first dataset; applying the first learning model to a second dataset thereby generating a first learning model output; training, using an unsupervised learning process, a second learning model with the first learning model output thereby generating a clustering output of the second learning model; determining a bias assessment based on the clustering output; and training, using a third dataset, a bias assessment modified learning model using supervised learning.


