Cluster Targeting for Machine Learning Bias Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models often suffer from biases and errors due to large, unclean datasets, making it difficult to accurately represent less common situations and leading to mislabeling or undesirable outcomes, which are costly and time-consuming to correct.

Innovation Solution

A system and method that uses unsupervised learning to inspect and improve supervised machine learning models by clustering data, identifying biases, and allowing for subjective interpretation to refine the training process, enabling more precise data curation and ethical AI behavior.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If large datasets are used to train machine learning models, then model performance may improve, but data cleaning costs and time increase significantly

Engineering Contradiction:
Improvemodel performanceVSAvoiddata cleaning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by using unsupervised learning models to pre-inspect and cluster training data before supervised learning. This preliminary clustering identifies potential biases and edge cases in advance, allowing data scientists to clean and curate data more efficiently before the main training process, thereby reducing the time and cost of data cleaning while maintaining model performance.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary unsupervised learning model that acts as a mediator between raw training data and the supervised learning model. This intermediary clusters the data and identifies biases, serving as a bridge that enables more efficient data cleaning and preparation processes, ultimately reducing the time required for data cleaning while improving model reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If large datasets are used to train machine learning models, then more training data is available, but visibility into data quality and patterns decreases

Engineering Contradiction:
Improvetraining data volumeVSAvoiddata quality visibility
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent applies segmentation by using unsupervised learning to divide large datasets into distinct clusters based on similarities and patterns. This segmentation provides visibility into data quality and characteristics by organizing vast amounts of data into manageable groups, allowing data scientists to inspect and understand data patterns that would be invisible in raw large-scale datasets.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The unsupervised learning model serves as an intermediary that processes large datasets and transforms them into clustered representations with visible patterns. This intermediary layer maintains the quantity of training data while recovering information about data quality and patterns through clustering, thereby preventing loss of information about data characteristics.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of manufacture

If traditional supervised learning is used without inspection, then training process is straightforward, but biases and errors in data are amplified

Engineering Contradiction:
Improvetraining process simplicityVSAvoidmodel accuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies preliminary action by inserting an unsupervised learning inspection step before supervised learning. This preliminary clustering and bias identification maintains relative simplicity of the training process while preventing the amplification of biases and errors, thereby improving model accuracy without significantly complicating the overall training workflow.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If data cleaning is performed manually to correct biases, then model accuracy can be improved, but the process is slow and expensive

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata cleaning efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies self-service by using unsupervised learning models to automatically inspect, cluster, and identify biases in training data without requiring extensive manual intervention. This automated inspection process significantly improves data cleaning efficiency and productivity while maintaining the ability to achieve high model accuracy through targeted corrections based on unsupervised insights.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The unsupervised learning model acts as an intermediary that automates the data inspection process, replacing slow and expensive manual data cleaning. This intermediary provides automated bias identification and clustering, thereby improving productivity while still enabling the achievement of high model accuracy through the insights it provides.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20230297886A1Cluster targeting for use in machine learning
Publication Date: 2023.09.21 GRABANGO CO
  • US20230297886A1 patent drawing
  • US20230297886A1 patent drawing
  • US20230297886A1 patent drawing

AI summary

A system and method for training, using a supervised learning process, a first learning model with a first dataset; applying the first learning model to a second dataset thereby generating a first learning model output; training, using an unsupervised learning process, a second learning model with the first learning model output thereby generating a clustering output of the second learning model; determining a bias assessment based on the clustering output; and training, using a third dataset, a bias assessment modified learning model using supervised learning.