Confusion Metric Validation for ML Label Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning algorithms are hindered by faulty or problematic data, leading to flawed classifications due to the coexistence of new, old, and problematic taxonomies in software applications, which can hinder data retrieval and application performance.

Innovation Solution

A machine learning model is trained to predict labels for data items, validated using confusion metrics to identify and resolve problematic labels, and then retrained based on these resolutions to generate more accurate labels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If machine learning algorithms are trained with existing taxonomies, then classification capability is provided, but classification accuracy deteriorates due to faulty or problematic data

Engineering Contradiction:
Improveclassification accuracyVSAvoidfaulty or problematic data
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system performs preliminary validation of training data against a knowledge graph before training the machine learning model. This preliminary action identifies and filters out faulty or problematic data points, ensuring that only high-quality data is used for training, thereby improving classification accuracy without requiring post-training correction

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

A knowledge graph is introduced as an intermediary between the raw taxonomy data and the machine learning model. The knowledge graph validates and cleans the training data by checking for contradictions, typos, and problematic entries, acting as a mediator that purifies the data before it reaches the classifier

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple taxonomies coexist in the system, then taxonomy flexibility is maintained, but data retrieval performance deteriorates due to excess taxonomies

Engineering Contradiction:
Improvetaxonomy flexibilityVSAvoiddata retrieval performance
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system continuously monitors data retrieval performance and uses this feedback to identify problematic taxonomies. When performance deteriorates due to conflicting or redundant taxonomies, the system automatically adjusts by removing or merging problematic entries, maintaining flexibility while improving retrieval efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system extracts and removes problematic taxonomies from the overall taxonomy set. By identifying and taking out conflicting, redundant, or faulty taxonomies, the system maintains a cleaner, more efficient taxonomy structure that improves data retrieval performance while preserving legitimate taxonomy flexibility

Inventive Principle:
Principle #2Taking out (Extraction)

3Stability of the object's composition

If problematic taxonomies are not purged, then backward compatibility is maintained, but system performance deteriorates due to excess taxonomies

Engineering Contradiction:
Improvebackward compatibilityVSAvoidsystem performance
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

The system applies local quality improvement by selectively removing only problematic taxonomies while preserving valid ones. This localized approach maintains backward compatibility for legitimate use cases while improving system performance by eliminating conflicting or redundant taxonomies that hinder data retrieval

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11636387B2System and method for improving machine learning models based on confusion error evaluation
Publication Date: 2023.04.25 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11636387B2 patent drawing
  • US11636387B2 patent drawing
  • US11636387B2 patent drawing

AI summary

Embodiments described herein are directed to improving machine learning (ML) model-based techniques for automatically labeling data items based on identifying and resolving labels that are problematic. An ML model may be trained to predict labels for any given data item. The ML model may be validated to determine a confusion metric with respect to each distinct pair of labels predicted by the ML model. Each confusion metric indicates how a particular label is being mistaken for another particular label. The confusion metrics are analyzed to determine whether any of the ML model-generated labels are problematic (e.g., a label conflicts with another label, a label that is rarely predicted, a label that is incorrectly predicted, etc.). Steps for resolving the problematic labels are implemented, and the ML model is retrained based on the resolution steps. By doing so, the ML model generates a more accurate label for a data item.