Machine Learning Training Data Augmentation via Auxiliary Source Merger

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning algorithms trained with inaccurate or incomplete data can produce flawed outputs, which can have negative consequences in critical decision-making processes, as existing data sources often lack reliability and consistency, with large databases containing inaccuracies and smaller, curated databases being insufficient for comprehensive information retrieval.

Innovation Solution

The method involves selecting and merging key elements from different data sources by identifying a target performance metric, obtaining auxiliary data from separate sources, selecting candidate attribute types, calculating their impact on the performance metric, and determining a tradeoff between quality and performance to augment the training data set, thereby improving the fidelity of the data used to train machine learning algorithms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data from large databases is used to train machine learning algorithms, then the quantity of training data is increased, but the accuracy and reliability of the training data deteriorates due to inaccuracies in large databases

Engineering Contradiction:
Improvequantity of training dataVSAvoidaccuracy and reliability of training data
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent merges data from multiple auxiliary data sources with the original training data set. By combining data from multiple sources, the system achieves both increased data quantity and improved reliability through the diversity and complementary nature of different sources, resolving the contradiction between quantity and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies local quality by selectively integrating auxiliary data based on specific quality metrics and relevance to the machine learning task. Not all data is treated uniformly - instead, data from auxiliary sources is evaluated and integrated based on its specific quality characteristics and suitability for particular attributes, ensuring high reliability while expanding quantity.

Inventive Principle:
Principle #3Local quality

2Reliability

If auxiliary data from multiple sources is integrated to improve training data quality, then the accuracy of training data is improved, but the complexity of data processing and integration increases

Engineering Contradiction:
Improvequality of training dataVSAvoidcomplexity of data processing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing layer that manages the integration of auxiliary data from multiple sources. This intermediary system handles the complexity of data matching, quality assessment, and integration automatically, reducing the apparent complexity for end users while maintaining high data quality through systematic processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback mechanisms where the system evaluates the impact of integrated auxiliary data on machine learning model performance. This feedback loop allows the system to automatically adjust and optimize the integration process, managing complexity through iterative improvement and adaptive control rather than requiring complex upfront design.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If additional attribute types from auxiliary data sets are included in training data, then the comprehensiveness of training data is improved, but the quality control and validation of the data becomes more difficult

Engineering Contradiction:
Improvecomprehensiveness of training dataVSAvoiddifficulty of quality control and validation
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent replaces manual quality control and validation processes with automated computational methods. Machine learning algorithms and processing systems automatically assess the quality and relevance of additional attribute types from auxiliary data sources, making quality control scalable and manageable even as comprehensiveness increases.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent enables the data integration system to self-validate and self-regulate the quality of integrated data. Through automated quality metrics and validation rules, the system independently assesses and controls the quality of additional attribute types without requiring extensive external intervention, reducing the difficulty of quality control as comprehensiveness grows.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230057792A1Training data fidelity for machine learning applications through intelligent merger of curated auxiliary data
Publication Date: 2023.02.23 AT&T INTELLECTUAL PROPERTY I L P
  • US20230057792A1 patent drawing
  • US20230057792A1 patent drawing
  • US20230057792A1 patent drawing

AI summary

In one example, a method includes identifying a target performance metric of a machine learning algorithm, wherein the target performance metric is to be improved, obtaining a set of auxiliary data from a plurality of auxiliary data sources, wherein the plurality of auxiliary data sources is separate from a training data set used to train the machine learning algorithm, selecting a candidate attribute type from the set of auxiliary data, identifying a quality metric for the candidate attribute type, calculating a change in the target performance metric when data values associated with the candidate attribute type are included in the training data set, determining that a tradeoff between the target performance metric and the quality metric of the candidate attribute type is satisfied by inclusion of the data values in the training data set, and training the machine learning algorithm using the training data set augmented with the data value.