Machine Learning Training Data Augmentation via Auxiliary Source Merger
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning algorithms trained with inaccurate or incomplete data can produce flawed outputs, which can have negative consequences in critical decision-making processes, as existing data sources often lack reliability and consistency, with large databases containing inaccuracies and smaller, curated databases being insufficient for comprehensive information retrieval.
Innovation Solution
The method involves selecting and merging key elements from different data sources by identifying a target performance metric, obtaining auxiliary data from separate sources, selecting candidate attribute types, calculating their impact on the performance metric, and determining a tradeoff between quality and performance to augment the training data set, thereby improving the fidelity of the data used to train machine learning algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data from large databases is used to train machine learning algorithms, then the quantity of training data is increased, but the accuracy and reliability of the training data deteriorates due to inaccuracies in large databases
Solution Approach 1:
The patent merges data from multiple auxiliary data sources with the original training data set. By combining data from multiple sources, the system achieves both increased data quantity and improved reliability through the diversity and complementary nature of different sources, resolving the contradiction between quantity and accuracy.
Solution Approach 2:
The patent applies local quality by selectively integrating auxiliary data based on specific quality metrics and relevance to the machine learning task. Not all data is treated uniformly - instead, data from auxiliary sources is evaluated and integrated based on its specific quality characteristics and suitability for particular attributes, ensuring high reliability while expanding quantity.
2Reliability
If auxiliary data from multiple sources is integrated to improve training data quality, then the accuracy of training data is improved, but the complexity of data processing and integration increases
Solution Approach 1:
The patent introduces an intermediary processing layer that manages the integration of auxiliary data from multiple sources. This intermediary system handles the complexity of data matching, quality assessment, and integration automatically, reducing the apparent complexity for end users while maintaining high data quality through systematic processing.
Solution Approach 2:
The patent implements feedback mechanisms where the system evaluates the impact of integrated auxiliary data on machine learning model performance. This feedback loop allows the system to automatically adjust and optimize the integration process, managing complexity through iterative improvement and adaptive control rather than requiring complex upfront design.
3Adaptability or versatility
If additional attribute types from auxiliary data sets are included in training data, then the comprehensiveness of training data is improved, but the quality control and validation of the data becomes more difficult
Solution Approach 1:
The patent replaces manual quality control and validation processes with automated computational methods. Machine learning algorithms and processing systems automatically assess the quality and relevance of additional attribute types from auxiliary data sources, making quality control scalable and manageable even as comprehensiveness increases.
Solution Approach 2:
The patent enables the data integration system to self-validate and self-regulate the quality of integrated data. Through automated quality metrics and validation rules, the system independently assesses and controls the quality of additional attribute types without requiring extensive external intervention, reducing the difficulty of quality control as comprehensiveness grows.
Data Source
AI summary
In one example, a method includes identifying a target performance metric of a machine learning algorithm, wherein the target performance metric is to be improved, obtaining a set of auxiliary data from a plurality of auxiliary data sources, wherein the plurality of auxiliary data sources is separate from a training data set used to train the machine learning algorithm, selecting a candidate attribute type from the set of auxiliary data, identifying a quality metric for the candidate attribute type, calculating a change in the target performance metric when data values associated with the candidate attribute type are included in the training data set, determining that a tradeoff between the target performance metric and the quality metric of the candidate attribute type is satisfied by inclusion of the data values in the training data set, and training the machine learning algorithm using the training data set augmented with the data value.


