Incomplete Training Data Loss Class Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine-learning-based systems often face accuracy and reliability issues due to incomplete and erroneous training data, which can lead to suboptimal performance in real-world applications.
Innovation Solution
The method involves mapping incomplete training data into loss classes and employing specific loss-generation methods for each class during training, allowing for effective and efficient use of such data by classifying and adjusting losses based on additional information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If incomplete training data is used to train machine-learning-based systems, then training efficiency is improved, but accuracy and reliability deteriorate
Solution Approach 1:
The patent segments incomplete training data into different loss classes based on the type and degree of incompleteness. Each loss class is then processed using specialized loss functions tailored to its characteristics, allowing the system to efficiently train on incomplete data while maintaining accuracy by addressing each type of incompleteness appropriately.
Solution Approach 2:
The patent modifies the loss function parameters and training process based on the specific characteristics of incomplete data. By adjusting loss weights, sampling strategies, and training parameters according to the data's incompleteness profile, the system optimizes both training efficiency and model reliability for incomplete datasets.
2Device complexity
If traditional training methods are used on incomplete data, then processing simplicity is maintained, but model accuracy deteriorates
Solution Approach 1:
The patent introduces loss classes as an intermediary layer between the raw incomplete training data and the model training process. This intermediary structure organizes incomplete data by its deficiency characteristics and applies appropriate handling strategies, improving accuracy without significantly complicating the overall training workflow.
Solution Approach 2:
The patent implements dynamic training strategies that adapt to the specific characteristics of incomplete data. The system dynamically adjusts loss weights, sampling probabilities, and training parameters based on the identified loss class, allowing flexible handling of various incomplete data scenarios while maintaining processing efficiency.
Data Source
AI summary
The current document is directed to methods and systems that effectively and efficiently employ incomplete training data to train machine-learning-based systems. Incomplete training data, as one example, may include training data with erroneous or inaccurate input-vector/label pairs. In currently disclosed methods and systems, Incomplete training data is mapped to loss classes based on addition training-data information and specific, different additional-information-dependent loss-generation methods are employed for training data of different loss classes during machine-learning-based-system training so that incomplete training data can be effectively and efficiently used.


