Selective Data Labeling for Multi-Task Learning with Active Forgetting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data labeling techniques for multi-task learning are inefficient, leading to inaccurate labeling, increased costs, and degraded performance due to the accumulation of unuseful data, and existing methods fail to effectively identify and discard less useful data, resulting in suboptimal dataset construction.
Innovation Solution
A device and method utilizing active learning and active forgetting to selectively label data, identifying useful data through active learning and removing unuseful data through active forgetting, optimizing dataset composition and reducing labeling costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is continuously accumulated and expanded to construct a dataset, then the quantity of training data increases, but the dataset contains unuseful data that degrades model performance
Solution Approach 1:
The patent implements active forgetting to discard unuseful data from the training dataset. The system identifies and removes data samples that do not contribute to model performance improvement, thereby maintaining a high-quality training set that prevents performance degradation while keeping the dataset size manageable.
Solution Approach 2:
The patent employs a feedback mechanism where model performance is continuously evaluated after training on expanded datasets. This feedback loop identifies when adding more data stops improving or degrades performance, triggering the active forgetting process to remove unuseful data and optimize the dataset composition.
2Adaptability or versatility
If all tasks are labeled in multi-task learning, then comprehensive task coverage is achieved, but labeling costs substantially increase
Solution Approach 1:
The patent applies local quality by selectively labeling data based on task-specific usefulness. Instead of uniformly labeling all data for all tasks, the system identifies and labels only those data samples that are useful for specific tasks, thereby reducing overall labeling costs while maintaining comprehensive task coverage through targeted labeling.
Solution Approach 2:
The patent segments the multi-task learning problem into individual task evaluations. The system separately assesses the usefulness of data samples for each task and selectively labels them accordingly, rather than treating all tasks uniformly. This segmentation approach optimizes labeling resource allocation across multiple tasks.
3Measurement precision
If pixel-level or sophisticated labeling is performed for multi-task learning, then labeling accuracy improves, but construction costs substantially increase
Solution Approach 1:
The patent applies partial action by performing sophisticated labeling only on data samples that are identified as useful for specific tasks. Instead of applying high-cost pixel-level or sophisticated labeling to all data, the system selectively applies these labeling methods only where necessary, thereby maintaining labeling accuracy for critical samples while substantially reducing overall construction costs.
4Quantity of substance
If unuseful data is accumulated within a limited budget, then data quantity increases, but budget allocation becomes inefficient
Solution Approach 1:
The patent implements a feedback mechanism that continuously evaluates the usefulness of accumulated data samples. This feedback system identifies unuseful data before budget exhaustion occurs, enabling real-time optimization of budget allocation by redirecting resources from unuseful data collection to useful data acquisition, thereby improving overall budget allocation efficiency.
Solution Approach 2:
The patent replaces manual or mechanical data accumulation processes with an intelligent system that automatically evaluates data usefulness. This substitution uses machine learning models and automated evaluation metrics to identify and prioritize useful data, replacing inefficient manual data collection approaches and significantly improving budget allocation efficiency.
Data Source
AI summary
A device for constructing a dataset for multiple tasks through selective labeling using active learning and active forgetting includes a data collector configured to collect data for multi-task learning and a classifier that includes a deep learning model for performing the multi-task and is configured to classify useful data useful for performing a specific task through active learning using the deep learning model among the collected data, and unuseful data not useful for performing the specific task through active forgetting using the deep learning model among the identified useful data.


