Pseudo-Labeling for Multi-Task Neural Network Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Generating training data sets with matched ground-truth data for multi-task learning networks is time-consuming and costly, and requires additional labeling tasks when new tasks are added.
Innovation Solution
A training data generation apparatus comprising a first network, a second network, and a target network that perform supervised and unsupervised learning, ensemble learning, and fusion to generate pseudo label data without requiring additional labeling tasks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If training data sets with ground-truth data are collected for multi-task learning networks, then the network can be trained effectively, but it requires a lot of time and money
Solution Approach 1:
The patent uses pseudo-labeling to create synthetic ground-truth data by copying and adapting labels from similar samples or previous iterations. The system generates pseudo-labels for unlabeled data, effectively creating training data without manual annotation, thus reducing time and cost while maintaining training effectiveness
Solution Approach 2:
The patent performs preliminary labeling on a subset of data to create initial training sets, then uses these to generate pseudo-labels for the remaining data. This preliminary action enables subsequent automated labeling processes to proceed efficiently without requiring complete manual annotation upfront
2Measurement precision
If ground-truth data is matched for each task in multi-task learning, then the network learns accurately, but additional labeling tasks are required when new tasks are added
Solution Approach 1:
The patent creates a universal pseudo-labeling framework that can handle multiple tasks simultaneously. The same pseudo-labeling mechanism works across different task types (classification, detection, segmentation), eliminating the need for task-specific labeling procedures and reducing overall system complexity
Solution Approach 2:
The system performs self-labeling by generating its own training data through pseudo-labeling. When new tasks are added, the system automatically generates pseudo-labels for those tasks using its existing capabilities, without requiring external labeling resources or complex additional labeling workflows
3Loss of information
If manual labeling is performed for training data, then ground-truth data is obtained, but the process is time-consuming and costly
Solution Approach 1:
The patent replaces manual copying of ground-truth labels with automated pseudo-label generation. The system copies relevant features and patterns from labeled samples to generate pseudo-labels for unlabeled samples, maintaining data quality while dramatically improving productivity
Solution Approach 2:
The patent introduces pseudo-labels as an intermediary between manual labels and final training data. These pseudo-labels serve as intermediate annotations that bridge the gap between limited manual labeling and complete training data sets, enabling efficient semi-supervised learning
Data Source
AI summary
A training data generation apparatus may include a first network and a second network that individually learn a first image based on supervised learning and perform ensemble learning on a second image based on unsupervised learning. The apparatus also may include a fusion network that obtains a fusion output value based on the ensemble learning results of the first network and the second network. The apparatus also may include a target network that learns the second image to imitate the fusion output value.


