Feature Space Training for Labeled Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for generating labeled data from unlabeled data often result in discrepancies, leading to poor-quality labeled data, which is costly and time-consuming to collect, and limits the accuracy and efficiency of machine learning models across different data domains.
Innovation Solution
An information processing device trains a feature space where data within the same domain is closer and data from different domains is farther apart, allowing for the generation of high-quality labeled data sets by integrating labeled data within a predetermined range in this feature space, using unlabeled data to reduce collection costs and improve analysis accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If labeled data is collected through manual annotation, then data quality is high, but collection cost and time increase significantly
Solution Approach 1:
The patent uses unlabeled data as copies or substitutes for labeled data in the feature space. By training a model to map unlabeled data into the same feature space where labeled data resides, the system can generate pseudo-labeled data that mimics the characteristics of manually annotated data without requiring actual manual annotation, thus reducing time and cost while maintaining data quality
Solution Approach 2:
The patent transforms the parameter state of data by mapping unlabeled data into the feature space through learned transformations. The model changes the representation parameters of unlabeled data to match the feature space characteristics, enabling automated label generation that preserves quality attributes while eliminating manual annotation requirements
2Measurement precision
If manual annotation is used for data labeling, then labeled data accuracy is high, but production cost increases
Solution Approach 1:
The system creates copies of labeled data patterns by projecting unlabeled data into the feature space. This allows the generation of accurate labeled data through automated mapping rather than expensive manual annotation, maintaining high accuracy while significantly reducing production costs
Solution Approach 2:
The system enables self-service labeling by allowing unlabeled data to automatically find its position in the feature space and generate its own labels through the trained model, eliminating the need for human annotators and associated costs while maintaining accuracy
3Measurement precision
If feature space is not trained, then data processing is fast, but cross-domain analysis accuracy is poor
Solution Approach 1:
The patent applies preliminary action by training the feature space mapping model in advance before actual data processing. This pre-training establishes the geometric relationships and domain boundaries in the feature space, enabling both high cross-domain analysis accuracy and efficient processing during deployment without requiring real-time complex computations
Data Source
AI summary
A non-transitory computer-readable recording medium stores a generation program for causing a computer to execute a process including: with data included in each of a plurality of data sets, training a feature space in which a distance between pieces of the data included in a same domain is shorter and the distance of the data between different domains is longer; and generating labeled data sets by integrating labeled data included within a predetermined range in the trained feature space, among a plurality of pieces of the labeled data.


