Feature Space Training for Labeled Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating labeled data from unlabeled data often result in discrepancies, leading to poor-quality labeled data, which is costly and time-consuming to collect, and limits the accuracy and efficiency of machine learning models across different data domains.

Innovation Solution

An information processing device trains a feature space where data within the same domain is closer and data from different domains is farther apart, allowing for the generation of high-quality labeled data sets by integrating labeled data within a predetermined range in this feature space, using unlabeled data to reduce collection costs and improve analysis accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If labeled data is collected through manual annotation, then data quality is high, but collection cost and time increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidcollection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent uses unlabeled data as copies or substitutes for labeled data in the feature space. By training a model to map unlabeled data into the same feature space where labeled data resides, the system can generate pseudo-labeled data that mimics the characteristics of manually annotated data without requiring actual manual annotation, thus reducing time and cost while maintaining data quality

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the parameter state of data by mapping unlabeled data into the feature space through learned transformations. The model changes the representation parameters of unlabeled data to match the feature space characteristics, enabling automated label generation that preserves quality attributes while eliminating manual annotation requirements

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual annotation is used for data labeling, then labeled data accuracy is high, but production cost increases

Engineering Contradiction:
Improvelabeled data accuracyVSAvoidcollection cost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system creates copies of labeled data patterns by projecting unlabeled data into the feature space. This allows the generation of accurate labeled data through automated mapping rather than expensive manual annotation, maintaining high accuracy while significantly reducing production costs

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system enables self-service labeling by allowing unlabeled data to automatically find its position in the feature space and generate its own labels through the trained model, eliminating the need for human annotators and associated costs while maintaining accuracy

Inventive Principle:
Principle #25Self-service

3Measurement precision

If feature space is not trained, then data processing is fast, but cross-domain analysis accuracy is poor

Engineering Contradiction:
Improvecross-domain analysis accuracyVSAvoiddata processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by training the feature space mapping model in advance before actual data processing. This pre-training establishes the geometric relationships and domain boundaries in the feature space, enabling both high cross-domain analysis accuracy and efficient processing during deployment without requiring real-time complex computations

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230259827A1Computer-readable recording medium storing generation program, generation method, and information processing device
Publication Date: 2023.08.17 FUJITSU LTD
  • US20230259827A1 patent drawing
  • US20230259827A1 patent drawing
  • US20230259827A1 patent drawing

AI summary

A non-transitory computer-readable recording medium stores a generation program for causing a computer to execute a process including: with data included in each of a plurality of data sets, training a feature space in which a distance between pieces of the data included in a same domain is shorter and the distance of the data between different domains is longer; and generating labeled data sets by integrating labeled data included within a predetermined range in the trained feature space, among a plurality of pieces of the labeled data.