Learning Dataset Generation Device for Supervised Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for generating new learning datasets, such as data augmentation and deep convolutional generative adversarial networks, are inefficient in creating sufficient variation when the existing dataset is limited, especially in supervised learning, where high labor and time costs are associated with preparing learning datasets.

Innovation Solution

A method involving sorting a learning dataset into subsets based on output signals, generating new input signals through a learning device that alternates between input and output signal groups, and integrating these new signals to create a new learning dataset, thereby expanding the dataset with increased variation at reduced costs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data augmentation is used to expand the learning dataset, then the quantity of learning data is increased, but the variation of the object itself does not change sufficiently

Engineering Contradiction:
Improvequantity of learning dataVSAvoidvariation of object
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent creates new learning data by generating copies of existing learning data through a learning device. The device learns from the original learning data and generates new input signals that are copies with variations, thereby increasing both quantity and variation of the learning dataset without requiring manual creation of new data.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If deep convolutional generative adversarial network is used to generate new still images, then new learning data with variation can be generated, but it takes time to converge and requires sufficiently high quality and quantity of existing learning data

Engineering Contradiction:
Improvevariation of learning dataVSAvoidconvergence time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments the learning data into multiple subsets and processes each subset independently through the learning device. This segmentation allows the system to generate varied learning data more efficiently by parallel processing, reducing the overall convergence time compared to processing the entire dataset as a single unit.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If manual linking of input signals and output signals is performed, then the quality of learning dataset is maintained, but time and labor costs are high

Engineering Contradiction:
Improvequality of learning datasetVSAvoidpreparation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The learning device performs self-service by automatically generating new learning data with proper input-output signal pairing without requiring manual intervention. The device learns the relationship between input and output signals from the original learning data and autonomously creates new paired data, maintaining quality while eliminating manual labor and time costs.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11551080B2Learning dataset generation method, new learning dataset generation device and learning method using generated learning dataset
Publication Date: 2023.01.10 KOKUSAI DENKI ELECTRIC INC
  • US11551080B2 patent drawing
  • US11551080B2 patent drawing
  • US11551080B2 patent drawing

AI summary

Even if an existing learning dataset is limited, a new learning dataset with sufficient variation is generated. Therefore, for each of a plurality of learning data subsets, new input signals are generated from input signals of a plurality of pieces of learning data, and a plurality of pieces of new learning data that are respectively combinations of the new input signals and output signals of the corresponding learning data subset are generated. The input signals of the plurality of pieces of the learning data included in the corresponding learning data subset are divided into a first signal group and a second signal group, and the new input signals are generated by a learning device that is generated by performing learning by the first signal group set as an input signal set and the second signal group set as an output signal set.