Learning Model Intermediate Feature Quantities for Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning models face performance issues when training data is insufficient, leading to potential failure in output results, especially in business scenarios where customer data privacy is a concern, and conventional methods of data augmentation result in increased data volume without ensuring learning accuracy.

Innovation Solution

A learning method that first learns to make intermediate feature quantities similar to reference feature quantities generated from augmented training data, using these reference feature quantities in subsequent learning processes to enhance accuracy while reducing data volume.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data augmentation is performed to expand the range of training data, then the coverage of applied data types is improved, but the volume of training data increases

Engineering Contradiction:
Improvecoverage of applied data typesVSAvoidvolume of training data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts intermediate feature quantities from the data augmentation process and stores them separately in a database. This allows the system to retain the beneficial feature representations without storing the complete augmented training data, thus reducing data volume while maintaining adaptability to various data types

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces intermediate feature quantities as an intermediary representation between raw training data and model output. These feature quantities serve as a compressed summary that captures essential information from augmented data without requiring storage of the full augmented datasets, resolving the contradiction between coverage and volume

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If intermediate feature quantities from multiple training datasets are stored to increase available data volume, then the learning accuracy is improved, but the data storage requirement increases

Engineering Contradiction:
Improvelearning accuracyVSAvoiddata storage requirement
Core Design Contradiction:
ReliabilityVSVolume of stationary object

Solution Approach 1:

The patent extracts only the intermediate feature quantities from the complete training datasets and stores these compressed representations in a database. This extraction approach maintains the essential learning information needed for accuracy while significantly reducing the storage requirement compared to storing complete datasets

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the storage format from raw training data to intermediate feature quantities, changing the parameter representation. This transformation preserves the informational content necessary for learning accuracy while reducing the data volume and storage requirements

Inventive Principle:
Principle #35Parameter changes

3Reliability

If customer data is retained to enable continued use in learning models, then the learning accuracy is improved, but the data privacy risk increases

Engineering Contradiction:
Improvelearning accuracyVSAvoiddata privacy risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and stores only the intermediate feature quantities derived from customer data, not the original customer data itself. This extraction allows the system to maintain learning accuracy using processed feature representations while minimizing privacy risks by not retaining sensitive原始 data

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a copy of the essential information (intermediate feature quantities) from the original customer data for continued learning use. This copy approach enables the system to maintain learning models without retaining the original sensitive customer data, thus reducing privacy exposure

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11409988B2Method, recording medium, and device for utilizing feature quantities of augmented training data
Publication Date: 2022.08.09 FUJITSU LTD
  • US11409988B2 patent drawing
  • US11409988B2 patent drawing
  • US11409988B2 patent drawing

AI summary

A learning device learns at last one parameter of a learning model such that each intermediate feature quantity becomes similar to a reference feature quantity, the each intermediate feature quantity being calculated as a result of inputting a plurality of sets of augmentation training data to a first neural network in the learning model, the plurality of augmentation training data being generated by performing data augmentation based on same first original training data. The learning device learns at last one parameter of a second network, in the learning model, using second original training data, which is different than the first original training data, and using the reference feature quantity.