Learning Model Intermediate Feature Quantities for Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning models face performance issues when training data is insufficient, leading to potential failure in output results, especially in business scenarios where customer data privacy is a concern, and conventional methods of data augmentation result in increased data volume without ensuring learning accuracy.
Innovation Solution
A learning method that first learns to make intermediate feature quantities similar to reference feature quantities generated from augmented training data, using these reference feature quantities in subsequent learning processes to enhance accuracy while reducing data volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If data augmentation is performed to expand the range of training data, then the coverage of applied data types is improved, but the volume of training data increases
Solution Approach 1:
The patent extracts intermediate feature quantities from the data augmentation process and stores them separately in a database. This allows the system to retain the beneficial feature representations without storing the complete augmented training data, thus reducing data volume while maintaining adaptability to various data types
Solution Approach 2:
The patent introduces intermediate feature quantities as an intermediary representation between raw training data and model output. These feature quantities serve as a compressed summary that captures essential information from augmented data without requiring storage of the full augmented datasets, resolving the contradiction between coverage and volume
2Reliability
If intermediate feature quantities from multiple training datasets are stored to increase available data volume, then the learning accuracy is improved, but the data storage requirement increases
Solution Approach 1:
The patent extracts only the intermediate feature quantities from the complete training datasets and stores these compressed representations in a database. This extraction approach maintains the essential learning information needed for accuracy while significantly reducing the storage requirement compared to storing complete datasets
Solution Approach 2:
The patent transforms the storage format from raw training data to intermediate feature quantities, changing the parameter representation. This transformation preserves the informational content necessary for learning accuracy while reducing the data volume and storage requirements
3Reliability
If customer data is retained to enable continued use in learning models, then the learning accuracy is improved, but the data privacy risk increases
Solution Approach 1:
The patent extracts and stores only the intermediate feature quantities derived from customer data, not the original customer data itself. This extraction allows the system to maintain learning accuracy using processed feature representations while minimizing privacy risks by not retaining sensitive原始 data
Solution Approach 2:
The patent creates a copy of the essential information (intermediate feature quantities) from the original customer data for continued learning use. This copy approach enables the system to maintain learning models without retaining the original sensitive customer data, thus reducing privacy exposure
Data Source
AI summary
A learning device learns at last one parameter of a learning model such that each intermediate feature quantity becomes similar to a reference feature quantity, the each intermediate feature quantity being calculated as a result of inputting a plurality of sets of augmentation training data to a first neural network in the learning model, the plurality of augmentation training data being generated by performing data augmentation based on same first original training data. The learning device learns at last one parameter of a second network, in the learning model, using second original training data, which is different than the first original training data, and using the reference feature quantity.


