Generative Cooperative Network for Neural Network Training Data Synthesis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for generating training datasets for neural networks require significant resources and manual effort, making it difficult to provide large datasets for training, especially for applications where such resources are not feasible.
Innovation Solution
A training dataset generator model is used to generate datasets using Gaussian noise as input, which is then evaluated and adjusted based on the performance of the trained model, allowing for the iterative improvement of dataset generation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If large training datasets are built manually through crowdsourcing, then the quality and size of training data is improved, but the computational resources, storage resources, and manual effort required increase significantly
Solution Approach 1:
The patent uses a trained generator model to copy and synthesize training data samples automatically, replacing the need for manual crowdsourcing. The generator model learns from a small reference dataset and generates synthetic samples that replicate the statistical properties and patterns of real data, thereby producing large-scale training datasets without proportional increases in manual effort or resource requirements
Solution Approach 2:
The system enables self-service data generation where the trained generator model autonomously produces training datasets without requiring ongoing manual intervention. Once the generator is trained on a reference dataset, it can independently generate unlimited synthetic training samples, eliminating the need for continuous crowdsourcing and reducing both manual effort and operational costs
2Quantity of substance
If large training datasets are built manually through crowdsourcing, then the quality and size of training data is improved, but the manual effort required increases significantly
Solution Approach 1:
The patent replaces the mechanical process of manual data labeling and crowdsourcing with an automated machine learning system. The generator model, once trained, automatically generates synthetic training data through computational processes, substituting human manual effort with algorithmic automation that can produce datasets at scale without proportional increases in manual labor
Solution Approach 2:
The system enables self-service data generation where the trained generator model autonomously produces training datasets without requiring ongoing manual intervention. Once the generator is trained on a reference dataset, it can independently generate unlimited synthetic training samples, eliminating the need for continuous crowdsourcing and reducing both manual effort and operational costs
3Adaptability or versatility
If conventional techniques are used to build training datasets, then the training data can be provided for standard applications, but the resources and effort required become prohibitive for several applications
Solution Approach 1:
The patent creates a universal generator model that can adapt to different application domains and data types. By training the generator on domain-specific reference datasets, the same automated system can generate training data for multiple applications including but not limited to image recognition, natural language processing, and speech recognition, providing versatile support across different AI tasks without requiring separate manual crowdsourcing processes for each application
Solution Approach 2:
The system adapts to different applications by changing the parameters and characteristics of the reference training dataset used for generator training. By adjusting the input data distribution, domain knowledge, and training parameters, the generator can produce synthetic data tailored to specific application requirements while maintaining the same automated generation infrastructure, thereby achieving application flexibility without proportional resource increases
Data Source
AI summary
A generative cooperative network (GCN) comprises a dataset generator model and a learner model. The dataset generator model generates training datasets used to train the learner model. The trained learner model is evaluated according to a reference training dataset. The dataset generator model is modified according to the evaluation. The training datasets, the dataset generator model, and the leaner model are stored by the GCN. The trained learner model is configured to receive input and to generate output based on the input.


