Virtual Training Sets for Neural Network Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The development of neural networks is hindered by the time-consuming process of creating training sets, with small training sets leading to low accuracy and the lack of explainability due to the 'black box' nature of neural networks, especially in extrapolation modes, and the proprietary and expensive nature of commercial training data.
Innovation Solution
A system and method that generates improved training sets for convolutional neural networks using a General Logic Gate Module (GLGM) to produce virtual training sets, which are then combined with real-world data to create hybrid training sets, enhancing the network's range and traceability, and incorporating sensitivity and importance parameters for better feature identification and backpropagation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-world training data is used, then training accuracy is improved, but data acquisition cost and time increase
Solution Approach 1:
The patent creates virtual training sets by synthesizing data that copies the essential characteristics and patterns of real-world data. The system generates artificial training examples that mimic the statistical properties, feature distributions, and relationships found in actual operational data, allowing the neural network to be trained without requiring extensive collection of real-world samples.
Solution Approach 2:
The patent extracts key features, patterns, and statistical properties from available real-world data or domain knowledge, then uses these extracted elements to generate comprehensive virtual training sets. By separating the essential training requirements from the need for extensive real data collection, the system achieves accurate training with reduced data acquisition overhead.
2Measurement precision
If neural network complexity is increased, then prediction accuracy is improved, but explainability decreases
Solution Approach 1:
The patent introduces virtual training sets as an intermediary that bridges the gap between complex neural network models and explainability requirements. By training on synthetically generated data with known ground truths and controlled variations, the system can achieve high accuracy while maintaining the ability to trace and explain predictions through the training data generation process and feature importance analysis.
3Reliability
If training set size is increased, then neural network performance is improved, but development time increases
Solution Approach 1:
The patent performs preliminary actions by pre-generating comprehensive virtual training sets that encompass a wide range of possible scenarios, edge cases, and operational conditions. This advance preparation creates a robust training foundation that improves network performance while reducing the iterative development cycle, as the training data is systematically generated rather than collected and prepared over time.
Solution Approach 2:
The patent utilizes parameter changes in the virtual data generation process to efficiently create diverse training samples. By varying parameters such as feature distributions, noise levels, and scenario conditions during synthetic data generation, the system produces large volumes of high-quality training data rapidly, improving network reliability without proportionally increasing development time.
Data Source
AI summary
One embodiment of the present invention provides a computer implemented method for generating a training set to train a convolutional neural network comprising the steps of providing prediction space data to a General Logic Gate Module (GLGM). Prediction space expert judgement is also provided to the GLGM and to a sensitivity and importance module. The GLGM determines or outputs state possibilities. The state possibilities are provided to the sensitivity and importance module and to the feature extraction module. Feature extraction algorithms are applied to the state possibilities within the feature extraction module to produce a training possibility set that is a virtual training possibility set. The training possibility set is provided to a state inferential module and to a final training set. From the state inferential module a possibility ranking is generated that is independent of the convolutional neural network and further the output from the state inferential module is provided to a sensitivity and importance module for analysis. A sensitivity parameter and an importance parameter is determined from the output from the sensitivity and importance module. The state possibility ranking is provided to the final training set. The sensitivity parameter and importance parameter are provided to a final training set and a training set structure metric. A convolutional neural network input layer is generated from the final training set informed by one or more of the state possibility ranking, the sensitivity parameter, the importance parameter and the training possibility set. A convolutional neural network layer design is generated from the training set structure metric.


