Deep Learning Model Overfitting Reduction via Supplemental Benign Training Examples
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep learning systems for pathogenicity classification often overfit on position frequency matrices, leading to difficulty in distinguishing benign from deleterious amino acid missense variants, which diminishes their ability to generalize and predict pathogenicity accurately.
Innovation Solution
The introduction of supplemental benign training examples that share the same position frequency matrices as missense training examples, but are labeled as either benign or pathogenic, forces the system to differentiate between them based on other features, reducing overfitting and improving training outcomes by contrasting against pathogenic or unlabeled examples.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the deep learning system is trained using position frequency matrices for pathogenicity classification, then the model can process and classify variants, but the system overfits on the position frequency matrices, leading to difficulty in distinguishing benign from deleterious amino acid missense variants
Solution Approach 1:
The patent applies preliminary action by introducing supplemental benign training examples before the main training process. These supplemental examples are constructed to share the same position frequency matrices as missense training examples but are labeled as benign or pathogenic, forcing the system to differentiate based on other features. This preliminary differentiation step reduces overfitting on the position frequency matrices and improves generalization ability before the model is used for actual pathogenicity classification.
2Measurement precision
If the model is trained to accurately classify pathogenicity using position frequency matrices, then classification precision improves, but the model's ability to generalize to new variants diminishes
Solution Approach 1:
The patent introduces supplemental benign training examples as a preliminary step before main training. These examples force the model to learn differentiation beyond position frequency matrices, thereby improving adaptability to new variants while maintaining classification precision.
Solution Approach 2:
The patent changes the training parameters by introducing supplemental benign training examples with specific labeling (benign or pathogenic) that differ from standard training approaches. This parameter change forces the model to utilize additional features beyond position frequency matrices, thereby improving generalization to new variants while maintaining classification accuracy.
3Reliability
If supplemental benign training examples are introduced to reduce overfitting, then model generalization improves, but training data requirements and complexity increase
Solution Approach 1:
The supplemental benign training examples are introduced as a preliminary step with a clear, structured approach. The examples are systematically constructed to share position frequency matrices with missense training examples, which organizes the increased complexity in a manageable way while achieving improved generalization.
Solution Approach 2:
The patent systematically changes training parameters by introducing supplemental benign training examples with specific labeling schemes. This structured parameter change improves generalization ability while keeping the increased complexity organized and manageable through consistent labeling and data construction protocols.
Data Source
AI summary
The technology disclosed includes systems and methods to reduce overfitting of neural network-implemented models that process sequences of amino acids and accompanying position frequency matrices. The system generates supplemental training example sequence pairs, labelled benign, that include a start location, through a target amino acid location, to an end location. A supplemental sequence pair supplements a pathogenic or benign missense training example sequence pair. It has identical amino acids in a reference and an alternate sequence of amino acids. The system includes logic to input with each supplemental sequence pair a supplemental training position frequency matrix (PFM) that is identical to the PFM of the benign or pathogenic missense at the matching start and end location. The system includes logic to attenuate the training influence of the training PFMs during training the neural network-implemented model by including supplemental training example PFMs in the training data.


