Deep Learning Model Overfitting Reduction via Supplemental Benign Training Examples

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning systems for pathogenicity classification often overfit on position frequency matrices, leading to difficulty in distinguishing benign from deleterious amino acid missense variants, which diminishes their ability to generalize and predict pathogenicity accurately.

Innovation Solution

The introduction of supplemental benign training examples that share the same position frequency matrices as missense training examples, but are labeled as either benign or pathogenic, forces the system to differentiate between them based on other features, reducing overfitting and improving training outcomes by contrasting against pathogenic or unlabeled examples.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the deep learning system is trained using position frequency matrices for pathogenicity classification, then the model can process and classify variants, but the system overfits on the position frequency matrices, leading to difficulty in distinguishing benign from deleterious amino acid missense variants

Engineering Contradiction:
Improvepathogenicity classification accuracyVSAvoidmodel generalization ability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by introducing supplemental benign training examples before the main training process. These supplemental examples are constructed to share the same position frequency matrices as missense training examples but are labeled as benign or pathogenic, forcing the system to differentiate based on other features. This preliminary differentiation step reduces overfitting on the position frequency matrices and improves generalization ability before the model is used for actual pathogenicity classification.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If the model is trained to accurately classify pathogenicity using position frequency matrices, then classification precision improves, but the model's ability to generalize to new variants diminishes

Engineering Contradiction:
Improvevariant classification accuracyVSAvoidmodel generalization to new variants
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces supplemental benign training examples as a preliminary step before main training. These examples force the model to learn differentiation beyond position frequency matrices, thereby improving adaptability to new variants while maintaining classification precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the training parameters by introducing supplemental benign training examples with specific labeling (benign or pathogenic) that differ from standard training approaches. This parameter change forces the model to utilize additional features beyond position frequency matrices, thereby improving generalization to new variants while maintaining classification accuracy.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If supplemental benign training examples are introduced to reduce overfitting, then model generalization improves, but training data requirements and complexity increase

Engineering Contradiction:
Improvemodel generalization abilityVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The supplemental benign training examples are introduced as a preliminary step with a clear, structured approach. The examples are systematically constructed to share position frequency matrices with missense training examples, which organizes the increased complexity in a manageable way while achieving improved generalization.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent systematically changes training parameters by introducing supplemental benign training examples with specific labeling schemes. This structured parameter change improves generalization ability while keeping the increased complexity organized and manageable through consistent labeling and data construction protocols.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10540591B2Deep learning-based techniques for pre-training deep convolutional neural networks
Publication Date: 2020.01.21 ILLUMINA INC
  • US10540591B2 patent drawing
  • US10540591B2 patent drawing
  • US10540591B2 patent drawing

AI summary

The technology disclosed includes systems and methods to reduce overfitting of neural network-implemented models that process sequences of amino acids and accompanying position frequency matrices. The system generates supplemental training example sequence pairs, labelled benign, that include a start location, through a target amino acid location, to an end location. A supplemental sequence pair supplements a pathogenic or benign missense training example sequence pair. It has identical amino acids in a reference and an alternate sequence of amino acids. The system includes logic to input with each supplemental sequence pair a supplemental training position frequency matrix (PFM) that is identical to the PFM of the benign or pathogenic missense at the matching start and end location. The system includes logic to attenuate the training influence of the training PFMs during training the neural network-implemented model by including supplemental training example PFMs in the training data.