Machine Learning Variant Triage Bypassing Sanger Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genetic screening methods, particularly Next Generation Sequencing (NGS), face challenges in efficiently confirming high-confidence variants, leading to increased costs, labor, and turnaround time for diagnostic results.
Innovation Solution
A computer-implemented method utilizing a two-tier machine learning process to select variants for bypassing Sanger sequencing, reducing the need for confirmatory testing by evaluating variant types and quality features to determine true positive probability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If Sanger sequencing confirmation is performed on all NGS-detected variants, then measurement precision and reliability are improved, but productivity is reduced due to labor-intensive manual confirmation processes
Solution Approach 1:
The patent replaces the manual mechanical process of Sanger sequencing confirmation with an automated machine learning-based classification system. The ML model automatically evaluates NGS variants and determines which ones require Sanger confirmation, eliminating the need for manual review of all variants and significantly increasing throughput while maintaining accuracy through algorithmic decision-making
Solution Approach 2:
The system enables self-service by allowing the ML model to autonomously triage variants without human intervention. The model independently assesses variant characteristics, applies classification criteria, and generates recommendations for confirmation, making the system self-sufficient in the variant triage process while human experts only review final recommendations
2Reliability
If Sanger sequencing is performed on all detected variants, then reliability of diagnostic results is improved, but loss of time increases due to extended turnaround time
Solution Approach 1:
The patent extracts only the necessary subset of variants that require Sanger confirmation by using the ML model to identify high-confidence variants that can be reported directly. This selective approach removes unnecessary confirmation steps for reliable variants while maintaining thorough confirmation for uncertain cases, thereby reducing overall turnaround time without compromising diagnostic reliability
Solution Approach 2:
The ML model performs preliminary classification of variants before Sanger sequencing is conducted. By pre-evaluating variant characteristics and predicting which ones need confirmation, the system avoids performing time-consuming Sanger sequencing on variants that are already high-confidence, thus reducing total processing time while ensuring reliable variants are identified early
3Productivity
If machine learning models are trained on prior concordance data, then productivity is improved by reducing Sanger confirmation needs, but measurement precision may be compromised due to biases in training data
Solution Approach 1:
The patent applies local quality by using different training strategies for different aspects of the ML model. The model is trained to recognize specific patterns in NGS data that correlate with true positives while incorporating techniques to mitigate biases in the training data. This localized approach to quality control in different parts of the training process maintains both productivity gains and measurement precision
Solution Approach 2:
The system incorporates feedback mechanisms where the ML model's predictions are evaluated against actual Sanger confirmation results. This feedback loop allows the model to learn from discrepancies and continuously improve its accuracy, ensuring that productivity gains from reduced confirmation do not come at the expense of long-term measurement precision
Data Source
AI summary
The present disclosure relates to a sequencing platform and workflow that leverages machine learning algorithms in genetic assays to bypass confirmatory Sange sequencing for high-confidence variants. Aspects are directed towards performing next generation sequencing (NGS) on nucleic acid obtained from a biological sample of a subject to generate sequencing data; extracting variant information from the sequencing data, wherein the information includes variant types and quality features; clustering variants into a subset of variants based on the variant types; generating a predicted status of each variant in the subset of variants based on the one or more quality features using a first machine learning model; generating a confirmatory status of each variant with an unknown status as the predicted status using a second machine learning model; and performing Sanger sequencing on nucleic acid molecules comprising variants with the absence status.


