Flow-Signal Variant Calling for Short Nucleic Acid Variants
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sequencing methods struggle with high single-signal errors and inefficiency in detecting genetic variants, particularly single nucleotide polymorphisms (SNPs) and insertions/deletions (indels), often requiring costly and time-consuming high-depth sequencing to overcome errors.
Innovation Solution
A method involving sequencing nucleic acid molecules using non-terminating nucleotides in separate nucleotide flows, generating flow signals at specific positions, and determining match scores to accurately detect short genetic variants by analyzing flow signals rather than traditional base signals, allowing for computationally efficient variant calling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-depth sequencing is used to overcome single-signal errors, then measurement precision is improved, but cost and time consumption increase
Solution Approach 1:
The patent segments the sequencing process into multiple flows, where each flow sequences a subset of nucleotides. By dividing the complete sequence into multiple flows and analyzing them separately, the system can identify variants more accurately without requiring excessive sequencing depth, thus resolving the contradiction between accuracy and time consumption
Solution Approach 2:
The patent introduces a flow dimension to the sequencing analysis, transforming single-signal base calls into multi-flow signal patterns. This dimensional transformation allows the system to distinguish true variants from errors by examining signal consistency across multiple flows, improving accuracy without proportionally increasing sequencing depth or time
2Measurement precision
If high-depth sequencing is used to overcome single-signal errors, then measurement precision is improved, but cost increases
Solution Approach 1:
By segmenting the sequencing into multiple flows that each cover portions of the genome, the patent reduces the total sequencing depth required while maintaining accuracy. This segmentation approach lowers the quantity of sequencing cycles needed, thereby reducing cost without sacrificing variant detection precision
Solution Approach 2:
The patent creates multiple copies of sequence information across different flows, where each flow provides an independent observation of the same genomic region. This copying strategy allows the system to achieve high accuracy through consensus across multiple low-depth observations, reducing overall sequencing cost
3Device complexity
If reversible-terminator sequencing-by-synthesis is used, then device complexity is reduced, but measurement precision deteriorates due to single differentiated signal
Solution Approach 1:
The patent segments the base calling task across multiple flows, where each flow provides a differentiated signal for a subset of bases. By combining information from multiple segmented flows, the system achieves high measurement precision while maintaining the simplicity of reversible-terminator chemistry, resolving the contradiction between device simplicity and accuracy
Solution Approach 2:
The patent adds a flow dimension to the sequencing data, transforming single-signal base calls into multi-flow signal patterns. This dimensional enrichment allows the simple reversible-terminator chemistry to produce complex, high-precision base calling results by analyzing signal patterns across multiple flows rather than relying on complex chemistry
Data Source
AI summary
Methods for detecting a short genetic variant in a test sample are described herein. In some exemplary methods, the short genetic variant is called using one or match scores, which are determined using one or more sequencing data sets obtained from a test nucleic acid molecule, wherein the test sequencing data sets are determined by sequencing the test nucleic acid molecule using non-terminating nucleotides provided in separate nucleotide flows according to a flow-cycle order. Also described herein are methods of sequencing a test nucleic acid molecule using two or more different flow-cycle orders and/or extended flow cycle orders having five or more nucleotide flows per flow cycle.


