Nucleic Acid Base Calling Using Forward-Cognate Sequence Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid sequencing methods face challenges in accurately determining the true base identity at a locus due to high false positive rates, particularly when dealing with modified bases like 5-methylcytosine, 5-hydroxymethylcytosine, and cytosine, which are prone to miscalls.
Innovation Solution
The method involves generating forward and cognate polynucleotides through various chemical and enzymatic reactions, such as deamination, methylation, and oxidation, followed by sequencing to determine the true base identity at a locus using a computer-based analysis that considers the identities of corresponding bases in the forward and cognate polynucleotides, thereby reducing false positive rates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current nucleic acid sequencing technologies are used to identify base variants, then sequencing can be performed, but false positive rates are high and extensive sequencing coverage is required
Solution Approach 1:
The patent introduces an intermediary computational analysis step that compares forward and reverse sequencing reads to identify and filter false positive base variants. This intermediary process acts as a mediator between raw sequencing data and final variant identification, using algorithms to distinguish true biological variants from sequencing errors by examining patterns in paired-end reads and applying quality filters.
Solution Approach 2:
The patent segments the sequencing analysis into distinct components: forward read analysis, reverse read analysis, and integrated comparison. By dividing the problem into separate analytical stages and examining each strand independently before combining results, the method can identify inconsistencies that indicate false positives while preserving true variants that appear consistently across both directions.
2Measurement precision
If extensive sequencing coverage is used to achieve reliable mutation detection, then sensitivity improves, but cost and complexity increase
Solution Approach 1:
The patent implements a feedback mechanism where the computational analysis of paired-end reads provides information about the reliability of detected variants. By using the reverse read as feedback to validate or refute findings from the forward read, the system can achieve high confidence in mutation detection at lower overall coverage levels, as each read pair mutually validates the other's findings.
Solution Approach 2:
The patent performs preliminary computational filtering and validation of base variants during the sequencing analysis process itself, rather than requiring additional sequencing passes. By pre-processing the data to identify and remove false positives through algorithmic analysis of read patterns and quality metrics, the method achieves reliable detection without needing excessive sequencing depth.
3Measurement precision
If chemical or enzymatic reactions are applied to differentiate bases before sequencing, then base identification accuracy improves, but process complexity increases
Solution Approach 1:
The patent replaces complex chemical and enzymatic differentiation procedures with computational analysis methods. Instead of using multiple chemical reactions to modify and differentiate bases before sequencing, the invention uses bioinformatics algorithms to analyze sequencing data patterns, distinguish base types through computational means, and identify variants without requiring extensive wet-lab chemical processing.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach significantly reduces false positive rates to 1 in 100,000 or lower, enabling accurate determination of true base identities, even for modified bases, with sensitivity up to 90% for mutations below 0.1% frequency.
Implementation Method 1
determining a first identity of a first base at a locus of the forward polynucleotide and a second identity of a second base at or proximal to a corresponding locus of the cognate polynucleotide using sequencing
Implementation Method 2
applying chemical or enzymatic reactions such as deamination, methylation, or oxidation to differentiate between these bases before sequencing
Implementation Method 3
applying chemical or enzymatic reactions such as deamination, methylation, or oxidation to differentiate between these bases before sequencing
Implementation Method 4
applying chemical or enzymatic reactions such as deamination, methylation, or oxidation to differentiate between these bases before sequencing
Implementation Method 5
the forward polynucleotide and cognate polynucleotide are linked as a double-stranded polynucleotide via Watson-Crick base pairing
Data Source
AI summary
Provided herein are methods, systems, and compositions for determining a base in a polynucleotide. In various aspects, the methods, systems, and compositions presented herein are useful for performing 4-base, 5-base, or 6-base sequencing of polynucleotide molecules, for example, from liquid biopsy samples or wherein the base is a low frequency mutation.


