Circular Pair-Locked DNA Sequencing for Modified Base Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sequencing technologies face challenges in accuracy, particularly with single molecule sequencing methods which are prone to errors due to weak signals and fast polymerase reactions, leading to low accuracy in assembling large genome sequences and detecting single-base modifications, and there is a need for high-throughput methods to determine DNA methylation profiles effectively.
Innovation Solution
The method involves creating a circular nucleic acid molecule with insert-sample units of known sequence, obtaining sequence data from multiple units, calculating scores for sequence comparison, accepting or rejecting repeats based on upstream and downstream scores, and compiling an accepted sequence set to determine the nucleic acid sample sequence and identify modified bases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If single molecule sequencing is used to achieve high throughput and low cost, then productivity is improved, but measurement precision deteriorates due to weak signals and fast polymerase reactions
Solution Approach 1:
The patent creates multiple copies (2-10 copies) of each original DNA molecule through rolling circle amplification. Each copy is sequenced independently, and the results are compared to identify and correct errors. This copying approach maintains the high throughput of single molecule sequencing while improving accuracy through redundancy
Solution Approach 2:
The patent implements a feedback mechanism where sequence data from multiple copies are compared against each other and against a reference genome. Discrepancies trigger re-sequencing or re-analysis, allowing the system to self-correct errors and improve measurement precision while maintaining productivity
2Measurement precision
If single molecule sequencing is used to detect single-base modifications, then measurement precision is improved, but reliability deteriorates due to low raw data accuracy
Solution Approach 1:
The patent generates multiple copies of each DNA molecule containing potential modified bases. By sequencing these copies and comparing the results, the system can distinguish true modified bases from sequencing errors, thereby improving reliability while maintaining the ability to detect single-base modifications
Solution Approach 2:
The patent performs preliminary enrichment and purification steps before sequencing to concentrate modified bases and remove contaminants. This preliminary action improves the signal-to-noise ratio, enhancing both measurement precision for modified base detection and overall data reliability
3Productivity
If fragmented reads are assembled to reconstruct whole genome sequence, then productivity is improved, but manufacturing precision deteriorates due to assembly difficulties from low accuracy reads
Solution Approach 1:
The patent creates multiple copies of each fragmented read through rolling circle amplification. These redundant copies provide overlapping sequence information that facilitates more accurate assembly by allowing the system to resolve ambiguities and correct errors in individual reads, thereby improving manufacturing precision while maintaining productivity
Solution Approach 2:
The patent adds a temporal dimension to the assembly process by incorporating sequence data from multiple time points (multiple sequencing runs of the same copies). This multi-dimensional approach provides additional constraints and information for the assembly algorithm, improving genome reconstruction accuracy without sacrificing productivity
Data Source
AI summary
Disclosed herein are methods of determining the sequence and/or positions of modified bases in a nucleic acid sample present in a circular molecule with a nucleic acid insert of known sequence comprising obtaining sequence data of at least two insert-sample units. In some embodiments, the methods comprise obtaining sequence data using circular pair-locked molecules. In some embodiments, the methods comprise calculating scores of sequences of the nucleic acid inserts by comparing the sequences to the known sequence of the nucleic acid insert, and accepting or rejecting repeats of the sequence of the nucleic acid sample according to the scores of one or both of the sequences of the inserts immediately upstream or downstream of the repeats of the sequence of the nucleic acid sample.


