Synthetic Long Read DNA Sequencing via Molecular Barcodes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current ultra-high throughput sequencing technologies generate short reads that are insufficient for answering biological questions, such as phasing variants on homologous chromosomes or performing de novo sequencing of mammalian-sized genomes, as they are too short to efficiently find structural variations and rearrangements.
Innovation Solution
A method that involves denaturing DNA molecules, attaching primers with barcodes, dividing and pooling samples to generate synthetic long reads by assembling short reads with identical barcodes into a single long read sequence, allowing for the creation of DNA sequences of various lengths, including up to entire chromosomes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short read sequencing is used to achieve high throughput and low cost, then productivity and ease of manufacture are improved, but the read length is insufficient for de novo sequencing and structural variation detection
Solution Approach 1:
The patent segments long DNA molecules into multiple short reads by fragmenting the DNA and sequencing each fragment separately. Each short read is tagged with a molecular barcode that identifies its origin from the parent long molecule. This segmentation allows high-throughput sequencing of many short fragments while the barcode information enables virtual reassembly into long reads, thus achieving both high productivity and effective long read length.
Solution Approach 2:
The patent implements a nested structure where multiple short reads are nested within a single long read through the use of molecular barcodes. Each short read contains a portion of the original long molecule's sequence, and the barcode nested within each short read allows computational reassembly to reconstruct the parent long read. This nested approach enables short read technologies to produce long read equivalents.
2Quantity of substance
If short reads are used for sequencing, then sequencing cost and throughput are improved, but the ability to phase variants and detect structural variations deteriorates
Solution Approach 1:
The patent uses molecular barcodes as feedback mechanisms to track and associate short reads with their parent long molecules throughout the sequencing process. The barcode information provides feedback that enables computational algorithms to correctly phase variants by identifying which short reads originate from the same parental allele. This feedback system preserves phasing information despite the use of short reads.
Solution Approach 2:
The molecular barcode acts as an intermediary that bridges the gap between short reads and long molecule information. The barcode is attached to multiple short reads derived from the same long molecule, serving as an intermediary identifier that allows computational reassembly and phasing without requiring direct long read sequencing. This intermediary enables information preservation through indirect association.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables the use of short read DNA sequencing for full de novo genome sequencing by generating synthetic long reads of at least 500 base pairs to several kilobases, effectively overcoming the limitations of short read lengths in existing technologies.
Implementation Method 1
extending primers attached to the DNA molecule with a polymerase
Implementation Method 2
denaturing a plurality of DNA molecules
Data Source
AI summary
The disclosure describes a method for sequencing long portions of DNA sequence by assembling a plurality of shorter polynucleotide reads. Generally, the method includes annealing a plurality of primers to a denatured DNA molecule, appending a barcode polynucleotide to the 5′ end of the primer, subjecting the DNA molecules to a plurality of cycles of (1) pooling, (2) dividing, and (3) appending a barcode polynucleotide to the 5′ end of the primer, sequencing the barcode polynucleotides and the genomic DNA, and assembling the short read polynucleotide sequences having identical barcode polynucleotides.


