Long Nucleic Acid Sequencing via Template Segmentation and Barcoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current nucleic acid sequencing technologies are limited by short read lengths, making them unreliable for identifying large genetic variants and structural variants, and require additional experimentation for phasing variants, while existing long-read technologies are inaccurate, costly, and not viable for whole genome sequencing.
Innovation Solution
A method involving stretching and priming of template nucleic acid molecules on a substrate with immobilized primers, followed by extension reactions and sequencing to generate mega-base range reads, accurately identifying various genetic variants and phasing them to appropriate chromosomes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If short-read sequencing technology is used, then sequencing throughput and accuracy for single nucleotide variants is improved, but read length is limited making it unreliable for identifying large insertions/deletions and structural variants
Solution Approach 1:
The patent segments the template nucleic acid molecule into multiple fragments by hybridizing them to primers distributed across a substrate. Each fragment is sequenced independently, and the barcodes associated with each primer allow computational assembly of the full-length sequence, effectively overcoming the read length limitation while maintaining high accuracy through multiple short reads covering the same region.
Solution Approach 2:
The patent introduces barcodes as intermediary elements that link short read sequences to their physical positions on the template molecule. These barcodes enable the reconstruction of long-range genomic information from short reads by serving as markers that track which fragment originated from which location, thus bridging the gap between short read capability and long-range sequencing needs.
2Length of moving object
If long-read sequencing technology is used, then read length is improved for identifying structural variants, but accuracy and throughput deteriorate while cost increases
Solution Approach 1:
The patent divides the long template molecule into multiple short fragments that can be sequenced with high accuracy using standard short-read technology. By sequencing many short fragments and assembling them using barcode information, the method achieves long-read equivalent functionality without sacrificing the accuracy advantages of short-read sequencing.
Solution Approach 2:
The patent changes the parameter of read length from a physical constraint to a computational solution. Instead of physically sequencing long continuous reads, the method uses computational assembly of multiple short reads with associated barcodes to reconstruct long-range sequences, thereby achieving long read lengths without the accuracy and cost penalties of existing long-read technologies.
3Productivity
If short-read sequencing is used, then cost and throughput are improved, but additional experimentation is required for phasing variants
Solution Approach 1:
The patent uses barcodes as intermediary markers that directly encode positional information on the template molecule. These barcodes enable phasing of variants by indicating which variants occur on the same physical molecule, eliminating the need for additional phasing experiments while maintaining high throughput sequencing.
Solution Approach 2:
The barcode system serves multiple functions simultaneously: it tracks the physical position of each fragment, enables assembly of long-range sequences, and provides phasing information for variants. This multi-functionality eliminates the need for separate phasing experiments, reducing overall experimental complexity while maintaining high throughput.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables the production of very long sequencing reads, accurately identifying genetic variants and phasing them, thereby improving the feasibility and utility of whole genome sequencing by providing detailed sequence information.
Implementation Method 1
covalently bonding initiator species to said surface
Implementation Method 2
contacting a stretched and primed template nucleic acid molecule with a substrate, said primed template nucleic acid molecule comprising one or more primer-binding sites, said substrate comprising primers immobilized thereon, each primer comprising: (i) a region complementary to a primer-binding site
Implementation Method 3
conducting extension reactions using said primers and said template nucleic acid molecule as a template, thereby generating extension products
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided herein are methods and compositions for the sequencing of long nucleic acids, such as DNA. The methods and compositions are suited for the spatial labeling and sequencing of long nucleic acid molecules.