Short Read Assembly into Long Sequences via Circular DNA
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Massively parallel DNA sequencing platforms produce short, inaccurate reads that limit their utility for applications like de novo genome assembly and full-length cDNA sequencing due to their short lengths and high error rates.
Innovation Solution
The method involves circularizing target DNA fragments with adaptor molecules, amplifying, fragmenting, and adding defined sequences to create a library that allows for the clustering and assembly of short reads into longer subassemblies, enabling the sequencing of kilobase-scale DNA molecules.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If massively parallel DNA sequencing platforms are used, then productivity and cost-effectiveness are improved, but read length and accuracy deteriorate
Solution Approach 1:
The invention segments the sequencing process into two distinct phases: (1) massively parallel short-read sequencing of many DNA fragments, and (2) computational assembly of these short reads into long subassemblies using overlap-layout-consensus algorithms. This segmentation allows each phase to optimize for its specific strength while the computational assembly bridges the read length limitation.
Solution Approach 2:
The invention introduces computational assembly algorithms as an intermediary process between short-read sequencing and long-read applications. These algorithms act as a mediator that takes numerous short reads covering the same genomic region and reconstructs them into accurate long subassemblies, effectively translating short-read data into long-read equivalent information.
2Productivity
If massively parallel DNA sequencing platforms are used, then productivity is improved, but measurement precision deteriorates
Solution Approach 1:
The invention merges multiple short reads that originate from the same long DNA fragment into a single consensus sequence. By combining information from numerous overlapping short reads, the assembly algorithm produces a consensus sequence with accuracy superior to individual short reads, while maintaining high throughput.
Solution Approach 2:
The assembly process incorporates iterative feedback mechanisms where initial assemblies are refined through multiple rounds of overlap detection, consistency checking, and error correction. The algorithm uses feedback from read overlaps and consensus building to progressively improve sequence accuracy while processing large datasets efficiently.
3Productivity
If short reads are used for de novo genome assembly, then productivity is improved, but manufacturing precision deteriorates
Solution Approach 1:
The invention transitions the assembly problem from one dimension (individual read lengths) to another dimension (collective read coverage). Instead of being limited by the length of single reads, the system uses the collective information from thousands of short reads covering the same region, effectively achieving long-range contiguity through increased sampling depth rather than increased read length.
Data Source
AI summary
The invention provides methods for preparing DNA sequencing libraries by assembling short read sequencing data into longer contiguous sequences for genome assembly, full length cDNA sequencing, metagenomics, and the analysis of repetitive sequences of assembled genomes.


