Phased Read-Sets Using Long-Range DNA Linking for Genome Assembly
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High-throughput sequencing technologies face challenges in spanning large repetitive regions of genomes due to short read lengths and small insert sizes, leading to difficulties in de novo assembly and phasing information over long distances, particularly in complex datasets like environmental samples containing refractory microbes.
Innovation Solution
Generation of extremely long-range read pairs (XLRPs) that span genomic distances of hundreds of kilobases to megabases by utilizing chromatin conformation to covalently link distant DNA segments, enabling accurate assembly and phasing through in vitro proximity-ligation methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If short read lengths and small insert sizes are used in next generation sequencing, then sequencing throughput and cost-effectiveness are improved, but the ability to span large repetitive regions and achieve accurate de novo assembly deteriorates
Solution Approach 1:
The patent divides the sequencing approach into two complementary parts: standard short-read sequencing for high-throughput data generation and long-range read pair sequencing for structural context. This segmentation allows each method to optimize for its specific function while collectively solving the assembly accuracy problem.
Solution Approach 2:
The patent introduces long-range read pairs as an intermediary element that bridges the gap between short reads and complete genome assembly. These read pairs serve as mediators that provide contextual information about repetitive regions and chromosomal architecture, enabling accurate assembly without requiring excessively long individual reads.
2Loss of time
If short read lengths are used, then sequencing cost and time are reduced, but phasing information over long distances becomes indeterminable
Solution Approach 1:
The patent performs preliminary long-range read pair sequencing to establish phasing information and chromosomal context before completing the full de novo assembly. This preliminary action provides essential phasing data that guides subsequent assembly steps, reducing overall sequencing time while maintaining information quality.
Solution Approach 2:
The patent adds a temporal dimension to the sequencing process by performing iterative rounds of sequencing and assembly. In early rounds, long-range read pairs establish phasing information; in later rounds, standard reads fill in gaps. This dimensional approach to sequencing allows phasing to be determined without requiring all reads to be extremely long simultaneously.
3Ease of manufacture
If conventional sequencing methods are used on complex environmental samples, then data generation is simplified, but de novo assembly of highly complex datasets becomes intractable
Solution Approach 1:
The patent segments the assembly process into multiple manageable stages: initial assembly of individual contigs from short reads, followed by scaffolding using long-range read pairs, and final integration. This segmentation transforms the intractable problem of assembling entire complex genomes into a series of solvable sub-problems.
Solution Approach 2:
The patent introduces long-range read pairs as an intermediary data structure that simplifies the assembly of complex environmental samples. These read pairs provide chromosomal context that organizes fragmented contigs into coherent genomic structures, reducing the computational complexity of assembly without requiring overly complex sequencing protocols.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for high-quality, cost-effective de novo assembly and re-sequencing with improved phasing accuracy, providing comprehensive understanding of microbial communities and their impact on health and disease, while reducing data requirements.
Implementation Method 1
These techniques include in vitro proximity-ligation methods, e.g. for fecal metagenomics applications
Implementation Method 2
The disclosure enables distant segments to be brought together and covalently linked by chromatin conformation, thereby physically connecting previously distant portions of the DNA molecule
Data Source
AI summary
The disclosure provides methods to assemble genomes of eukaryotic or prokaryotic organisms. The disclosure provides methods for haplotype phasing and meta-genomics assemblies. The disclosure provides a streamlined method for accomplishing these tasks, such that intermediates need not be labeled by an affinity label to facilitate binding to a solid surface. The disclosure also provides methods and compositions for the de novo generation of scaffold information, linkage information, and genome information for unknown organisms in heterogeneous metagenomic samples or samples obtained from multiple individuals. Practice of the methods can allow de novo sequencing of entire genomes of uncultured or unidentified organisms in heterogeneous samples, or the determination of linkage information for nucleic acid molecules in samples comprising nucleic acids obtained from multiple individuals.


