Phased Read-Sets Using Long-Range DNA Linking for Genome Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

High-throughput sequencing technologies face challenges in spanning large repetitive regions of genomes due to short read lengths and small insert sizes, leading to difficulties in de novo assembly and phasing information over long distances, particularly in complex datasets like environmental samples containing refractory microbes.

Innovation Solution

Generation of extremely long-range read pairs (XLRPs) that span genomic distances of hundreds of kilobases to megabases by utilizing chromatin conformation to covalently link distant DNA segments, enabling accurate assembly and phasing through in vitro proximity-ligation methods.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short read lengths and small insert sizes are used in next generation sequencing, then sequencing throughput and cost-effectiveness are improved, but the ability to span large repetitive regions and achieve accurate de novo assembly deteriorates

Engineering Contradiction:
Improvesequencing throughputVSAvoidassembly accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent divides the sequencing approach into two complementary parts: standard short-read sequencing for high-throughput data generation and long-range read pair sequencing for structural context. This segmentation allows each method to optimize for its specific function while collectively solving the assembly accuracy problem.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces long-range read pairs as an intermediary element that bridges the gap between short reads and complete genome assembly. These read pairs serve as mediators that provide contextual information about repetitive regions and chromosomal architecture, enabling accurate assembly without requiring excessively long individual reads.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If short read lengths are used, then sequencing cost and time are reduced, but phasing information over long distances becomes indeterminable

Engineering Contradiction:
Improvesequencing timeVSAvoidphasing information
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The patent performs preliminary long-range read pair sequencing to establish phasing information and chromosomal context before completing the full de novo assembly. This preliminary action provides essential phasing data that guides subsequent assembly steps, reducing overall sequencing time while maintaining information quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent adds a temporal dimension to the sequencing process by performing iterative rounds of sequencing and assembly. In early rounds, long-range read pairs establish phasing information; in later rounds, standard reads fill in gaps. This dimensional approach to sequencing allows phasing to be determined without requiring all reads to be extremely long simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of manufacture

If conventional sequencing methods are used on complex environmental samples, then data generation is simplified, but de novo assembly of highly complex datasets becomes intractable

Engineering Contradiction:
Improvedata generation simplicityVSAvoidassembly complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent segments the assembly process into multiple manageable stages: initial assembly of individual contigs from short reads, followed by scaffolding using long-range read pairs, and final integration. This segmentation transforms the intractable problem of assembling entire complex genomes into a series of solvable sub-problems.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces long-range read pairs as an intermediary data structure that simplifies the assembly of complex environmental samples. These read pairs provide chromosomal context that organizes fragmented contigs into coherent genomic structures, reducing the computational complexity of assembly without requiring overly complex sequencing protocols.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach allows for high-quality, cost-effective de novo assembly and re-sequencing with improved phasing accuracy, providing comprehensive understanding of microbial communities and their impact on health and disease, while reducing data requirements.

Implementation Method 1

These techniques include in vitro proximity-ligation methods, e.g. for fecal metagenomics applications

Methodology Applied
Scientific EffectProximity ligation:

Implementation Method 2

The disclosure enables distant segments to be brought together and covalently linked by chromatin conformation, thereby physically connecting previously distant portions of the DNA molecule

Methodology Applied
Scientific EffectChromatin conformation:

Data Source

PatentUS20250305029A1Generation of phased read-sets for genome assembly and haplotype phasing
Publication Date: 2025.10.02 DOVETAIL GENOMICS LLC
  • US20250305029A1 patent drawing
  • US20250305029A1 patent drawing
  • US20250305029A1 patent drawing

AI summary

The disclosure provides methods to assemble genomes of eukaryotic or prokaryotic organisms. The disclosure provides methods for haplotype phasing and meta-genomics assemblies. The disclosure provides a streamlined method for accomplishing these tasks, such that intermediates need not be labeled by an affinity label to facilitate binding to a solid surface. The disclosure also provides methods and compositions for the de novo generation of scaffold information, linkage information, and genome information for unknown organisms in heterogeneous metagenomic samples or samples obtained from multiple individuals. Practice of the methods can allow de novo sequencing of entire genomes of uncultured or unidentified organisms in heterogeneous samples, or the determination of linkage information for nucleic acid molecules in samples comprising nucleic acids obtained from multiple individuals.