Nucleic Acid Tagging for Long-Range Genome Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sequencing technologies face challenges in generating accurate, high-quality, highly contiguous genome sequences due to the inability to span large repetitive regions and uncertainty in genomic rearrangements, leading to difficulties in de novo assembly and haplotype phasing.

Innovation Solution

The method involves crosslinking and covalently linking distant DNA segments using chromatin conformation to generate extremely long-range read pairs (XLRPs) that span genomic distances of hundreds of kilobases to megabases, allowing for high-quality assembly and chromosome-level phasing with reduced data requirements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short read sequencing is used, then sequencing cost is reduced and throughput is increased, but the ability to span repetitive regions and achieve accurate de novo assembly deteriorates

Engineering Contradiction:
Improvesequencing throughputVSAvoidassembly accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the genome sequencing problem into two complementary parts: (1) short read sequencing for high-throughput coverage and (2) long-range chromatin conformation capture (Hi-C) for structural information. The short reads are used for initial assembly while the Hi-C data provides long-range contact frequencies to resolve repetitive regions and scaffold contigs, thereby maintaining both high throughput and assembly accuracy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces chromatin conformation capture (Hi-C) data as an intermediary to bridge the gap between short reads and genomic structure. The Hi-C methodology captures physical proximity information through crosslinking and sequencing of chromatin interactions, serving as a mediator that connects dispersed short reads to their spatial relationships, enabling accurate de novo assembly without requiring long reads

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If long-range DNA sequencing is implemented, then repetitive region spanning and de novo assembly quality improve, but sequencing cost and complexity increase

Engineering Contradiction:
Improveassembly qualityVSAvoidsequencing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges two established methodologies into a unified workflow: standard short read sequencing and chromatin conformation capture (Hi-C). By combining these approaches, the system leverages the high throughput of short reads with the long-range structural information from Hi-C, achieving high-quality assembly without the complexity of entirely new long-read sequencing technologies

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent makes the Hi-C methodology universal by applying it to any genome sequencing project requiring de novo assembly or scaffolding. The same chromatin crosslinking and sequencing protocol can be applied across different organisms and genome types, providing a multi-functional solution that works for both simple and complex genomes, thereby reducing overall system complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If traditional short read sequencing is used, then sequencing cost is low, but haplotype phasing accuracy and long-distance variant association deteriorate

Engineering Contradiction:
Improvesequencing costVSAvoidphasing accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent uses chromatin conformation capture (Hi-C) data as an intermediary to recover haplotype phase information from short read sequencing. The physical proximity information captured in Hi-C contacts serves as a mediator that links variants across long distances, enabling accurate haplotype phasing and variant association without requiring expensive long-read or linked-read technologies

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables cost-effective de novo assembly and accurate phasing of heterozygous single nucleotide polymorphisms (SNPs) with high accuracy, achieving genome sequencing accuracy and haplotype phasing accuracy of 99% or greater, surpassing the capabilities of substantially more costly and laborious methods.

Implementation Method 1

crosslinking and covalently linking distant DNA segments using chromatin conformation

Methodology Applied
Scientific EffectCrosslinking:

Data Source

PatentUS20250305028A1Tagging nucleic acids for sequence assembly
Publication Date: 2025.10.02 DOVETAIL GENOMICS LLC
  • US20250305028A1 patent drawing
  • US20250305028A1 patent drawing
  • US20250305028A1 patent drawing

AI summary

Various approaches for generating long-distance contiguity information to facilitate contig assembly and phase determination are disclosed. Nucleic acids are assembled into complexes using binding moieties such that, when the nucleic acid backbones are cleaved, the ensuing fragments remain bound. Exposed ends are tagged and ligated either to one another or to tagging moieties such as oligo labels. Ligated junctions are sequenced, and the sequence information is used to assemble contigs into common scaffolds or to assign phase information. Various approaches to tagging the exposed ends are presented.