Short Read Assembly into Long Sequences via Circular DNA

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Massively parallel DNA sequencing platforms produce short, inaccurate reads that limit their utility for applications like de novo genome assembly and full-length cDNA sequencing due to their short lengths and high error rates.

Innovation Solution

The method involves circularizing target DNA fragments with adaptor molecules, amplifying, fragmenting, and adding defined sequences to create a library that allows for the clustering and assembly of short reads into longer subassemblies, enabling the sequencing of kilobase-scale DNA molecules.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If massively parallel DNA sequencing platforms are used, then productivity and cost-effectiveness are improved, but read length and accuracy deteriorate

Engineering Contradiction:
Improvesequencing throughputVSAvoidread length
Core Design Contradiction:
ProductivityVSLength of stationary object

Solution Approach 1:

The invention segments the sequencing process into two distinct phases: (1) massively parallel short-read sequencing of many DNA fragments, and (2) computational assembly of these short reads into long subassemblies using overlap-layout-consensus algorithms. This segmentation allows each phase to optimize for its specific strength while the computational assembly bridges the read length limitation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention introduces computational assembly algorithms as an intermediary process between short-read sequencing and long-read applications. These algorithms act as a mediator that takes numerous short reads covering the same genomic region and reconstructs them into accurate long subassemblies, effectively translating short-read data into long-read equivalent information.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If massively parallel DNA sequencing platforms are used, then productivity is improved, but measurement precision deteriorates

Engineering Contradiction:
Improvesequencing throughputVSAvoidsequence accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The invention merges multiple short reads that originate from the same long DNA fragment into a single consensus sequence. By combining information from numerous overlapping short reads, the assembly algorithm produces a consensus sequence with accuracy superior to individual short reads, while maintaining high throughput.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The assembly process incorporates iterative feedback mechanisms where initial assemblies are refined through multiple rounds of overlap detection, consistency checking, and error correction. The algorithm uses feedback from read overlaps and consensus building to progressively improve sequence accuracy while processing large datasets efficiently.

Inventive Principle:
Principle #23Feedback

3Productivity

If short reads are used for de novo genome assembly, then productivity is improved, but manufacturing precision deteriorates

Engineering Contradiction:
Improveassembly efficiencyVSAvoidassembly contiguity
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The invention transitions the assembly problem from one dimension (individual read lengths) to another dimension (collective read coverage). Instead of being limited by the length of single reads, the system uses the collective information from thousands of short reads covering the same region, effectively achieving long-range contiguity through increased sampling depth rather than increased read length.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250101413A1Sequence tag directed subassembly of short sequencing reads into long sequencing reads
Publication Date: 2025.03.27 UNIV OF WASHINGTON
  • US20250101413A1 patent drawing
  • US20250101413A1 patent drawing
  • US20250101413A1 patent drawing

AI summary

The invention provides methods for preparing DNA sequencing libraries by assembling short read sequencing data into longer contiguous sequences for genome assembly, full length cDNA sequencing, metagenomics, and the analysis of repetitive sequences of assembled genomes.