Synthetic Long Read DNA Sequencing via Molecular Barcodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current ultra-high throughput sequencing technologies generate short reads that are insufficient for answering biological questions, such as phasing variants on homologous chromosomes or performing de novo sequencing of mammalian-sized genomes, as they are too short to efficiently find structural variations and rearrangements.

Innovation Solution

A method that involves denaturing DNA molecules, attaching primers with barcodes, dividing and pooling samples to generate synthetic long reads by assembling short reads with identical barcodes into a single long read sequence, allowing for the creation of DNA sequences of various lengths, including up to entire chromosomes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If short read sequencing is used to achieve high throughput and low cost, then productivity and ease of manufacture are improved, but the read length is insufficient for de novo sequencing and structural variation detection

Engineering Contradiction:
Improvesequencing throughputVSAvoidread length
Core Design Contradiction:
ProductivityVSLength of moving object

Solution Approach 1:

The patent segments long DNA molecules into multiple short reads by fragmenting the DNA and sequencing each fragment separately. Each short read is tagged with a molecular barcode that identifies its origin from the parent long molecule. This segmentation allows high-throughput sequencing of many short fragments while the barcode information enables virtual reassembly into long reads, thus achieving both high productivity and effective long read length.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where multiple short reads are nested within a single long read through the use of molecular barcodes. Each short read contains a portion of the original long molecule's sequence, and the barcode nested within each short read allows computational reassembly to reconstruct the parent long read. This nested approach enables short read technologies to produce long read equivalents.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Quantity of substance

If short reads are used for sequencing, then sequencing cost and throughput are improved, but the ability to phase variants and detect structural variations deteriorates

Engineering Contradiction:
Improvenumber of readsVSAvoidphasing information
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent uses molecular barcodes as feedback mechanisms to track and associate short reads with their parent long molecules throughout the sequencing process. The barcode information provides feedback that enables computational algorithms to correctly phase variants by identifying which short reads originate from the same parental allele. This feedback system preserves phasing information despite the use of short reads.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The molecular barcode acts as an intermediary that bridges the gap between short reads and long molecule information. The barcode is attached to multiple short reads derived from the same long molecule, serving as an intermediary identifier that allows computational reassembly and phasing without requiring direct long read sequencing. This intermediary enables information preservation through indirect association.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enables the use of short read DNA sequencing for full de novo genome sequencing by generating synthetic long reads of at least 500 base pairs to several kilobases, effectively overcoming the limitations of short read lengths in existing technologies.

Implementation Method 1

extending primers attached to the DNA molecule with a polymerase

Methodology Applied
Scientific EffectDNA polymerization: Enzyme

Implementation Method 2

denaturing a plurality of DNA molecules

Methodology Applied
Scientific EffectDNA denaturation: Heating

Data Source

PatentUS20220119867A1Synthetic long read DNA sequencing
Publication Date: 2022.04.21 UNM RAINFOREST INNOVATIONS
  • US20220119867A1 patent drawing
  • US20220119867A1 patent drawing
  • US20220119867A1 patent drawing

AI summary

The disclosure describes a method for sequencing long portions of DNA sequence by assembling a plurality of shorter polynucleotide reads. Generally, the method includes annealing a plurality of primers to a denatured DNA molecule, appending a barcode polynucleotide to the 5′ end of the primer, subjecting the DNA molecules to a plurality of cycles of (1) pooling, (2) dividing, and (3) appending a barcode polynucleotide to the 5′ end of the primer, sequencing the barcode polynucleotides and the genomic DNA, and assembling the short read polynucleotide sequences having identical barcode polynucleotides.