Barcode-Directed DNA Read Assembly for Kilobase Sequencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current next-generation sequencing technologies have a maximum read length of around 250 bases, limiting the quality of genome assembly and preventing broader applications such as sequencing larger genomes and long-range haplotype analysis.

Innovation Solution

A method involving barcode assignment, amplification, fragmentation, and juxtaposing barcode-containing fragments to random short segments of the original DNA template molecule, followed by assembly to generate extended sequence reads.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If next-generation sequencing technologies are used with current read length limits, then sequencing speed and cost-effectiveness are maintained, but read length remains limited to around 250 bases, reducing genome assembly quality

Engineering Contradiction:
Improveread lengthVSAvoidsequencing process complexity
Core Design Contradiction:
Length of moving objectVSDevice complexity

Solution Approach 1:

The patent divides the original long DNA template into multiple shorter fragments that can be sequenced with current technology. Each fragment is tagged with a barcode that identifies its origin from the parent template molecule. After sequencing, these fragmented reads are computationally reassembled using the barcode information to reconstruct the extended original sequence, effectively achieving long-read sequencing through short-read technology.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces barcodes as intermediary molecular tags that bridge the gap between short sequenced fragments and the original long template. These barcodes serve as mediators that carry identity information through the fragmentation and sequencing process, enabling the computational reassembly of extended sequences from multiple short reads without requiring the entire long molecule to be sequenced in one continuous read.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If read length is increased to several kilobases, then genome assembly quality and haplotype analysis capability are improved, but current sequencing technologies cannot achieve this read length

Engineering Contradiction:
Improvegenome assembly qualityVSAvoidapplication range
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

By segmenting the long template into manageable fragments that fit within current sequencing read length limits while retaining barcode identifiers, the method enables reliable genome assembly and haplotype analysis using existing sequencing platforms. The segmentation allows the system to work around technological read length limitations while maintaining the ability to reconstruct long-range genomic information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal methodology that can be applied across different sequencing platforms and genome types. The barcode-based fragmentation and reassembly approach is platform-agnostic and can be used for various applications including genome assembly, haplotype analysis, and targeted region sequencing, making the solution broadly adaptable rather than limited to specific technologies or use cases.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If barcode tagging and fragmentation steps are added to extend read length, then sequencing coverage and assembly accuracy are improved, but processing time and workflow complexity increase

Engineering Contradiction:
Improvesequence assembly accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The barcode tagging is performed in advance during library preparation, before fragmentation and sequencing occur. This preliminary action ensures that each fragment carries its identity marker from the outset, eliminating the need for time-consuming post-sequencing identification steps and enabling parallel processing of multiple fragments throughout the workflow.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The barcode sequence serves as a copy of the parent template's identity information that is replicated across all fragments derived from that template. This copying mechanism allows rapid computational matching and reassembly without requiring complex physical manipulation or time-intensive verification steps, as the barcode copies provide direct identifiers for reconstitution.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20260049304A1Method for generating extended sequence reads
Publication Date: 2026.02.19 AGENCY FOR SCI TECH & RES
  • US20260049304A1 patent drawing
  • US20260049304A1 patent drawing
  • US20260049304A1 patent drawing

AI summary

The present invention provides an approach to increase the effective read length of commercially available sequencing platforms to several kilobases and be broadly applied to obtain long sequence reads from mixed template populations. A method for generating extended sequence reads of long DNA molecules in a sample, comprising the steps of: assigning a specific barcode sequence to each template DNA molecule in a sample to obtain barcode-tagged molecules; amplifying the barcode-tagged molecules; fragmenting the amplified barcode-tagged molecules to obtain barcode-containing fragments; juxtaposing the barcode-containing fragments to random short segments of the original DNA template molecule during the process of generating a sequencing library to obtain demultiplexed reads; and assembling the demultiplexed reads to obtain extended sequence reads for each DNA template molecule, is disclosed.