Paired Code Tags for De Novo Nucleic Acid Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current nucleic acid sequencing methods face challenges in assembling contiguous reads, especially through repetitive sequences, requiring larger DNA amounts and a reference genome for verification, and often result in lost order information during DNA fragmentation.

Innovation Solution

The method involves preparing a library of template nucleic acids by inserting unique barcodes into target nucleic acids, fragmenting them, and using paired barcode sequences to assemble sequencing data without a reference genome, preserving order information and enabling de novo assembly of target nucleic acids.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If DNA fragmentation is performed for sequencing, then sequencing throughput is improved, but order information is lost

Engineering Contradiction:
Improvesequencing throughputVSAvoidorder information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The DNA molecule is segmented into multiple fragments through controlled fragmentation, with each fragment receiving a unique molecular identifier (UMI) barcode. This segmentation enables parallel processing of multiple fragments simultaneously, improving sequencing throughput while the barcodes preserve the original positional information for later reassembly.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Unique molecular identifier (UMI) barcodes serve as intermediary tags that are attached to each DNA fragment during library preparation. These barcodes act as mediators that carry positional information through the fragmentation and sequencing process, enabling reconstruction of the original DNA sequence order after high-throughput sequencing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If traditional assembly methods are used for repetitive sequences, then assembly accuracy is improved, but a reference genome is required

Engineering Contradiction:
Improveassembly accuracyVSAvoidreference genome requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The method creates multiple copies of the original DNA sequence information through the use of UMI barcodes on different fragments. By sequencing multiple barcoded fragments that originate from the same genomic region, the system can reconstruct repetitive sequences de novo without requiring a reference genome, achieving high assembly accuracy through redundant information copying.

Inventive Principle:
Principle #26Copying

3Reliability

If larger DNA amounts are used for sequencing, then coverage is improved, but sample requirements increase

Engineering Contradiction:
ImprovecoverageVSAvoidDNA amount
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The UMI barcode system serves multiple functions simultaneously: it tracks individual DNA molecules through fragmentation, enables high-throughput parallel sequencing, preserves positional information for assembly, and allows for de novo reconstruction without reference genomes. This multi-functionality achieves high coverage reliability without proportionally increasing DNA input requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240384262A1Linking sequence reads using paired code tags
Publication Date: 2024.11.21 ILLUMINA INC
  • US20240384262A1 patent drawing
  • US20240384262A1 patent drawing
  • US20240384262A1 patent drawing

AI summary

Artificial transposon sequences having code tags and target nucleic acids containing such sequences. Methods for making artificial transposons and for using their properties to analyze target nucleic acids.