Combinatorial DNA Barcoding for Long-Fragment Genome Assembly

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional sequencing methods are limited by signal degradation and signal-to-noise ratios, leading to inefficient sequencing and assembly of complete sequences from shorter read lengths.

Innovation Solution

A method involving the fragmentation of double-stranded nucleic acids using Controlled Random Enzymatic (CoRE) techniques, where nucleotides are replaced with dNTP analogs, followed by enzymatic treatment to create gapped DNA, and then gap translation to form blunt-ended fragments, reducing GC bias and coverage bias.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of moving object

If conventional sequencing methods are used, then sequencing can be performed with standard techniques, but signal degradation limits the read length to only a few tens of nucleotides

Engineering Contradiction:
Improveread lengthVSAvoidsignal quality
Core Design Contradiction:
Length of moving objectVSReliability

Solution Approach 1:

The invention segments the original DNA molecule into multiple overlapping fragments, each of which can be sequenced independently to a sufficient depth. By creating a library of overlapping fragments and sequencing each to high coverage, the complete original sequence can be reconstructed through computational assembly, effectively overcoming the read length limitation while maintaining accuracy through redundancy

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention performs preliminary fragmentation and library preparation before sequencing, creating a collection of overlapping DNA fragments with known relationships. This preliminary action allows the sequencing process to work with manageable fragment lengths while the computational assembly step later reconstructs the complete sequence, separating the physical sequencing constraint from the logical sequence reconstruction

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If conventional sequencing methods are used, then standard protocols can be applied, but signal-to-noise ratios are insufficient for single-molecule sequencing

Engineering Contradiction:
Improvesignal-to-noise ratioVSAvoidsequencing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The invention creates multiple copies of the original DNA sequence through fragmentation and amplification, generating a library of identical sequence information distributed across many fragments. By sequencing each fragment independently and assembling them computationally, the signal can be reconstructed with high precision through consensus sequencing, effectively amplifying the signal-to-noise ratio through redundancy rather than direct physical amplification of the original molecule

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The invention introduces computational assembly algorithms as an intermediary step between fragment sequencing and final sequence determination. This computational mediator integrates signals from multiple fragmented sequences, resolves ambiguities, and reconstructs the complete original sequence with high confidence, effectively separating the low-signal fragment sequencing from the high-precision final sequence determination

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If DNA is fragmented using conventional methods, then fragmentation can be achieved, but GC bias and coverage bias occur in the resulting fragments

Engineering Contradiction:
Improvefragmentation processVSAvoidfragment uniformity
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The invention changes the chemical parameters of the fragmentation process by using transposase enzymes that insert adapters at random positions through a mechanism independent of sequence composition. This parameter change from sequence-dependent fragmentation (conventional methods) to enzyme-mediated random insertion (transposase) eliminates GC bias and produces uniformly distributed fragments across the entire genome, improving fragment uniformity while maintaining operational simplicity

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method produces reproducibly controlled fragments with reduced bias, enabling efficient sequencing and assembly of complete sequences, particularly through Long Fragment Read (LFR) sequencing.

Implementation Method 1

amplifying the DNA in the separate aliquots in the presence of a population of dNTPs that includes dNTP analogs, such that a number of nucleotides in the DNA are replaced by dNTP analogs

Methodology Applied
Scientific EffectEnzyme catalysis: Enzyme

Implementation Method 2

removing the dNTP analogs to form gapped DNA

Methodology Applied
Scientific EffectEnzymatic removal: Enzyme

Implementation Method 3

treating the gapped DNA to translate the gaps until gaps on opposite strands converge, thereby creating blunt-ended DNA fragments

Methodology Applied
Scientific EffectEnzymatic translation: Enzyme

Data Source

PatentUS12529163B2Library of DNA fragments tagged with combinatorial oligonucleotide bar codes for use in genome sequencing
Publication Date: 2026.01.20 COMPLETE GENOMICS INC
  • US12529163B2 patent drawing
  • US12529163B2 patent drawing
  • US12529163B2 patent drawing

AI summary

This disclosure provides methods and compositions for long fragment read sequencing. Technology is described for preparing long fragments of genomic DNA, for processing genomic DNA for long fragment read sequencing methods, as well as software and algorithms for processing and analyzing sequence data. Combinatorial oligonucleotide bar codes are used to label fragments from nearby portions of the genome, which facilitate computational assembly of sequence reads to obtain the genome sequence. This improves efficiency and accuracy of sequencing, whereby an entire sequence can be obtained from fragments that constitute a lower coverage amount of the genome.