Multimeric Barcoding Reagents for Sequencing Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sequencing machines face limitations due to finite raw readlengths and accuracy, and biophysical challenges from experimental DNA samples like FFPE, which hinder scientific and medical applications.

Innovation Solution

Multimeric barcoding reagents and methods for labeling and sequencing nucleic acids, enabling synthetic long reads and improved sequencing accuracy by linking barcode molecules to target nucleic acids, facilitating efficient sequencing of fragmented and damaged DNA.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If molecular barcoding is used to enable redundant sequencing and improve accuracy, then sequencing accuracy is improved, but device complexity and reagent complexity increase

Engineering Contradiction:
Improvesequencing accuracyVSAvoidbarcoding reagent complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The barcode molecule is divided into multiple segments (first barcode segment, second barcode segment, etc.) that can be independently synthesized and then assembled. This segmentation allows for modular design where each segment can be optimized separately, reducing the overall complexity of generating and managing complete barcode sequences while enabling redundant sequencing for improved accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Barcode molecules and their segments are pre-synthesized and prepared in advance before the actual sequencing experiment. This preliminary action allows for quality control and validation of barcode sequences beforehand, ensuring high accuracy in the final sequencing results without adding complexity during the actual sequencing process.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If molecular barcoding is used to enable digital molecular counting, then measurement precision is improved, but loss of information increases due to stochastic noise

Engineering Contradiction:
Improvemolecular counting accuracyVSAvoidstochastic sequence noise
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

Each target nucleic acid molecule is assigned a unique barcode copy that serves as its identifier. By creating multiple copies of the same target molecule with identical barcodes and sequencing them redundantly, the true signal can be distinguished from stochastic noise through consensus analysis, thereby improving molecular counting accuracy while filtering out random errors.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses the barcode information as feedback to identify and group redundant sequences originating from the same original molecule. This feedback mechanism allows the sequencing algorithm to focus on consistent signals across multiple reads while discarding stochastic variations, thus reducing information loss due to noise.

Inventive Principle:
Principle #23Feedback

3Length of moving object

If synthetic long reads are generated by assembling sub-sequences, then read length is improved, but device complexity increases

Engineering Contradiction:
Improveread lengthVSAvoidsequencing system complexity
Core Design Contradiction:
Length of moving objectVSDevice complexity

Solution Approach 1:

Multiple shorter sub-sequences are nested within a single synthetic long read by using a shared barcode molecule as the organizing framework. The barcode acts as a container that links together multiple sub-sequences, allowing them to be assembled into a longer contiguous sequence without requiring complex physical long-read sequencing technology.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The barcode molecule serves as an intermediary that connects multiple sub-sequences to each other and to the final assembled long read. This intermediary element simplifies the assembly process by providing clear linkage information, reducing the computational and experimental complexity compared to direct long-read sequencing approaches.

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

Enhances sequencing accuracy and ability to handle challenging samples by reducing stochastic noise and improving digital molecular counting, thereby overcoming limitations of existing sequencing technologies.

Implementation Method 1

The first barcoded oligonucleotide comprises a barcode region annealed to a barcode molecule

Methodology Applied
Scientific EffectAnnealing: Annealing

Implementation Method 2

barcode region annealed to the barcode molecule

Methodology Applied
Scientific EffectComplementary base pairing:

Implementation Method 3

a target region capable of annealing or ligating to a sub-sequence of the target nucleic acid

Methodology Applied
Scientific EffectAnnealing: Annealing

Implementation Method 4

a target region capable of annealing or ligating to a sub-sequence of the target nucleic acid

Methodology Applied
Scientific EffectLigation:

Data Source

PatentUS20220411786A1Reagents, kits and methods for molecular barcoding
Publication Date: 2022.12.29 CS GENETICS
  • US20220411786A1 patent drawing
  • US20220411786A1 patent drawing
  • US20220411786A1 patent drawing

AI summary

Multimeric barcoding reagents for labelling a target nucleic acid comprise: first and second barcode molecules linked together, wherein each of the barcode molecules comprises a nucleic acid sequence comprising a barcode region; and first and second barcoded oligonucleotides. The multimeric barcoding reagents enable spatial sequencing. A single multimeric barcoding reagent can be used to label sub-sequences of an intact nucleic acid molecule or co-localised fragments of a nucleic acid molecule. The labelled sub-sequences can be sequenced and the sequencing data processed to determine the sequence of sub-sequences from a single intact nucleic acid molecule or from co-localised fragments of a nucleic acid molecule. Corresponding libraries, kits, methods and uses are provided.