Identifier Sequence Design for Nanopore Sequencing Error Correction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current single-cell RNA sequencing methods face challenges in accurately identifying and correcting barcode and unique molecular identifier (UMI) sequences due to high error rates in long-read nanopore sequencing, which limits the adoption of long-read sequencing for single-cell analysis.

Innovation Solution

The method involves using identifier sequences built from pre-selected nucleotide blocks with mixtures of sequences differing by at least two nucleotide substitutions, allowing for error detection and correction in long-read single-cell sequencing without the need for parallel short-read sequencing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If long-read nanopore sequencing is used for single-cell RNA sequencing, then true end-to-end transcript sequencing and splicing information are obtained, but high error rates (5-15%) prevent accurate barcode and UMI sequence identification

Engineering Contradiction:
Improvesplicing informationVSAvoidbasecalling accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The identifier sequence is divided into multiple nucleotide blocks, where each block differs from others by at least two nucleotide substitutions. This segmentation allows individual blocks to be independently validated against the defined set, enabling error detection and correction while preserving the full identifier functionality for accurate cell and transcript identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A defined set of valid nucleotide blocks is predetermined and established before sequencing. During data analysis, each nucleotide block from the sequenced identifier is compared against this pre-defined set to identify and correct sequencing errors, rather than attempting error correction after the fact.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If parallel short-read Illumina sequencing is used to error correct long-read Nanopore sequencing, then assignment rates increase from 6% to 67%, but time and cost increase due to constructing and sequencing two separate libraries

Engineering Contradiction:
Improveassignment rateVSAvoidsequencing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs self-validation by comparing each nucleotide block against the pre-defined set of valid blocks. This self-service error detection mechanism eliminates the need for external error correction using parallel short-read sequencing, thereby reducing both time and cost while maintaining high assignment rates.

Inventive Principle:
Principle #25Self-service

3Productivity

If pre-synthesised nucleotide blocks with at least two nucleotide substitutions between each block are used, then error detection and correction is enabled without parallel sequencing, but the complexity of designing and synthesising the identifier sequences increases

Engineering Contradiction:
ImproveyieldVSAvoididentifier sequence design
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The identifier sequence is segmented into multiple nucleotide blocks, each drawn from a defined set where every block differs from others by at least two nucleotide substitutions. This segmentation structure enables error detection while distributing the design complexity across manageable blocks rather than requiring complex overall sequence design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The design specifies a minimum parameter (at least two nucleotide substitutions) between each nucleotide block in the defined set. This parameter change ensures sufficient sequence divergence for error detection while maintaining a manageable number of possible blocks, balancing design complexity with error correction capability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240102091A1oligonucleotides
Publication Date: 2024.03.28 OXFORD UNIVERSITY INNOVATION LTD
  • US20240102091A1 patent drawing
  • US20240102091A1 patent drawing
  • US20240102091A1 patent drawing

AI summary

The invention relates to methods of adding identifier sequences to polynucleotides of an array. The identifier sequences comprise a plurality of nucleotide blocks. Also provided are arrays of polynucleotides having identifier sequences, microparticles comprising said arrays, a plurality of 5 said microparticles, surfaces comprising said arrays, kits and methods for generating libraries using the array, methods for determining the accuracy of sequencing or amplification an array, and methods of analysing said libraries.