Identifier Sequence Design for Nanopore Sequencing Error Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current single-cell RNA sequencing methods face challenges in accurately identifying and correcting barcode and unique molecular identifier (UMI) sequences due to high error rates in long-read nanopore sequencing, which limits the adoption of long-read sequencing for single-cell analysis.
Innovation Solution
The method involves using identifier sequences built from pre-selected nucleotide blocks with mixtures of sequences differing by at least two nucleotide substitutions, allowing for error detection and correction in long-read single-cell sequencing without the need for parallel short-read sequencing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If long-read nanopore sequencing is used for single-cell RNA sequencing, then true end-to-end transcript sequencing and splicing information are obtained, but high error rates (5-15%) prevent accurate barcode and UMI sequence identification
Solution Approach 1:
The identifier sequence is divided into multiple nucleotide blocks, where each block differs from others by at least two nucleotide substitutions. This segmentation allows individual blocks to be independently validated against the defined set, enabling error detection and correction while preserving the full identifier functionality for accurate cell and transcript identification.
Solution Approach 2:
A defined set of valid nucleotide blocks is predetermined and established before sequencing. During data analysis, each nucleotide block from the sequenced identifier is compared against this pre-defined set to identify and correct sequencing errors, rather than attempting error correction after the fact.
2Reliability
If parallel short-read Illumina sequencing is used to error correct long-read Nanopore sequencing, then assignment rates increase from 6% to 67%, but time and cost increase due to constructing and sequencing two separate libraries
Solution Approach 1:
The system performs self-validation by comparing each nucleotide block against the pre-defined set of valid blocks. This self-service error detection mechanism eliminates the need for external error correction using parallel short-read sequencing, thereby reducing both time and cost while maintaining high assignment rates.
3Productivity
If pre-synthesised nucleotide blocks with at least two nucleotide substitutions between each block are used, then error detection and correction is enabled without parallel sequencing, but the complexity of designing and synthesising the identifier sequences increases
Solution Approach 1:
The identifier sequence is segmented into multiple nucleotide blocks, each drawn from a defined set where every block differs from others by at least two nucleotide substitutions. This segmentation structure enables error detection while distributing the design complexity across manageable blocks rather than requiring complex overall sequence design.
Solution Approach 2:
The design specifies a minimum parameter (at least two nucleotide substitutions) between each nucleotide block in the defined set. This parameter change ensures sufficient sequence divergence for error detection while maintaining a manageable number of possible blocks, balancing design complexity with error correction capability.
Data Source
AI summary
The invention relates to methods of adding identifier sequences to polynucleotides of an array. The identifier sequences comprise a plurality of nucleotide blocks. Also provided are arrays of polynucleotides having identifier sequences, microparticles comprising said arrays, a plurality of 5 said microparticles, surfaces comprising said arrays, kits and methods for generating libraries using the array, methods for determining the accuracy of sequencing or amplification an array, and methods of analysing said libraries.


