Single cell RNA sequencing methods and systems

The method corrects barcode and probe errors in nanopore sequencing data to achieve accurate single cell RNA analysis, addressing the incompatibility of commercial kits with non-Illumina sequencing data and improving sequencing throughput.

WO2026060095A1PCT designated stage Publication Date: 2026-03-19ROCHE SEQUENCING SOLUTIONS INC
View PDF 16 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing commercial RNA sequencing kits optimized for Illumina sequencing struggle with non-Illumina sequencing data, leading to poor results due to differences in accuracy and error profiles, hindering the effective use of alternative sequencing techniques like nanopore sequencing for single cell RNA analysis.

Method used

A method for single cell RNA sequencing using nanopore sequencers involves capturing mRNA, creating cDNA with barcodes and UMIs, amplifying, synthesizing Xpandomer molecules, and correcting barcode and probe errors in the sequencing data using bioinformatics pipelines, allowing for accurate analysis.

Benefits of technology

The method enables accurate and reliable single cell RNA sequencing with nanopore sequencers, producing results comparable to Illumina sequencing, thereby enhancing sequencing throughput and resolving subpopulations and cell-cell interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025045917_19032026_PF_FP_ABST
    Figure US2025045917_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A method includes obtaining data representing molecular bases of DNA or RNA and identifying and correcting each of one base pair substitution, insertion, or deletion errors within the sequencing data pertaining to the at least one barcode, UMI, or runway of molecular bases leading to a UMI. The method includes validating the sequencing data for bioinformatics analysis based on the errors identified within the sequencing data pertaining to the at least one barcode, UMI, or runway leading to a UMI, wherein the validating applies a greater threshold for an error within a portion of the sequence data including a BMI, barcode, or runway compared to portions of the sequencing data that does not include a BMI, barcode, or runway respectively.
Need to check novelty before this filing date? Find Prior Art

Description

SINGLE CELL RNA SEQUENCING METHODS AND SYSTEMSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims benefit of priority to U.S. Provisional Application No. 63 / 694925, filed September 16, 2024, the entire contents of which is herein incorporated by reference.INCORPORATION BY REFERENCE

[0002] All publications and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.FIELD

[0003] Embodiments of the disclosure relate generally to systems and methods for RNA sequencing, and more particularly, to systems and methods for RNA sequencing using a nanopore sequencer.BACKGROUND OF THE INVENTION

[0004] RNA sequencing of a sample facilitates an analysis of the gene expression profile, i.e. the transcriptome, of a group of cells in the sample. This technique has been extended to single cell RNA sequencing (scRNA-Seq), which allows the transcriptome of single cells to be determined. This information can be used in a variety of applications, such as for cell type identification, which can be used to design personalized treatments for tumors or cancers.

[0005] Multiple systems and kits are commercially available to perform scRNA-Seq (e.g., lOx Genomics and Parse) and spatial transcriptome mapping at single cell resolution (e.g., Curio Biosciences). These commercial kits generally utilize in their workflows Illumina sequencers to generate the RNA sequencing data.

[0006] However, new next-generation sequencing (NGS) techniques, such as nanopore sequencing, are being developed as alternatives to Illumina based sequencing in order to provide lower costs and higher throughput. Researchers continuously seek to sequence more reads per cell and more cells per experiment to better resolve subpopulations and discover biological insights. Furthermore, spatial transcriptomics provides spatial context of gene expression within tissues, and gains in sequencing throughput per spatial barcode enable higher resolution of colocalization patterns and cell-cell interactions.

[0007] The sequencing data generated by these alternative sequencing techniques can differ from Illumina sequencing data, for example with different accuracy and error profdes. Consequently, use of the non-Illumina sequencing data in the commercial kits that have been optimized for use with Illumina sequencing data can result in relatively poor results.

[0008] Therefore, it would desirable to provide systems and methods for modifying and / or adapting the methods and protocols of the alternative sequencing techniques to work with existing commercial RNA sequencing kits.SUMMARY OF THE INVENTION

[0009] The present invention relates generally to systems and methods for RNA sequencing, and more particularly, to systems and methods for RNA sequencing using a nanopore sequencer.

[0010] In some embodiments, a method for performing single cell RNA analysis is provided. The method includes receiving a sample of mRNA; capturing on a solid surface the mRNA; creating cDNA from the captured mRNA using reverse transcriptase, wherein the cDNA comprises a barcode, a runway in front of the barcode, and a unique molecular identifier (UMI); releasing the cDNA from the solid surface; amplifying the cDNA; synthesizing Xpandomer molecules from the cDNA; sequencing the Xpandomer molecules using a nanopore sequencer; writing the sequencing data into a FASTQ file; identify barcode errors, wherein the barcode errors include 1 base pair substitution errors and 1 bp insertion and deletion errors; correct 1 bp barcode errors in the sequencing data; and analyze the corrected sequencing data with a bioinformatics pipeline.

[0011] In some embodiments, the method further includes identifying UMIs having base qualities less than 10; and adjusting the UMI base qualities that are below 10 to 10.

[0012] In some embodiments, the runway is less than or equal to two base pairs in length.

[0013] In some embodiments, a method for performing single cell RNA analysis is provided. The method includes receiving a sample of single cells having mRNA; hybridizing the mRNA in the single cells to probes; partitioning the single cells in a gel bead in emulsion; extending the probes, wherein the extended probes comprise a probe barcode, a cell barcode, and UMI; amplifying the extended probes; synthesizing Xpandomer molecules from the amplified probes; sequencing the Xpandomer molecules using a nanopore sequencer; writing the sequencing data into a FASTQ file; correcting probe barcodes in the sequencing data; and analyze the corrected sequencing data with a bioinformatics pipeline.

[0014] In some embodiments, correcting probe barcodes includes allowing the probe barcodes to be 1-2 base pair edit distance away from a known list of valid probe barcodes.

[0015] In some embodiments, the method further includes correcting cell barcodes in the sequencing data by allowing and correcting cell barcodes with a 1 bp Indel.SUMMARY OF THE INVENTION

[0016] The novel features of the invention are set forth with particularity in the claims that follow. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:

[0017] FIG. 1 includes a top view of an embodiment of a nanopore sensor device having an array of nanopore cells.

[0018] FIG. 2 illustrates the elements of a nanopore sensor device.

[0019] FIGS. 3A-3B depict an embodiment of a device performing nucleotide sequencing with the nanopore -enabled SBS technique (panel 3A), and an embodiment of a cell performing sequencing by expansion (SBX) (panel 3B), respectively.

[0020] FIG. 4 illustrates an embodiment of the workflow for the 10X single cell 3’ gene expression kit.

[0021] FIG. 5 illustrates a portion of an embodiment of the bioinformatics workflow for the 10X single cell 3’ gene expression kit.

[0022] FIG. 6 illustrates an embodiment of a read without a runway and a read with a runway.

[0023] FIG. 7 illustrates an embodiment of the Parse Evercode Workflow for both SBX and Illumina sequencing platforms.

[0024] FIG. 8 illustrates an embodiment of the procedure for constructing and correcting reads in the Parse workflow.

[0025] FIG. 9 illustrates an embodiment of the workflow for the 10X Chromium Single Cell Gene Expression Flex Kit.

[0026] FIG. 10 illustrates an embodiment of a portion of the bioinformatics workflow for the 10X Chromium Single Cell Gene Expression Flex Kit.

[0027] FIG. 11 illustrates an embodiment of the Curio Seeker workflow.

[0028] FIG. 12 illustrates an embodiment of the Visium HD Probe Capture workflow the Visium HD Probe Capture workflow.

[0029] FIG. 13 illustrates embodiments of different types of reads generated by the different scRNA kits described herein.

[0030] FIG. 14 illustrates that using SBX sequencing reads with the Parse kits results in similar predicted cell types as using Illumina sequencing reads with the Parse kits.

[0031] FIG. 15 illustrates that use of SBX sequencing reads with the Curio kit and platform results in similar spatial mapping of predicted clusters as use of Illumina sequencing reads with the Curio Kit and platform.

[0032] FIG. 16 illustrates an embodiment of a computer system that can be used to execute the bioinformatics piplelines described herein.DETAILED DESCRIPTION OF THE INVENTION

[0033] A. Definitions

[0034] Unless defined otherwise, technical and scientific terms used herein have the same meaning as commonly understood by a person of ordinary skill in the art. Methods, devices, and materials similar or equivalent to those described herein can be used in the practice of disclosed techniques. The following terms are provided to facilitate understanding of certain terms used frequently and are not meant to limit the scope of the present disclosure. Abbreviations used herein have their conventional meaning within the chemical and biological arts.

[0035] When a feature or element is herein referred to as being “on” another feature or element, it can be directly on the other feature or element or intervening features and / or elements may also be present. In contrast, when a feature or element is referred to as being “directly on” another feature or element, there are no intervening features or elements present. It will also be understood that, when a feature or element is referred to as being “connected," “attached” or “coupled” to another feature or element, it can be directly connected, attached or coupled to the other feature or element or intervening features or elements may be present. In contrast, when a feature or element is referred to as being “directly connected," “directly attached” or “directly coupled” to another feature or element, there are no intervening features or elements present. Although described or shown with respect to one embodiment, the features and elements so described or shown can apply to other embodiments. It will also be appreciated by those of skill in the art that references to a structure or feature that is disposed “adjacent” another feature may have portions that overlap or underlie the adjacent feature.

[0036] Terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. For example, as used herein, the singular forms “a," “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items and may be abbreviated as “ / ."

[0037] Spatially relative terms, such as “under," “below," “lower," “over," “upper” and the like, may be used herein for ease of description to describe one element or feature’s relationship to another element(s) or feature(s) as illustrated in the figures. It will be understood that the spatially relative terms are intended to encompass different orientations of the device in use or operation in addition to the orientation depicted in the figures. For example, if a device in the figures is inverted, elements described as “under” or “beneath” other elements or features would then be oriented “over” the other elements or features. Thus, the exemplary term “under” can encompass both an orientation of over and under. The device may be otherwise oriented (rotated 90 degrees or at other orientations) and the spatially relative descriptors used herein interpreted accordingly. Similarly, the terms “upwardly," “downwardly," “vertical," “horizontal” and the like are used herein for the purpose of explanation only unless specifically indicated otherwise.

[0038] Although the terms “first” and “second” may be used herein to describe various features / elements (including steps), these features / elements should not be limited by these terms, unless the context indicates otherwise. These terms may be used to distinguish one feature / element from another feature / element. Thus, a first feature / element discussed below could be termed a second feature / element, and similarly, a second feature / element discussed below could be termed a first feature / element without departing from the teachings of the present disclosure.

[0039] Throughout this specification and the claims which follow, unless the context requires otherwise, the word “comprise," and variations such as “comprises” and “comprising” means various components can be co-jointly employed in the methods and articles (e.g., compositions and apparatuses including device and methods). For example, the term “comprising” will be understood to imply the inclusion of any stated elements or steps but not the exclusion of any other elements or steps.

[0040] As used herein in the specification and claims, including as used in the examples and unless otherwise expressly specified, all numbers may be read as if prefaced by the word “about” or “approximately,” even if the term does not expressly appear. The phrase “about” or “approximately” may be used when describing magnitude and / or position to indicate that the value and / or position described is within a reasonable expected range of values and / or positions. For example, a numeric value may have a value that is + / - 0.1% of the stated value (or range of values), + / - 1% of the stated value (or range of values), + / - 2% of the stated value (or range of values), + / - 5% of the stated value (or range of values), + / - 10% of the stated value (or range of values), etc. Any numerical values given herein should also be understood to include about or approximately that value, unless the context indicates otherwise. For example, if the value “10” is disclosed, then “about 10” is also disclosed. Any numerical range recited herein is intended to include all sub-ranges subsumed therein. It is also understood that when a value is disclosed that “less than or equal to” the value, “greater than or equal to the value” and possible ranges between values are also disclosed, as appropriately understood by the skilled artisan. For example, if the value “X” is disclosed the “less than or equal to X” as well as “greater than or equal to X” (e.g., where X is a numerical value) is also disclosed. It is also understood that the throughout the application, data is provided in a number of different formats, and that this data, represents endpoints and starting points, and ranges for any combination of the data points. For example, if a particular data point “10” and a particular data point “15” are disclosed, it is understood that greater than, greater than or equal to, less than, less than or equal to, and equal to 10 and 15 are considered disclosed as well as between 10 and 15. It is also understood that each unit between two particular units are also disclosed. For example, if 10 and 15 are disclosed, then 11, 12, 13, and 14 are also disclosed.

[0041] For the descriptions herein and the appended claims, the singular forms “a,” and “an” include plural referents unless the context clearly indicates otherwise. Thus, for example, reference to “a protein” includes more than one protein, and reference to “a compound” refers to more than one compound. The use of “comprise,” “comprises,” “comprising” “include,” “includes,” and “including” are interchangeable and not intended to be limiting. It is to be further understood that where descriptions of various embodiments use the term “comprising,” those skilled in the art would understand that in some specific instances, an embodiment can be alternatively described using language “consisting essentially of’ or “consisting of.”

[0042] Where a range of values is provided, unless the context clearly dictates otherwise, it is understood that each intervening integer of the value, and each tenth of each intervening integer of the value, unless the context clearly dictates otherwise, between the upper and lower limit of that range, and any other stated or intervening value in that stated range, is encompassed within the disclosure. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges, and are also encompassed within the disclosure, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding (i) either or (ii) both of those included limits are also included in the disclosure. For example, “1 to 50” includes “2 to 25,” “5 to 20,” “25 to 50,” “1 to 10,” etc.

[0043] Generally, the nomenclature used herein and the techniques and procedures described herein include those that are well understood and commonly employed by those of ordinary skill in the art, such as the common techniques and methodologies described in e.g., Green and Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Vols. 1-3, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y., 2012 (hereinafter “Sambrook”); and Current Protocols in Molecular Biology, F. M. Ausubel et al., eds., originally published in 1987 in book form by Greene Publishing Associates, Inc. and John Wiley & Sons, Inc., and regularly supplemented through 2011, and now available in journal format online as Current Protocols in Molecular Biology, Vols. 00 - 130, (1987-2020), published by Wiley & Sons, Inc. in the Wiley Online Library(hereinafter “Ausubel”).

[0044] A "nanopore" refers to a pore, channel or passage formed or otherwise provided in a membrane. A membrane can be an organic membrane, such as a lipid bilayer, or a synthetic membrane, such as a membrane formed of a polymeric material. The nanopore can be disposed adjacent or in proximity to a sensing circuit or an electrode coupled to a sensing circuit, such as, for example, a complementary metal oxide semiconductor (CMOS) or field effect transistor (FET) circuit. In some examples, a nanopore has a characteristic width or diameter on the order of 0.1 nanometers (nm) to about 1000 nm. In some implementations, a nanopore may be a protein.

[0045] A “nucleic acid” refers to deoxyribonucleotides or ribonucleotides and polymers thereof in either single- or double-stranded form. The term encompasses nucleic acids containing known nucleotide analogs or modified backbone residues or linkages, which are synthetic, naturally occurring, and non-naturally occurring, which have similar binding properties as the reference nucleic acid, and which are metabolized in a manner similar to the reference nucleotides. Examples of such analogs include, without limitation,phosphorothioates, phosphoramidites, methyl phosphonates, chiral-methyl phosphonates, 2- O-methyl ribonucleotides, and peptide-nucleic acids (PNAs). Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions) and complementary sequences, as well as the sequence explicitly indicated. Specifically, degenerate codon substitutions can be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)). The term nucleic acid can be used interchangeably with gene, cDNA, mRNA, oligonucleotide, and polynucleotide.

[0046] The term “nucleotide,” in addition to referring to the naturally occurring ribonucleotide or deoxyribonucleotide monomers, can be understood to refer to related structural variants thereof, including derivatives and analogs, that are functionally equivalent with respect to the particular context in which the nucleotide is being used (e.g., hybridization to a complementary base), unless the context clearly indicates otherwise.

[0047] The term "tag" refers to a detectable moiety that can be atoms or molecules, or a collection of atoms or molecules. A tag can provide an optical, electrochemical, magnetic, or electrostatic (e.g., inductive, capacitive) signature, which signature can be detected with the aid of a nanopore. Typically, when a nucleotide is attached to the tag it is called a "tagged nucleotide." The tag can be attached to the nucleotide via the phosphate moiety.

[0048] The term “template” refers to a single stranded nucleic acid molecule that is copied into a complementary strand of DNA nucleotides for DNA synthesis. In some cases, a template can refer to the sequence of DNA that is copied during the synthesis of mRNA.

[0049] “Xpandomer” or “XP” refers to a polymer synthesized by transcription of the sequence of a nucleic acid template. The transcribed sequence is encoded along the XP backbone in high signal-to-noise reporters that are separated by ~10 nm and are designed for high-signal -to-noise, well-differentiated responses. These differences provide significant performance enhancements in sequence read efficiency and accuracy of XPs relative to natural DNA. XPs are used in to carry out sequencing- by-expansion (“SBX”) and the building blocks of XPs are XNTPs, which are non-natural nucleotide analogs used in XP synthesis to transcribe the sequence of a nucleic acid template. XNTPs are expandable, 5' triphosphate modified non-natural nucleotide analogs compatible with template dependent enzymatic polymerization.

[0050] The term “signal value” refers to a value of the sequencing signal output from a sequencing cell. According to certain embodiments, the sequencing signal is an electrical signal that is measured and / or output from a point in a circuit of one or more sequencing cells e.g., the signal value is (or represents) a voltage or a current. The signal value can represent the results of a direct measurement of voltage and / or current and / or may represent an indirect measurement, e.g., the signal value can be a measured duration of time for which it takes a voltage or current to reach a specified value. A signal value can represent any measurable quantity that correlates with the resistivity of a nanopore and from which the resistivity and / or conductance of the nanopore (threaded and / or unthreaded) can be derived. As another example, the signal value can correspond to a light intensity, e.g., from a fluorophore attached to a nucleotide being added to a nucleic acid with a polymerase.

[0051] The term “osmolarity,” also known as osmotic concentration, refers to a measure of solute concentration. Osmolarity measures the number of osmoles of solute particles per unit volume of solution. An osmole is a measure of the number of moles of solute that contribute to the osmotic pressure of a solution. Osmolarity allows the measurement of the osmotic pressure of a solution and the determination of how the solvent will diffuse across a semipermeable membrane (osmosis) separating two solutions of different osmotic concentration. The term “osmolyte” refers to any soluble compound that when dissolved into a solution increases the osmolarity of that solution.

[0052] Exemplary nanopore systems, circuitry, and sequencing operations are described below, as well as methods of using such systems and components in a sequencing workflow. Embodiments of the disclosure can be implemented in numerous ways, including as a process, a system, and a computer program product embodied on a computer readable storage medium and / or a processor, such as a processor configured to execute instructions stored on and / or provided by a memory coupled to the processor.

[0053] B. Nanopore-based devices

[0054] Nanopore-based devices for detecting nucleic acids have been developed for rapid sequencing and various designs and methods of use are known in the art. See e.g., US9494554B2, US9567630B2, US9557294B2, US9605309B2, each of which hereby incorporated by reference herein. These devices generally include a plurality of electrochemical cells and each electrochemical cell is in fluid communication with the cells in the array (or plurality) and comprises a chamber containing a nanopore embedded in a membrane. The membrane acts to separate the cell chamber into two sub-chambers, referred to as the cis and trans sides of the cell, each of which contain an electrode.

[0055] Within the electrochemical cell, the nanopore embedded in the membrane is disposed in proximity to an electrode coupled to a sensing circuit, such as, for example, a complementary metal -oxide semiconductor (CMOS) or field effect transistor (FET) circuit. When a voltage potential is applied (via the electrodes) across a nanopore immersed in a conducting fluid, a small current attributed to the flow of ions through the nanopore can be observed. This ion flow is sensitive to the pore size, and thus, molecules entering the pore affect the ion flow and the voltage measured through this sensor circuit.

[0056] Electrochemical cells for nanopore-based sequencing of nucleic acids are typically used in a massively parallel fashion in which thousands of such cells are configured as an array in a single device often referred to as a chip (or device). A typically nanopore -based sequencing device device incorporates an array of one million or more electrochemical cells, and may include 1000 rows by 1000 columns of such cells (see e.g., devices fabricated by Roche Sequencing Solutions, Santa Clara, CA, USA). Methods for fabricating and using such nanopore array devices can also be found in U.S. Patent Application Publication Nos. 2013 / 0244340 Al, US 2013 / 0264207 Al, US2014 / 0134616 Al, 2015 / 0368710 Al, and 2018 / 0057870 Al, and published International Application WO 2019 / 166457 Al, each of which is hereby incorporated by reference herein. Each well in the array is manufactured using a semiconductor manufacturing process that provides surface modifications that allow for constant contact with biological reagents and conductive salts. Each well can support a phospholipid bilayer membrane with a nanopore -polymerase conjugate embedded therein. The electrode at each well is individually addressable by computer interface. All reagents used are introduced into a simple flow cell above the array device using a computer- controlled syringe pump. The device supports analog to digital conversion and reports electrical measurements from all electrodes independently at a rate of over 200 points per second, e.g., over 500 points per second, and in a specific embodiment, over 1000 points per second. Nanopore measurements can be made asynchronously at each of 8 M addressable nanopore -containing membranes in the array at least once every millisecond (msec) and recorded on the interfaced computer. Further description of exemplary electrochemical cells useful for nanopore-based nucleic acid assays, such as sequencing, including chamber and electrode materials, buffer solutions, sensing circuitry, array devices, and their use in various applications is provided below.

[0057] FIG. 1 is a top view of an embodiment of a nanopore sensor device 100 having an array 140 of nanopore cells 150. Each nanopore cell 150 includes a control circuit integrated on a silicon substrate of nanopore sensor device 100. In some embodiments, side walls 136are included in array 140 to separate groups of nanopore cells 150 so that each group can receive a different sample for characterization. Each nanopore cell can be used to sequence a nucleic acid. In some embodiments, nanopore sensor device 100 includes a cover plate 130. In some embodiments, nanopore sensor device 100 also includes a plurality of electrical contacts 110 (e.g., pins, wires, or solder bumps) for interfacing with other circuits, such as a computer processor.

[0058] In some embodiments, nanopore sensor device 100 includes multiple devices in the same package, such as, for example, a Multi -device Module or System -in-Package. The devices can include, for example, a memory, a processor, a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), data converters, a high-speed I / O interface, etc.

[0059] In some embodiments, nanopore sensor device 100 is coupled to (e.g., docked to) a workstation 120, which can include various components for carrying out (e.g., automatically carrying out) various embodiments of the processes disclosed herein. These processes can include, for example, analyte delivery mechanisms, such as pipettes for delivering lipid suspension or other membrane structure suspension, analyte solution, and / or other liquids, suspension or solids. The workstation components can further include robotic arms, one or more computer processors, and / or memory. A plurality of analytes of interest can be detected on array 140 of nanopore cells 150. In some embodiments, each nanopore cell 150 is individually addressable.

[0060] Nanopore cells 150 in nanopore sensor device 100 can be implemented in many different ways. For example, in some embodiments, tags of different sizes and / or chemical structures are attached to different nucleotides in a nucleic acid molecule to be sequenced. In some embodiments, a complementary strand to a template of the nucleic acid molecule to be sequenced may be synthesized by hybridizing differently polymer-tagged nucleotides with the template. In some implementations, the nucleic acid molecule and the attached tags both move through the nanopore, and an ion current passing through the nanopore can indicate the nucleotide that is in the nanopore because of the particular size and / or structure of the tag attached to the nucleotide. In some implementations, only the tags are moved into the nanopore. There can also be many different ways to detect the different tags in the nanopores.

[0061] In a specific embodiment, the nanopore cells are used for SBX protocols and the compositions and methods described herein are optimized for SBX chemistry, as described in more detail herein.

[0062] FIG. 2 illustrates an embodiment of an example nanopore cell 200 in a nanopore sensor device, such as nanopore cell 150 in nanopore sensor device 100 of FIG. 1, that can be used to characterize a xpandomer, polynucleotide or a polypeptide. Nanopore cell 200 can include a well 205 formed of dielectric layers 201 and 204; a membrane, such as a lipid bilayer 214 formed over well 205; and a sample chamber 215 on lipid bilayer 214 and separated from well 205 by lipid bilayer 214. Well 205 can contain a volume of electrolyte 206 containing a nanopore, and sample chamber 215 can hold bulk electrolyte 208 containing a nanopore, e.g., a soluble protein nanopore transmembrane molecular complexes (PNTMC), and the analyte of interest (e.g., a nucleic acid molecule to be sequenced).

[0063] Nanopore cell 200 can include a working electrode 202 at the bottom of well 205 and a counter electrode 210 disposed in sample chamber 215. A signal source 228 can apply a voltage signal between working electrode 202 and counter electrode 210. A single nanopore (e.g., a PNTMC) can be inserted into lipid bilayer 214 by an electroporation process caused by the voltage signal, thereby forming a nanopore 216 in lipid bilayer 214. The individual membranes (e.g., lipid bilayers 214 or other membrane structures) in the array can be neither chemically nor electrically connected to each other. Thus, each nanopore cell in the array can be an independent sequencing machine, producing data unique to the single polymer molecule associated with the nanopore that operates on the analyte of interest and modulates the conductance of the otherwise low conductivity lipid bilayer.

[0064] As shown in FIG. 2, nanopore cell 200 can be formed on a substrate 230, such as a silicon substrate. Dielectric layer 201 can be formed on substrate 230. Dielectric material used to form dielectric layer 201 can include, for example, glass, oxides, nitrides, and the like. An electric circuit 222 for controlling electrical stimulation and for processing the signal detected from nanopore cell 200 can be formed on substrate 230 and / or within dielectric layer 201. For example, a plurality of patterned metal layers (e.g., metal 1 to metal 6) can be formed in dielectric layer 201, and a plurality of active devices (e.g., transistors) can be fabricated on substrate 230. In some embodiments, signal source 228 is included as a part of electric circuit 222. Electric circuit 222 can include, for example, amplifiers, integrators, analog-to-digital converters, noise filters, feedback control logic, and / or various other components. Electric circuit 222 can be further coupled to a processor 224 that is coupled to a memory 226, where processor 224 can analyze the sequencing data to determine sequences of the polymer molecules that have been sequenced in the array.

[0065] Working electrode 202 can be formed on dielectric layer 201, and can form at least a part of the bottom of well 205. In some embodiments, working electrode 202 is a metalelectrode. For non-faradaic conduction, working electrode 202 can be made of metals or other materials that are resistant to corrosion and oxidation, such as, for example, platinum, gold, titanium nitride, and graphite. For example, working electrode 202 can be a platinum electrode with electroplated platinum. In another example, working electrode 202 can be a titanium nitride (TiN) working electrode. Working electrode 202 can be porous, thereby increasing its surface area and a resulting capacitance associated with working electrode 202. Because the working electrode of a nanopore cell can be independent from the working electrode of another nanopore cell, the working electrode can be referred to as cell electrode in this disclosure.

[0066] Dielectric layer 204 can be formed above dielectric layer 201. Dielectric layer 204 forms the walls surrounding well 205. Dielectric material used to form dielectric layer 204 can include, for example, glass, oxide, silicon mononitride (SiN), polyimide, or other suitable hydrophobic insulating material. The top surface of dielectric layer 204 can be silanized. The silanization can form a hydrophobic layer 220 above the top surface of dielectric layer 204. In some embodiments, hydrophobic layer 220 has a thickness of about 1.5 nanometer (nm).

[0067] Well 205 formed by the dielectric layer walls 204 includes volume of electrolyte 206 above working electrode 202. Volume of electrolyte 206 can be buffered and can include one or more of the following: ammonium chloride, lithium chloride (LiCl), sodium chloride (NaCl), potassium chloride (KC1), lithium glutamate, sodium glutamate, potassium glutamate, lithium acetate, sodium acetate, potassium acetate, calcium chloride (CaC12), strontium chloride (SrC12), manganese chloride (MnC12), and magnesium chloride (MgC12). In some embodiments, volume of electrolyte 206 has a thickness of about two microns (pm).

[0068] As also shown in FIG. 2, a membrane can be formed on top of dielectric layer 204 and spanning across well 205. In some embodiments, the membrane includes a lipid monolayer 218 formed on top of hydrophobic layer 220. As the membrane reaches the opening of well 205, lipid monolayer 208 can transition to lipid bilayer 214 that spans across the opening of well 205.

[0069] As shown, lipid bilayer 214 is embedded with a single nanopore 216, e.g., formed by a single PNTMC. As described above, nanopore 216 can be formed by inserting a single PNTMC into lipid bilayer 214 by electroporation. Nanopore 216 can be large enough for passing at least a portion of the analyte of interest and / or small ions (e.g., Na+, K+, Ca2+, CI- ) between the two sides of lipid bilayer 214.

[0070] Sample chamber 215 is over lipid bilayer 214, and can hold a solution of the analyte of interest for characterization. The solution can be an aqueous solution containing bulkelectrolyte 208 and buffered to an optimum ion concentration and maintained at an optimum pH to keep the nanopore 216 open. Nanopore 216 crosses lipid bilayer 214 and provides the only path for ionic flow from bulk electrolyte 208 to working electrode 202. In addition to nanopores (e.g., PNTMCs) and the analyte of interest, bulk electrolyte 208 can further include one or more of the following: ammonium chloride, lithium chloride (Li Cl), sodium chloride (NaCl), potassium chloride (KC1), lithium glutamate, sodium glutamate, potassium glutamate, lithium acetate, sodium acetate, potassium acetate, calcium chloride (CaC12), strontium chloride (SrC12), manganese chloride (MnC12), and magnesium chloride (MgC12).

[0071] Counter electrode (CE) 210 can be an electrochemical potential sensor. In some embodiments, counter electrode 210 is shared between a plurality of nanopore cells, and can therefore be referred to as a common electrode. In some cases, the common potential and the common electrode can be common to all nanopore cells, or at least all nanopore cells within a particular grouping. The common electrode can be configured to apply a common potential to the bulk electrolyte 208 in contact with the nanopore 216. Counter electrode 210 and working electrode 202 can be coupled to signal source 228 for providing electrical stimulus (e.g., voltage bias) across lipid bilayer 214, and can be used for sensing electrical characteristics of lipid bilayer 214 (e.g., resistance, capacitance, and ionic current flow). In some embodiments, nanopore cell 200 can also include a reference electrode 212.

[0072] In some embodiments, various checks are made during creation of the nanopore cell as part of calibration. Once a nanopore cell is created, further calibration steps can be performed, e.g., to identify nanopore cells that are performing as desired (e.g., one nanopore in the cell). Such calibration checks can include physical checks, voltage calibration, open channel calibration, and identification of cells with a single nanopore.

[0073] Nanopore cells, e.g., cells 150 in nanopore sensor device 100, can enable parallel sequencing using a single molecule nanopore based sequencing by synthesis (Nano-SBS) technique, described below.

[0074] C. Exemplary Nanopore Sequencing Chemistry

[0075] Nanopore-based sequencing-by-synthesis (“SBS”) uses a polymerase (or other strandextending enzyme) covalently linked to a nanopore to synthesize a DNA strand complementary to a target sequence template (i.e., a copy strand). The nanopore embedded in a membrane in an electrochemical cell is used to concurrently detect the identity of each nucleotide monomer as it is added to that growing strand. See e.g., US Pat. Publ. Nos. 2013 / 0244340 Al, 2013 / 0264207 Al, 2014 / 0134616 Al, 2015 / 0368710 Al, and 2018 / 0057870 Al, and published International Application WO 2019 / 166457 Al. Eachadded nucleotide monomer is detected by monitoring signals due to changes in ion flow through the nanopore as a tag moiety attached to each added nucleotide monomer enters the nanopore and alter the ion flow. For optimal performance, the tag moiety should reside in the nanopore for a sufficient amount of time to provide for a detectable, identifiable, and reproducible signal associated with altering ion flow through the nanopore (relative to the baseline “open current” flow), such that the specific nucleotide associated with the tag can be distinguished unambiguously from the other tagged nucleotides in the SBS solution.

[0076] FIG. 3A illustrates an embodiment of a nanopore cell 300 performing nucleotide sequencing using the Nano-SBS technique. In the Nano-SBS technique, a template 332 to be sequenced (e.g., a nucleotide acid molecule or another analyte of interest) and a primer can be introduced into bulk electrolyte 308 in the sample chamber of nanopore cell 300. As examples, template 332 can be circular or linear. A nucleic acid primer can be hybridized to a portion of template 332 to which four differently polymer-tagged nucleotides 338 can be added.

[0077] In some embodiments, an enzyme (e.g., a polymerase 334, such as a DNA polymerase) is associated with nanopore 316 for use in synthesizing a complementary strand to template 332. For example, polymerase 334 can be covalently attached to nanopore 316. Polymerase 334 can catalyze the incorporation of nucleotides 338 onto the primer using a single stranded nucleic acid molecule as the template. Nucleotides 338 can comprise tag species (“tags”) with the nucleotide being one of four different types: A, T, G, or C. When a tagged nucleotide is correctly complexed with polymerase 334, the tag can be pulled (e.g., loaded) into the nanopore by an electrical force, such as a force generated in the presence of an electric field generated by a voltage applied across lipid bilayer 314 and / or nanopore 316. The tail of the tag can be positioned in the barrel of nanopore 316. The tag held in the barrel of nanopore 316 can generate a unique ionic blockade signal 340 due to the tag’s distinct chemical structure and / or size, thereby electronically identifying the added base to which the tag is attached.

[0078] Alternatively, the nanopore devices contemplated by the disclosure are used for Sequencing by Expansion ("SBX") as shown in FIG. 3B, a nanopore-based nucleic acid sequencing method that uses a biochemical process to transcribe the sequence of DNA onto a measurable polymer molecule referred to as an xpandomer (“XP molecule” or “XP”). See e.g., U.S. Pat. No. 7,939,259, entitled, “High Throughput Nucleic Acid Sequencing by Expansion;” and PCT publication WO2020236526A1, entitled “Translocation control elements, reporter codes, and further means for translocation control for use in nanoporesequencing.” In the SBX process, a target nucleic acid sequence is encoded along the backbone XP sequence with reporter constructs that are separated by ~10 nm that are designed to provide high signal-to-noise, well-differentiated response signals during nanopore translocation. The enhanced signal-to-noise provided by the different response signals provides significantly increased sequence read efficiency and accuracy of XPs relative to native nucleic acid molecules.

[0079] SBX chemistry sequences nucleic acids by creating an XP from a nucleic acid template. This is achieved by encoding the nucleic acid information on a surrogate polymer of extended length which is easier to detect. The surrogate polymer, i.e., XP, is formed by template directed synthesis which preserves the original genetic information of the target nucleic acid, while also increasing linear separation of the individual elements of the sequence data.

[0080] In one embodiment, a method is disclosed for sequencing a target nucleic acid, comprising: a) providing a daughter strand produced by a template-directed synthesis, the daughter strand comprising a plurality of subunits coupled in a sequence corresponding to a contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise a tether, at least one probe or nucleobase residue, and at least one selectively cleavable bond; b) cleaving the at least one selectively cleavable bond to yield an XP of a length longer than the plurality of the subunits of the daughter strand, the XP comprising the tethers and reporter elements for parsing genetic information in a sequence corresponding to the contiguous nucleotide sequence of all or a portion of the target nucleic acid; and c) detecting the reporter elements of the XP.

[0081] In more specific embodiments, the reporter elements for parsing the genetic information may be associated with the tethers of the XP, with the daughter strand prior to cleavage of the at least one selectively cleavable bond, and / or with the XP after cleavage of the at least one selectively cleavable bond. The XP may further comprise all or a portion of the at least one probe or nucleobase residue, and the reporter elements for parsing the genetic information may be associated with the at least one probe or nucleobase residue or may be the probe or nucleobase residues themselves. Further, the selectively cleavable bond may be a covalent bond, an intra-tether bond, a bond between or within probes or nucleobase residues of the daughter strand, and / or a bond between the probes or nucleobase residues of the daughter strand and a target template.

[0082] In further embodiments, oligomer substrate constructs for use in a template directed synthesis for sequencing a target nucleic acid are disclosed. Oligomer substrate constructscomprise a first probe moiety joined to a second probe moiety, each of the first and second probe moieties having an end group suitable for the template directed synthesis, and a tether having a first end and a second end with at least the first end of the tether joined to at least one of the first and second probe moieties, wherein the oligomer substrate construct when used in the template directed synthesis is capable of forming a daughter strand comprising a constrained XP and having a plurality of subunits coupled in a sequence corresponding to the contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise a tether, the first and second probe moieties and at least one selectively cleavable bond.

[0083] In another embodiment, monomer substrate constructs for use in a template directed synthesis for sequencing a target nucleic acid are disclosed. Monomer substrate constructs comprise a nucleobase residue with end groups suitable for the template directed synthesis, and a tether having a first end and a second end with at least the first end of the tether joined to the nucleobase residue, wherein the monomer substrate construct when used in the template directed synthesis is capable of forming a daughter strand comprising a constrained XP and having a plurality of subunits coupled in a sequence corresponding to the contiguous nucleotide sequence of all or a portion of the target nucleic acid, wherein the individual subunits comprise a tether, the nucleobase residue and at least one selectively cleavable bond.

[0084] In yet further embodiments, template-daughter strand duplexes are disclosed comprising a daughter strand duplexed with a template strand, as well as to methods for forming the same from the template strand and the oligomer or monomer substrate constructs.

[0085] In some embodiments, a unique molecular identifier (UMI) is a type of molecular barcode that can be attached to the XP molecule. The UMI can typically be attached to one end of the XP molecule, such as the starting end, for example. The UMI can be constructed using XNTPs that encode for the four nucleotides (e.g. A, T, G, and C). The size of the UMI can be between about 5 and 20 nucleotides, but can be longer or shorter depending on how much information needs to be encoded into the UMI. For example, in some embodiments, the UMI is less than 5, 10, 15, 20, 25, 30, 35, 40, 45, and 50 nucleotides in length. The UMI is typically attached to the nucleic acid molecules purified from the sample before the PCR amplification process takes place. Since each nucleic acid molecule receives a different UMI, it enables various analysis steps such as deduplication and clustering.

[0086] In some embodiments, cell barcodes can also be added to the XP molecule in a similar manner as the addition of the UMI. The cell barcode, however, will be the same for all nucleic acids purified from an individual cell.

[0087] Other types of barcodes can also be added, such as sample barcodes, for example.

[0088] In the embodiment shown in Fig. 3B, the XP molecule, 345, includes a translocational control element (TCE, 350) which serves to arrest XP translocation through a nanopore (347). The TCE (350) is surrounded by reporter codes 351 and 352. In this embodiment, TCE (350) has a larger physical bulk relative to that of the reporter codes 351 and 352. XP translocation through the nanopore 347 is arrested when TCE 350 encounters the pore aperture 353. In certain embodiments, both the bulk of the TCE and the charge densities of the reporter codes (i.e., the local electric field at the arrest site) contribute to translocation arrest. During the pause, reporter code 354 is held in the barrel of the nanopore and blocks the flow of current through the pore in a characteristic and detectable manner. To overcome the arrest, a voltage pulse is applied to the system, which forces the TCE to enter and pass through the pore. Translocation then resumes until the next TCE encounters the pore aperture. In one embodiment, SBX chemistry is facilitated by the inclusion of additives that, e.g., enhance the translocation rate of XP molecules through a pore, including but not limited to, stabilizers such as EDTA and redox reagents. In a specific embodiment, concentrations from about lOmM to about 300mM of one or more organic and inorganic redox-capable species are included, e.g., ferri / ferro- cyanide, metal bipyridine compounds such as iron tris- bipyridine or cobalt tris-bipyridine, and modified ferrocenes.

[0089] In the SBS and SBX processes, an electrical signal, e.g., resistance or conductance, of the nanopore including the loaded (threaded) tag or XP can be measured via a signal value (e.g., voltage or a current passing through the nanopore), thereby providing an identification of the species and thus the nucleotide at the position of the template nucleic acid. In some embodiments, a direct current (DC) signal is applied to the nanopore cell (e.g., so that the direction in which the species moves through the nanopore is not reversed). However, operating a nanopore sensor for long periods of time using a direct current can change the composition of the electrode, unbalance the ion concentrations across the nanopore, and have other undesirable effects that can affect the lifetime of the nanopore cell. Applying an alternating current (AC) waveform can reduce the electro-migration to avoid these undesirable effects and have certain advantages as described below. The nucleic acid sequencing methods described herein are fully compatible with applied AC voltages, and therefore an AC waveform can be used to achieve these advantages.

[0090] D. Single Cell RNA Sequencing using

[0091] Numerous technological advances have enabled transcriptome profiling of single cells at increasing scale. Researchers continuously seek to sequence more reads per cell and morecells per experiment to better resolve subpopulations and discover biological insights. Furthermore, spatial transcriptomics provides spatial context of gene expression within tissues, and gains in sequencing throughput per spatial barcode enable higher resolution of colocalization patterns and cell -cell interactions. By leveraging a novel next generation sequencing (NGS) chemistry called sequencing-by-expansion (SBX), it is possible to achieve unprecedented read throughput for analysis of existing kits for single cell RNA sequencing (scRNA-seq) and spatial transcriptomics, including the lOx Genomics Single Cell 3’ Gene Expression kit, lOx Genomics Single Cell 5’ Gene Expression kit, lOx Chromium Single Cell Gene Expression Flex kit, and Parse Biosciences Evercode WT kit, Curio Seeker Spatial Mapping Kit, and the Visium HD Spatial Gene Expression Kit.

[0092] Given that most existing analysis methods are tailored for Illumina paired-end sequencing-by-synthesis data, we characterized the baseline performance of existing tools on SBX data and demonstrated improved usable read throughput with chemistry and informatic optimizations. For Oxford nanopore sequencing data, cell barcode assignment is often guided by matched Illumina sequencing. While a method to call cell barcodes on Oxford nanopore data without matched Illumina sequencing was later developed, it requires full length reads (defined by an Illumina Readl sequence and TSO) and local alignment of read elements surrounding the barcode and UMI (part of the Illumina Readl sequence and a polyT). Due to the unique nature of the SBX workflow (which can be shortened relative to Illumina workflows) and the error profile of SBX reads, methods have been developed that improve the usable read throughput of single cell and spatial transcriptomic workflows on SBX.

[0093] Various cDNA libraries of peripheral blood mononuclear cells (PBMCs) isolated from healthy human donors were prepared for SBX sequencing using a shortened workflow compared to Illumina sequencing prep. For probe-based scRNA-seq, we analyzed human PBMC and mouse samples. The probe-based approach enables scRNA-seq on clinically relevant sample types, such as formalin fixed and paraffin embedded tissue, which is the most common sample type in clinical settings. We also performed spatial transcriptomic analysis of a human brain sample using the Curio Seeker Spatial Mapping Kit and the Visium HD Spatial Gene Expression Kit.

[0094] For baseline comparisons, SBX data was pre-processed to mimic the paired-end read format expected by available bioinformatics pipelines, such as Cell Ranger, split-pipe, and Curio Seeker. Since SBX had lower valid barcode rates and higher transcriptome mapping rates than Illumina, we compared overall sequencing throughput using the number of useful reads attained, defined as reads with valid barcodes that confidently mapped to the referencetranscriptome. We then implemented chemistry or informatics optimizations to increase the fraction of useful reads. Below are tables summarizing key metrics, such as the percentage of useful reads and useful read throughput per hour of sequencing relative to Illumina. Baseline results are compared with incremental improvements in the analysis, such as correcting cell barcodes by allowing 1 bp insertion or deletion (Indel) differences from observed valid cell barcodes and ignoring the UMI base quality fdter. Because some existing bioinformatics pipelines already correct for 1 bp substitution errors, the modified workflow may not need to correct for this error and can instead use the correction feature in the existing bioinformatics pipeline. However, in some embodiments, the modified workflow can correct for both Indel errors and substitution errors.

[0095] Addition of a runway in the preparation of the SBX library also improved valid barcode rates, since the first 1-2 bases of an SBX read can have lower base quality. For example, the runway can be a short sacrificial or dummy XNP sequence that is added to the beginning of the XP molecule and in front of the barcodes. In some embodiments, the runway can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases long, or can be less than 2, 3, 4, 5, 6, 7, 8, 9, or 10 bases long. The runway shifts the critical cell barcode bases to positions without base quality issues, thereby increasing the overall percentage of useful reads.

[0096] Across all kits that were not probe-based (which would have a fixed insert length), SBX reads were substantially longer (e.g., mean >300 SBX bases vs. 90 Illumina bases covering the cDNA insert), which could enable isoform detection and higher resolution gene expression profiling. Generally, downstream results such as cell type annotation from SBX were similar to those obtained with Illumina sequencing, demonstrating the successful application of SBX for scRNA-seq and spatial transcriptomics. Using matched Illumina sequencing, we found >98% agreement in detected cell barcodes and >88% in predicted PBMC cell types. For a library of 10k human peripheral blood mononuclear cells (PBMCs), SBX and Illumina sequencing resulted in a similar number of detected genes per cell, highly correlated expression levels, and 98% agreement in detected PBMC cell types. For a 16-plex library of a total of 200k mouse cells, the cell and UMI distributions across the 16 probe sets between Illumina and SBX were well-correlated. For the human brain sample, SBX and Illumina sequencing produced visually very similar spatial maps and comparable distributions of human brain cell type clusters, classes, and subclasses.

[0097] SBX sequencing is not only compatible with available workflows for spatial transcriptomics and scRNA-seq, but also a higher throughput technology that can help accelerate basic and clinical research.

[0098] 1. Example 1

[0099] FIG. 4 illustrates the workflow for the 10X single cell 3’ gene expression kit, which starts with obtaining a sample with mRNA 400 from an individual cell, which can be done using an instrument from 10X Genomics. The mRNA is captured and converted to cDNA using reverse transcriptase (RT), and a barcode and UMI and a polyT insert can be added to the cDNA fragment 402. A representation of the captured oligo and modified cDNA is shown 403. After template switch 404 using a template switch oligo (TSO) and transcript extension 406, the cDNA is released 408. The cDNA is amplified 410, using PCR for example, and can be cleaned up using SPRI beads 412 to obtain fragments with a consistent size distribution. From there, the workflow can proceed to the standard Illumina workflow, or can enter a shortened SBX workflow that includes strand enrichment 414 and then SBX sequencing 416.

[0100] As shown in FIG. 5, the raw sequencing read data that is generated from the sequencer can be stored in a FASTQ file 500, which is the starting point of the bioinformatics portion of the workflow. The 10X bioinformatics software, Cell Ranger 506, can be used to quickly identify an initial observed valid barcode list on a subsample 502 of read data that has been trimmed and split 504. However, Cell Ranger only corrects barcodes with one bp substitution difference from an observed valid barcode, and does not correct invalid barcodes with a one bp Indel difference from a valid observed barcode. Due to differences between Illumina and SBX sequencing, it may be desirable to correct for one bp Indels in the barcode as well, which is shown in FIG. 5. The read data can be trimmed and split 508 and then the barcodes with one bp Indel difference can be identified (by comparison to the valid barcode) and then corrected 510 by, for example, inserting the missing bp for a deletion error or by deleting the added bp for an insertion error, so that the barcode will be valid when sent to Cell Ranger for further processing 512. If there are multiple valid barcode candidates, select by highest observed frequency. In some embodiments, correction for barcode Indel errors is only optionally done when the barcorde errors are determined to occur in the sample at above a threshold frequency, such being in the top 3% of observed valid barcodes. In other embodiments, the threshold frequency can be 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10%.

[0101] In addition, Cell Ranger discards reads with a UMI base with base quality less than 10. In some embodiments, this base quality filter is bypassed by replacing all UMI base qualities that are less than 10 with 10 so that no reads are discarded using the UMI base quality filter 511.

[0102] As shown in FIG. 6, addition of a runway 604 can be done to improve the valid barcode rate because the first 1 to 2 bases of the SBX read can have lower base quality. For comparison, one read 600 does not have the runway, while the other read 602 has the runway 604.1 Ox Genomics Single Cell SBX SBX SBX Matched3 ’ Gene Expression kit baseline corrected BC corrected BC Illumina and UMI BQE z ffective read length 398 _no_ xno 1l1lOo(mean)Reads confidently81 8o / o 81 8o / o 81 8o / o80.2% mapped to transcnptomeValid barcode rate 72.9% 87.0% 87.0% 87.0%Valid barcode and UMI 55.5%rate% useful reads 45.4% 54.1% 72.2% 79.2%Number of reads* into .. .. .. .. .. ..o.„ „ „ 8.00 B 8.00 B 8.00 B 25 BDpairsCell RangerNumber of useful reads 3.63 B 4.33 B 5.78 B 19.8 BInstrument run time 2 hr 2 hr 2 hr 48 hrUseful reads per hour 1.82 B 2.17 B 2.89 B 413 MUseful read per hourrelative to IlluminaTable 1A. Number of reads extrapolated to 8M array for SBX and 25B 2x150 NovaSeqX for Illumina, showing the effect of correcting cell barcodes and ignoring the UMI base quality filter.1 Ox Genomics Single Cell SBX SBX SBX SBX Matched3 ’ Gene Expression kit baseline runway runway and runway and Illumina corrected BC correctedBC andUMI BQEffective read length 3984 | 2 4 | 2 4 |2 | 18(mean)Reads confidently81 8o / o 83 Q% 83 0% 83.0% 80.2% mapped to transcnptomeValid barcode rate 72.9% 93.4% 97.6% 97.6% 87.0%Valid barcode and UMI 55.5 % 77 2% 80 4% 98 0% 98 7% rate% useful reads 45.4% 64.1% 66.7% 81.4% 79.2%Number of reads* into g.QO B 6.72 B 6.72 B 6 72 B25 BCell Ranger pairsNumber of useful reads 3.63 B 4.31 B 4.48 B 5.47 B 19.8 BInstrument run time 2 hr 2 hr 2 hr 2 hr 48 hrUseful reads per hour 1.82 B 2.15 B 2.24 B 2.74 B 413 MUseful read per hour4 4X 5 2X 5 4X 6,6Xrelative to IlluminaTable IB. Number of reads extrapolated to 8M array for SBX and 25B 2x150 NovaSeqX for Illumina, showing the effect of adding a runway, correcting cell barcodes and ignoring the UMI base quality fdter.

[0103] For the 10X 3’ kit, we obtained 1.8 billion useful reads per hour of SBX sequencing, which was 4.4X throughput compared to an Illumina NovaSeqX run. With the addition of a runway or informatic read corrections, SBX useful read per hour throughput could be pushed to 5.2-7. OX compared to Illumina. Notably, the same informatic optimizations of allowing a Ibp Indel difference in the cell barcode from the observed valid barcode list and removing the UMI base quality threshold only improved Illumina useful read percentages by up to 0.4%.

[0104] 2. Example 2

[0105] FIG. 7 illustrates the Parse Evercode Workflow for both SBX and Illumina sequencing platforms. Starting with mRNA from a single cell 700, RT is used to generate cDNA 702, and barcodes are ligated to the cDNA 704. The cDNA is hybridized to a bead 706, and atemplate switch using TSO is performed 708. The cDNA is amplified 710, and cleaned up using SPRI beads 712. From this point, the workflow can proceed using the Illumina workflow, or proceed with the SBX workflow, which involves strand enrichment 714 and then SBX sequencing 716. The final modified cDNA fragment is shown 718.

[0106] To improve SBX performance on the Parse pipeline, SBX read can be corrected using the procedure shown in FIG. 8. First, identify linker22 800. The 8 bases on the right of linker22 is barcodel (BC1), and the 8 bases on the left of linker22 is barcode2 (BC2) 802. Next, identify linker30 804. The 8 bases to the left of linker30 is barcode3 (BC3) 806. The 10 bases to the left of BC3 is the UMI 808. If there are fewer than 10 bases, take the first 10 bases of the read as the UMI. This completes SBX read construction.Parse Biosciences Evercode WT SBX SBX Matched kit baseline corrected IlluminaEffective read length (mean) 444 444 160Reads mapped to transcriptome 65.5% 65.5% 60.5%Valid barcode rate 63.3% 82.8% 78.6%% useful reads 41.4% 54.1% 47.6%Number of reads* into split-pipe 5.96 B 5.96 B 25 B pairsNumber of useful reads 2.47 B 3.22 B 11.9 BInstrument run time 2 hr 2 hr 48 hrUseful reads per hour 1.23 B 1.61 B 247.9 MUseful read per hour relative to _nv, _vIlluminaTable 2. Number of reads extrapolated to 8M array for SBX and 25B 2x150 NovaSeqX for Illumina using corrected barcodes.

[0107] For the Parse kit, even though the split-pipe pipeline has a barcode edit distance parameter, increasing this threshold from the default of 2 to 3 only improved the useful read percentage from 41% to 44%. We implemented an approach to iteratively parse each barcode in the read based on the surrounding linker sequences, as described above. This improved the percentage of useful reads from 41% to 54%. We obtained 1.2-1.6 billion useful reads per hour of SBX sequencing, which was still 5-6.5X throughput compared to Illumina. The read correction approach allows for shifts in linker sequences to better extract the barcode positions, and could not be applied to Illumina data given the fixed read length of Illumina reads.

[0108] FIG. 14 illustrates that using SBX sequencing reads with the Parse kits results in similar predicted cell types as using Illumina sequencing reads with the Parse kits.

[0109] 3. Example 3

[0110] FIG. 9 illustrates an embodiment of the workflow for the 10X Chromium Single Cell Gene Expression Flex Kit. The workflow starts with the fixation of single cells to preserve the mRNA the cells. Next, the mRNA within the cells are hybridized to probes which include a probe barcode as well as a construct of the mRNA sequence. A 10X instrument can be used to perform single cell partitioning of the cells in the sample 904, and additional barcodes and UMI can be ligated in a probe extension step 906. The modified probe can thenbe amplified 908 and cleaned up using SPRI beads 910. From this point, the workflow can proceed either to an Illumina workflow, or enter an SBX workflow which includes strand enrichment 912 and SBX sequencing 914. A representation of the modified probe is shown 916.[oni] As shown in FIG. 10, the raw sequencing read data that is generated from the sequencer can be stored in a FASTQ file 1000, which is the starting point of the bioinformatics portion of the workflow. The reads can be trimmed and split 1002 into two reads (R1 and R2) which can be stored in separate FASTQ files 1004. The 10X bioinformatics software, Cell Ranger 1006, can be used to align the sequences, which can be stored in a BAM file 1008. Next, the R2 reads can be retrieved from the R2 FASTQ file and can be aligned against a list of known probe sequences (currently about 55k different probes) using VSEARCH. Next, the probe sequences can be corrected, where needed 1010. The R2 reads and any reads from Cell Ranger that were not assigned to cell barcodes can be adjusted or modified to include the full sequence of the best probe match (optionally with greater than 70, 75, 80, 85, 90, or 95% identity). Then the probe barcodes can be corrected 1012, which involves checking that the 8 bp (or other size) probe barcode is in the expected position against the list of allowed probe barcodes (128). Allow edit 1-2 and correct. Cell Ranger can be rerun with the updated R2 reads 1014, and optionally with corrected Gel Bead in Emulsion (GEM) barcode in R1 reads, where cell barcodes with 1 Indel are allowed and edited. lOx Chromium Single Cell SBX SBX SBX SBX MatchedGene Expression Flex kit baseline corrected corrected corrected Illumina(runway probe probe and probe, included) probe BC probe BC,GEM BCReads confidently mapped tog2 6% 94 2%,4 2%,4 2% % 3%probe setValid barcode rate 77.7% 80.9% 90.3% 94.0% 95.1%Valid GEM barcode rate 92.6% 92.6% 92.6% 96.5% 97.5%Valid probe barcode rate 82.5% 86.0% 96.5% 96.5% 96.7%% useful reads 72.3% 76.4% 85.7% 89.4% 93.7%Number of reads* into Cell „ 10 B„ 16.2 B 16.2 B 16.2 B 16.2 BRanger pairsNumber of useful reads 11.71 B 12.38 B 13.89 B 14.48 B 9.4 BInstrument run time 2 hr 2 hr 2 hr 2 hr 22 hrUseful reads per hour 5.85 B 6.19 B 6.94 B 7.24 B 426 MUseful read per hour relative u7X 14.5x i6JX17.0X to IlluminaTable 3. Number of reads extrapolated to 8M array for SBX and 10B 2x100 NovaSeqX for Illumina using runway, corrected probe, probe barcode correction, and cell barcode correction.

[0112] For the 10X Flex kit, we implemented probe sequence realignment using VSEARCH on reads that initially failed cell barcode assignment by the Cell Ranger pipeline. The SBX read was re-split and corrected to include the complete probe sequence of the best probe match with over 80% identity. Additional informatic optimizations were applied, including probe barcode correction and GEM barcode correction. Probe barcodes were allowed to be 1- 2 edit distance away from the known list of 128 probe barcodes. The more stringent edit distance of 1 was applied for barcodes that were known to be edit 2 apart from other allowed barcodes. GEM barcode correction entailed allowing 1 bp Indel differences from the observed valid GEM barcode list. The throughput of useful reads (i.e., those with valid barcodes and mapped to the transcriptome), were 13.7-17X higher on SBX than Illumina per hour of sequencing.

[0113] 4. Example 4

[0114] FIG. 11 illustrates an embodiment of the Curio Seeker workflow. The workflow starts with a tissue sample 1100, which is typically a thin slice (i.e., 10 um) that has been affixed to a slide. Next, the tissue sample can be hybridized to beads 1102 and RT can be added 1104 to initiate single strand synthesis 1106 of the cDNA, which can be released 1108, amplified 1110, and cleaned 1112. From here, the workflow can enter the standard Illumina workflow, or can enter the SBX workflow, which involves strand enrichment 1114 and SBX sequencing 1116. A representation of the cDNA construct is shown 1118.Curio Seeker Spatial Mapping SBX SBX MatchedKit baseline corrected IlluminaEffective read length (mean) 330 330 100Reads uniquely mapped to „ / .1l0 / / „ o / 5.110 / / O Oi- 1.0 / / 0 transcnptomeReads with proper read structure 19.9% 75.6% 78.7%% useful reads 8.3% 38.7% 38.0%Number of reads* into Curio „ . „ „ . „ . „ „„ . 7.1 B 7.1 B 10 B pairsSeekerNumber of useful reads 0.59 B 2.74 B 3.8 BInstrument run time 2 hr 2 hr 18 hrUseful reads per hour 294 M 1.37 B 211 MUseful read per hour relative to , ... . _..... .H1.4X 6.5XIlluminaTable 4. Number of reads extrapolated to 8M array for SBX and 10B 2x50 NovaSeqX for Illumina with

[0115] For the Curio spatial kit, reads with Indels or more than one substitution in the linker sequence would be considered improper by the Curio Seeker pipeline and discarded.Therefore, we implemented an informatic correction to search and correct linker sequences, allowing the pipeline to properly recognize the surrounding bead barcode and UMI sequences. This essentially allows for shifted barcode start and end positions. The last nine bases of Read 1 were also replaced with a polyT to remove the requirement of a near perfect polyT at the end of Read 1 for a read to be considered proper. This improved the percentage of useful reads from 8.3% to 38.7%.

[0116] More specifically, the correction is done by searching for linkers with up to 3 substitutions or Indels in the read. If found, replace with perfect linker sequence if there were errors. If linker starts at position 10+, shift back read by appropriate number of bases (Curio pipeline currently requires linker to be at positions 9-26). Make the last 9 bases of the read Ts (pipeline requires 8 of 9 last bases to be T).

[0117] FIG. 15 illustrates that use of SBX sequencing reads with the Curio kit and platform results in similar spatial mapping of predicted clusters as use of Illumina sequencing reads with the Curio Kit and platform.

[0118] 5. Example 5

[0119] FIG. 12 illustrates an embodiment of the Visium HD Probe Capture workflow. The workflow starts with a tissue sample 1200, which can be a thin slice fixed to a slide. Next, probes are hybridized 1202, released 1204, and captured on a slide 1206. After probe extension 1208 and denaturation 1210, the workflow can proceed with a Illumina workflow, or SBX workflow, which involves strand enrichment 1212 and SBX sequencing. A representation of the probe construct is shown 1216.Visium HD Spatial Gene Expression Kit SBX baseline SBX Public(runway corrected unmatched included) probe IlluminaReads confidently mapped to probe set 84.4% 92.4% 98.2%Valid barcode rate 86.5% 87.0% 98.7%% useful reads 73.9% 80.8% 96.8%Number of reads* into Space Ranger 25.3 B 35.7 B 10 B pairsNumber of useful reads 18.7 B 28.8 B 9.7 BInstrument run time 5 hr 5 hr 18 hrUseful reads per hour 3.74 B 5.76 B 539 MUseful read per hour relative to Illumina 6.9X 10.7XTable 5. Number of reads extrapolated to 8M array for SBX and 10B 2x50 NovaSeqX for Illumina

[0120] For the 10X Visium HD kit, we implemented probe sequence realignment using VSEARCH on reads that initially failed assignment by the Space Ranger pipeline. The SBX read was re-split and corrected to include the complete probe sequence of the best probe match with over 80% identity. Additional informatic optimizations were applied, including removal of a polyA tail trimming step that typically operated on a third of reads. This trimming step could render a probe read sequence less than the expected 50 bp length and not included in molecule counts in Space Ranger. With these changes, the throughput of useful reads improved from 3.74 B per hour to 5.76 B per hour, which was 6.9-10.7X higher Illumina per hour of sequencing.

[0121] 6. Summary

[0122] The following table summarizes the data pipeline optimizations that resulted in an increase in effective SBX read usage by about 10 to 30 percent.input reads into pipeline

[0123] FIG. 13 illustrates the different types of reads generated by the different scRNA kits described herein. Splitting of SBX reads into Readl and Read2 enables analysis using existing bioinformatics pipelines from the manufacturer of the kits. Bioinformatics analysis may include, for example, base calling, gene expression (e.g., 375’ gene expression), antibody capture, CRISPR capture, variant calling, cell annotation, spatial transcriptome mapping, and others such as further described herein.

[0124] E. Computer Systems

[0125] Any of the computer systems mentioned herein can utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG. 9 in computer system 1110. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system includes multiple computer apparatuses, each being a subsystem, with internal components. A computer system can include desktop and laptop computers, tablets, mobile phones, and other mobile devices.

[0126] The subsystems shown in FIG. 16 are interconnected via a system bus 1180. Additional subsystems such as a printer 1174, keyboard 1178, storage device(s) 1179, monitor 1176 which is coupled to display adapter 1182, and others are shown. Peripherals and input / output (RO) devices, which couple to RO controller 1171, can be connected to the computer system by any number of means known in the art such as I / O port 1177 (e.g., USB, FireWire®). For example, RO port 1177 or external interface 1181 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 1600 to a wide area network such as theInternet, a mouse input device, or a scanner. The interconnection via system bus 1180 allows the central processor 1173 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 1172 or the storage device(s) 1179 (e.g., a fixed disk, such as a hard drive, or optical disk), as well as the exchange of information between subsystems. The system memory 1172 and / or the storage device(s) 1179 can embody a computer readable medium. Another subsystem is a data collection device 1175, such as a camera, microphone, accelerometer, and the like. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0127] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface 1181, by an internal interface, or via removable storage devices that can be connected and removed from one component to another component. In some embodiments, computer systems, subsystem, or apparatuses communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0128] Aspects of embodiments can be implemented in the form of control logic using hardware circuitry (e.g., an APSIC or FPGA) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor can include a single-core processor, multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present invention using hardware and a combination of hardware and software.

[0129] Any of the software components or functions described in this application can be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C #, Objective-C, Swift, or scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code can be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission. A suitable non-transitory computer readable medium can include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium can be any combination of such storage or transmission devices.

[0130] Such programs can also be encoded and transmited using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium can be created using a data signal encoded with such programs. Computer readable media encoded with the program code can be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium can reside on or within a single computer product (e.g., a hard drive, a CD, or an entire computer system), and can be present on or within different computer products within a system or network. A computer system can include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0131] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments can be directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective step or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or at different times or in a different order. Additionally, portions of these steps can be used with portions of other steps from other methods. Also, all or portions of a step can be optional. Additionally, any of the steps of any of the methods can be performed with modules, units, circuits, or other means of a system for performing these steps.

[0132] Although various illustrative embodiments are described above, any of a number of changes may be made to various embodiments without departing from the scope of the invention as described by the claims. For example, the order in which various described method steps are performed may often be changed in alternative embodiments, and in other alternative embodiments one or more method steps may be skipped altogether. Optional features of various device and system embodiments may be included in some embodiments and not in others. Therefore, the foregoing description is provided primarily for exemplary purposes and should not be interpreted to limit the scope of the invention as it is set forth in the claims.

[0133] The examples and illustrations included herein show, by way of illustration and not of limitation, specific embodiments in which the subject mater may be practiced. As mentioned, other embodiments may be utilized and derived there from, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. Such embodiments of the inventive subject mater may be referred to hereinindividually or collectively by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or inventive concept, if more than one is, in fact, disclosed. Thus, although specific embodiments have been illustrated and described herein, any arrangement calculated to achieve the same purpose may be substituted for the specific embodiments shown. This disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the above embodiments, and other embodiments not specifically described herein, will be apparent to those of skill in the art upon reviewing the above description.

Claims

1. CLAIMSWhat is claimed is:

1. A method for performing single cell RNA transcriptone sequencing, the method comprising: receiving a sample of mRNA; reverse transcribing the mRNA to generate cDNA; amplifying the cDNA; causing generation of a template-dependent polymerase-based replication of the cDNA with an expandable nucleoside triphosphate (X-NTP) substrate, wherein the replicated cDNA with X-NTP comprises at least one of a molecule-identifying barcode or unique molecular identifier (UMI); causing passage of the replicated cDNA with X-NTP through one or more nanopores of an electrically resistant membrane, wherein a voltage is applied across each of the one or more nanopores; measuring sequences of electromagnetic signals generated in response to passage of the replicated cDNA with X-NTP through each of the nanopores; converting the electromagnetic signals into sequencing data identifying molecular bases of the cDNA; identifying and correcting each of one base pair substitution, insertion, and deletion errors within the sequencing data pertaining to the at least one barcode or UMI; validating the sequencing data for bioinformatics analysis based on the errors identified within the sequencing data pertaining to the at least one barcode or UMI.

2. The method claim 1 wherein generating the replicated cDNA and X-NTP comprises creating a runway of molecules leading to the at least one barcode or UMI.

3. The method of claim 2 wherein the size of the runway is configured to correspond with an expected initial error rate associated with reading leading molecules of the cDNA and X-NTP, wherein the expected initial error rate is different than an expected error rate associated with reading non-leading molecules of the cDNA and X-NTP.

4. The method of claim 3 wherein the size of the runway is less than or equal to two base pairs in length.

5. The method of claim 1 wherein the validating comprises a greater threshold for an error within a portion of the sequence data including a BMI compared to a portion of the sequencing data that does not include a BMI.

6. The method of claim 5 wherein the validating selectively ignores errors within the portions of the sequencing data including a BMI.

7. The method of claim 1 wherein the validating comprises a barcode reading error threshold of between about 1% and 10% of top observed valid barcodes.

8. The method of claim 6 wherein the validating comprises a barcode reading error threshold of about 3% of top observed valid barcodes.

9. The method of claim 1 comprising storing the sequencing data within a standardized fde format for representing nucleotide sequences and sequence quality information, and wherein the identifying and correcting is processed using the standardized file format.

10. The method of claim 9 wherein the standardized file format comprises FASTQ.

11. A method for performing single cell RNA probe sequencing, the method comprising: receiving a sample of single cells with mRNA; hybridizing the mRNA in the single cells by binding it with a complementary target segment of nucleotides to form probes; partitioning the single cells in a gel bead in emulsion; extending the probes, wherein the extended probes comprise a probe barcode, a cell barcode, and UMI; amplifying the extended probes; causing generation of a template-dependent polymerase-based replication of the amplified probes with an expandable nucleoside triphosphate (X-NTP) substrate; causing passage of the replicated probes with X-NTP through one or more nanopores of an electrically resistant membrane, wherein a voltage is applied across each of the one or more nanopores;measuring sequences of electromagnetic signals generated in response to passage of the replicated probes with X-NTP through each of the nanopores; converting the electromagnetic signals into sequencing data identifying molecular bases of the probes; identifying and correcting each of 1 base pair substitution, insertion, and deletion barcode errors within the probes; validating the sequencing data for bioinformatics analysis based on the barcode errors identified within the sequencing data.

12. The method of claim 11 wherein the validating comprises an error threshold of a one or two base pair edit distance away from a known list of valid probe barcodes.

13. The method of claim 11 wherein the validating comprises an error threshold of a 1 base pair (bp) substitution, insertion, or deletion.

14. A system for performing single cell RNA transcriptone sequencing, the system comprising: a plurality of nanopores situated within an electrically resistant membrane, a voltage source configured to provide a voltage across each of the nanopores, and one or more detectors configured to measure electromagnetic signals across each of the plurality of nanopores; one or more processors programed and configured to: receive measurements from the detectors of electromagnetic signals across each of the plurality of nanopores; cause generation of a template-dependent polymerase-based replication of the cDNA with an expandable nucleoside triphosphate (X-NTP) substrate, wherein the replicated cDNA with X-NTP comprises at least one of a molecule-identifying barcode or unique molecular identifier (UMI); cause passage of the replicated cDNA with X-NTP through one or more nanopores of an electrically resistant membrane, wherein a voltage is applied across each of the one or more nanopores;measure sequences of electromagnetic signals generated in response to passage of the replicated cDNA with X-NTP through each of the nanopores; convert the electromagnetic signals into sequencing data identifying molecular bases of the cDNA; identify and correcting each of one base pair substitution, insertion, and deletion errors within the sequencing data pertaining to the at least one barcode or UMI; validate the sequencing data for bioinformatics analysis based on the errors identified within the sequencing data pertaining to the at least one barcode or UMI.

15. The system of claim 14 wherein generating the replicated cDNA and X-NTP comprises creating a runway of molecules leading to the at least one barcode or UMI.

16. The system of claim 15 wherein the size of the runway is configured to correspond with an expected initial error rate associated with reading leading molecules of the cDNA and X-NTP, wherein the expected initial error rate is different than an expected error rate associated with reading non-leading molecules of the cDNA and X-NTP.

17. The system of claim 14 wherein the size of the runway is less than or equal to two base pairs in length.

18. The system of claim 14 wherein the validating comprises a greater threshold for an error within a portion of the sequence data including a BMI compared to a portion of the sequencing data that does not include a BMI.

19. The system of claim 18 wherein the validating selectively ignores errors within the portions of the sequencing data including a BMI.

20. The system of claim 14 wherein the validating comprises a barcode reading error threshold of between about 1% and 10% of top observed valid barcodes.

21. The system of claim 20 wherein the validating comprises a barcode reading error threshold of about 3% of top observed valid barcodes.

22. The system of claim 14 wherein the one or more processors are further programmed and configured to store the sequencing data within a standardized file format for representingnucleotide sequences and sequence quality information, and wherein the identifying and correcting is processed using the standardized fde format.

23. The system of claim 22 wherein the standardized fde format comprises FASTQ.

24. A method of validating molecular sequencing data, the method comprising: obtaining data representing molecular bases of DNA or RNA; identifying and correcting each of one base pair substitution, insertion, or deletion errors within the sequencing data pertaining to the at least one barcode, UMI, or runway of molecular bases leading to a UMI; validating the sequencing data for bioinformatics analysis based on the errors identified within the sequencing data pertaining to the at least one barcode, UMI, or runway leading to a UMI, wherein the validating comprises a greater threshold for an error within a portion of the sequence data including a BMI, barcode, or runway compared to portions of the sequencing data that does not include a BMI, barcode, or runway respectively.

25. The method of claim 24 wherein the sequencing data comprises a runway leading to a BMI and wherein the threshold comprises a greater threshold for an error within the runway leading to the UMI compared to the portions of the sequencing data that does not include the runway.26 The method of claim 24 wherein the validating comprises a greater threshold for an error with a portion of the sequence data including a BMI compared to the portions of the sequencing data that does not include the BMI.

27. The method of claim 24 wherein the sequence data represents sequencing data from a nanopore sequencer.

28. A system for performing single cell RNA transcriptone sequencing, the system comprising: a plurality of nanopores situated within an electrically resistant membrane, a voltage source configured to provide a voltage across each of the nanopores, and one or more detectors configured to measure electromagnetic signals across each of the plurality of nanopores; one or more processors programed and configured to:obtain molecular sequencing data representing mRNA in the single cells bound with a complementary target segment of nucleotides to form probes extended with a probe barcode, a cell barcode, and UMI; identifying and correcting each of 1 base pair substitution, insertion, and deletion barcode errors within the probes; validating the sequencing data for bioinformatics analysis based on the barcode errors identified within the sequencing data.

29. A non-transitory computer readable medium comprising instructions which, when executed by one or more processors, cause execution of any of the methods of claims 1-13 and / or claims 24-27.

Citation Information

Patent Citations

  • Nanopore Based Molecular Detection and Sequencing

    US20130244340A1

  • DNA sequencing by synthesis using modified nucleotides and nanopore detection

    US20130264207A1

  • Nucleic acid sequencing using tags

    US20140134616A1

  • Chemical methods for producing tagged nucleotides

    US20150368710A1

  • Tagged nucleotides useful for nanopore detection

    US20180057870A1