Methods and kits

By using the LRL method to split the cell population before reverse transcription and label it with different barcodes to form a triple barcode construct, the problem of barcode encoding crosstalk in existing technologies is solved, and accurate labeling and measurement of single-cell transcriptomes are achieved.

CN120898002APending Publication Date: 2025-11-04OXFORD NANOPORE TECH LTD
View PDF 38 Cites 0 Cited by

Patent Information

Application Number
CN202480021003.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-24
Filing Date
2024-03-22
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing Split-Seq methods are prone to re-barcoding or crosstalk during reverse transcription, leading to confusion of intracellular transcripts and making it impossible to accurately label and measure the transcriptome of individual cells.

Method used

The LRL (linking + RT + linking) method was used. Cell populations were split and labeled with different barcodes before reverse transcription. By linking the first adaptor in the first round, reverse transcription in the second round, and linking the second adaptor in the third round, a triple barcode construct was formed, which reduced crosstalk and ensured that the RNA molecules of each cell were uniquely labeled.

Benefits of technology

It effectively reduces barcode encoding crosstalk, ensures unique markers for RNA molecules in each cell, improves the accuracy and integrity of transcriptome measurements, and supports transcriptome characterization of single cells or cell populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The present invention relates to a method of uniquely labeling RNA molecules in a population of cells, as well as kits for use in such methods. The methods and kits of the invention allow for the study of transcriptomes in a single cell or population of cells.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates to a method of uniquely labelling RNA molecules in a cell population, and a kit for use in such a method. The method and kit of the present invention allow the study of the transcriptome in a single cell or cell population. BACKGROUND

[0002] The transcriptome is the collection of all coding and non-coding RNA transcripts in a single cell or cell population. Data on the transcriptome can be used to study cell differentiation, carcinogenesis, transcriptional regulation and biomarker discovery, among others. Transcriptome data can also be used to identify the source of one or more cells, including its / their phylogeny.

[0003] Advances in next generation sequencing (NGS) have enabled the transcriptome to be effectively studied. For example, RNA-Seq uses NGS to measure the presence and quantity of RNA molecules in a cell and is able to analyse the changing transcriptome of a cell. Furthermore, WO 2019 / 060771 describes a method of uniquely labelling or barcoding RNA molecules within each cell of a cell population. This technique is known as “Split-Seq” and involves multiple rounds of separating (or splitting) the cells into individual samples, labelling the cells in each sample with a different barcode, and then pooling the samples. However, all of the separation / splitting and pooling steps are carried out after the RNA molecules in the cell population have been reverse transcribed into cDNA. The same method is described in Rosenberg et al., Science, 360(6385): 176-182. This method can be described as an RLL (RT + ligation + ligation) method.

[0004] Biology pores (and other nanopores) have great potential as direct electrical biosensors for polymers and various small molecules. In particular, nanopores have recently been focused on as a potential DNA sequencing technology. When an electrical potential is applied across a nanopore, the current changes in the case where an analyte (such as a nucleotide) transiently resides in the barrel for a certain period of time. Nanopore detection of nucleotides gives a known signature and duration of current change. In a chain sequencing method, a single polynucleotide chain is made to pass through the pore and the identity of the nucleotides is derived. Chain sequencing can involve the use of a molecular brake to control the motion of the polynucleotide through the pore. SUMMARY

[0005] The inventors have identified a new method of uniquely labelling RNA molecules in a population of cells. This method involves multiple rounds of dividing (or splitting) a population of cells into (or into) separate samples, labelling the cells in each sample with a different barcode, and then pooling the samples. In this new method, the first round of dividing (or splitting), labelling and pooling is performed prior to reverse transcription (RT). The second dividing (or splitting step) is also performed prior to RT. These method steps distinguish the new method of the invention from Split-Seq disclosed in WO 2019 / 060771 and Rosenberg et al., Science, 360(6385): 176-182. The new method can be described as an LRL (ligation + RT + ligation) method.

[0006] The new method can be summarised as follows. Following the first dividing of the population of cells, the RNA molecules in the first divided sample are labelled with a first adaptor comprising a primer site for RT and the reverse complement of a unique first barcode for each first sample. The first divided sample is then pooled and divided a second time prior to RT using a RT primer comprising a unique second barcode for each second sample. The second divided sample is then pooled and divided a third time prior to labelling with a second adaptor comprising a unique third barcode for each third sample. This provides a tri-barcode construct from all cells. The tri-barcode construct comprises a DNA sequence transcribed from the RNA molecule (during the second round).

[0007] The barcode used for each sample (i.e. for each first, second and third sample) is preferably different to any other barcode or all other barcodes used in the method. The number of cells in the original population is preferably less than the number of first samples multiplied by the number of second samples multiplied by the number of third samples. If the number of possible barcode combinations used is greater than the number of cells in the population, then it is very likely that an RNA molecule from each cell is labelled with a unique combination of three barcodes (i.e. the barcode construct produced from each cell will have a unique combination of three barcodes).

[0008] The tri-barcode construct can then be isolated, amplified and characterised or sequenced. This allows the transcriptome of each individual cell in the population or multiple cells in the population to be measured and / or quantified. The method can be used in combination with nanopore sequencing, but this is not essential.

[0009] The method of the invention has several key advantages. For example, by separating the first and third labelling steps from the second RT step, the likelihood of rebarcoding or cross-talk upon pooling is reduced. This is a common problem associated with RLL Split-Seq, as disclosed in WO 2019 / 060771 and Rosenberg et al., Science, 360 (6385): 176-182, where two consecutive rounds of ligation are used. In this case, cross-talk means that a fragment fails to be RT primed or adaptor ligated in its individual well, but upon pooling, presents an excess of primers / adaptors and reagents that allow the fragment to be further labelled with a barcode that was not present in the previous labelling round. This can lead to transcripts within a cell being exposed to a mixed population of barcodes, and thus presenting as transcripts from different cells when considering the combined barcode. By the RT step, the RT primer primes internally within the transcript, and then allows reverse transcription to initiate and displace the previous RT product or produce a truncated RT product, which is also technically possible. This would be an example of rebarcoding, where the identity of the barcode already incorporated can be overwritten with a new barcode. Again, this would lead to transcripts presenting as if they were from different cells when considering the combined barcode. The method of the invention avoids these drawbacks of the RLL method.

[0010] The first adaptor is ligated to the 3’ end of the RNA transcript in the first round, allowing RT of the entire RNA transcript in the second round with specific RT primers, without the need for poly(T) VN reverse transcription primers that prime easily off-target. This minimises intron overlap and off-target amplification. It also allows cDNA synthesis of full-length poly(A) tails in eukaryotic RNA for end-to-end RT of RNA transcripts.

[0011] At the end of the first round, a portion of the first adaptor can be destroyed prior to cell pooling, and this reduces the likelihood of rebarcoding or cross-talk upon pooling.

[0012] Additional advantages of the method of the invention are discussed below with reference to specific examples.

[0013] The invention provides a method of uniquely labelling RNA molecules in a population of cells, the method comprising:

[0014] (a) dividing the population of cells into a plurality of first samples;

[0015] (b) ligating a first adaptor to the 3’ end of the RNA molecules in the plurality of first samples, wherein the first adaptor comprises a primer site for reverse transcription (RT), and wherein the first adaptor used in each first sample comprises the reverse complement of a different first barcode;

[0016] (c) pooling the plurality of first samples and dividing the pool into a plurality of second samples;

[0017] (d) reverse transcribing RNA molecules in the plurality of second samples using the primer site and an RT primer to form double-stranded constructs, wherein the RT primer used in each second sample comprises a different second barcode;

[0018] (e) pooling the plurality of second samples and dividing the pool into a plurality of third samples; and

[0019] (f) ligating a second adaptor to the double-stranded constructs in the plurality of third samples to form triple-barcode constructs, wherein the second adaptor used in each third sample comprises a different third barcode.

[0020] The present application also provides a method of characterizing the transcriptome of a single cell or two or more cells in a cell population, the method comprising:

[0021] (a) performing the method of the present application on the cell population; and

[0022] (b) characterizing RNA molecules from the single cell or two or more cells using their unique labels.

[0023] The present application also provides a kit for uniquely labeling RNA molecules in a cell population, the kit comprising two or more first adaptors, wherein the two or more first adaptors comprise a primer site for RT, and wherein each of the two or more first adaptors comprises the reverse complement of a different first barcode. BRIEF DESCRIPTION OF DRAWINGS

[0024] It is to be understood that the drawings are only for the purpose of illustrating certain embodiments of the application and are not intended to be limiting.

[0025] Figure 1 LRL first round barcode encoding by ligation of first adaptors is shown.

[0026] Figure 2 LRL second round barcode encoding by RT is shown.

[0027] Figure 3 LRL third round barcode encoding by ligation is shown.

[0028] Figure 4 LRL third round adaptor blocking is shown.

[0029] Figure 5It is shown how SSC-based rehydration buffers produce more intact total RNA than DPBS-based rehydration buffers. The RIN (RNA Integrity Number) score for SSC samples - a measure of the 28S to 18S rRNA peak ratio - is 2x that of DPBS samples. RNA integrity is critical throughout the cell fixation process because RNA transcripts must be preserved for cDNA synthesis as part of the combinatorial barcode encoding method. This experiment shows that methanol fixation and rehydration do not have a large impact on RNA integrity.

[0030] Figure 6 It is shown how RNA integrity is preserved upon incubation at 37°C compared to unincubated controls, with comparable RIN scores for both. RNA integrity is critical throughout the CDRA ligation and USER digestion process because RNA transcripts must be preserved for cDNA synthesis as part of the combinatorial barcode encoding method.

[0031] Figure 7 First round barcode encoding competition assay results are shown. Sequencing reads are aligned to BC01 or BC02 of the CDRA adapter against the primary Y-axis. PCR yield is aligned against the secondary Y-axis.

[0032] Figure 8 Second round barcode encoding competition assay results are shown. Sequencing reads are aligned to BC03 or BC04 of the RT barcode primer against the primary Y-axis. PCR yield is aligned against the secondary Y-axis.

[0033] Figure 9 Third round barcode encoding competition assay results are shown, where no blocking oligo was used. Sequencing reads are aligned to BC05 or BC06 of the third round ligation adapter against the primary Y-axis. PCR yield is aligned against the secondary Y-axis.

[0034] Figure 10 Third round barcode encoding competition assay results are shown, where the RLL method standard blocking oligo was used in the LRL method. Sequencing reads are aligned to BC05 or BC06 of the third round ligation adapter against the primary Y-axis. PCR yield is aligned against the secondary Y-axis.

[0035] Figure 11 Third round barcode encoding competition assay results are shown, where the LRL method was used in place of the blocking oligo. Sequencing reads are aligned to BC05 or BC06 of the third round ligation adapter against the primary Y-axis. PCR yield is aligned against the secondary Y-axis.

[0036] Figure 12Third round barcoding competition assay results are shown. Non-barcoded primary third round ligation competing with BC06 was compared, where no blocking oligo, RLL method standard blocking oligo (std), or LRL method alternative blocking oligo (alt) was used. Sequencing reads were aligned to BC05 or BC06 of the third round ligation adapter against the primary Y-axis. PCR yield was aligned against the secondary Y-axis.

[0037] SEQUENCE LISTING

[0038] SEQ ID NOs: 1-18 are oligonucleotides used in the examples. DETAILED DESCRIPTION

[0039] It should be understood that different applications of the disclosed products and methods can be tailored to specific needs in the art. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting.

[0040] All publications, patents, and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any contradictory disclosure.

[0041] Definitions

[0042] Where the indefinite articles "a" or "an" are used, as in "a" or "an" unless specifically stated otherwise, or the definite article "the" is used, this includes plural referents unless the context clearly indicates otherwise. In addition, the terms first, second, third, etc. as used in this description and in the claims are used for clarity not necessarily for describing a sequential or chronological order. It should be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the application described herein are capable of functioning in other sequences than the one described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the application. Unless specifically defined herein, all terms used herein have the same meaning as those which are understood by one of ordinary skill in the art. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016) for definitions and terms of art in the field. The definitions provided herein should not be construed to have a scope less than that understood by one of ordinary skill in the art.

[0043] When referring to measurable values such as amounts, durations, etc., the term "about" as used herein means covering deviations of ± 20% or ± 10%, more preferably ± 5%, even more preferably ± 1% and still more preferably ± 0.1% from the specified value, as such deviations are suitable for performing the disclosed methods.

[0044] A "nucleotide sequence," "DNA sequence," or "nucleic acid molecule" as used herein refers to a polymeric form of nucleotides of any length, whether ribonucleotides or deoxyribonucleotides. The term refers only to the primary structure of the molecule. Thus, the term includes double-stranded and single-stranded DNA, as well as RNA. The term "nucleic acid" as used herein is a single- or double-stranded covalently linked sequence of nucleotides, in which the 3' and 5' ends of each nucleotide are linked by a phosphodiester bond. A polynucleotide can be composed of deoxyribonucleotide bases or ribonucleotide bases. Nucleic acids can be manufactured synthetically or isolated from natural sources. Nucleic acids can further include modified DNA or RNA, for example, DNA or RNA that has been methylated, or RNA that has been subjected to post-translational modifications, for example, 5' capping with 7-methylguanosine, 3' processing such as cleavage and polyadenylation, and splicing. Nucleic acids can also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), and peptide nucleic acid (PNA). The size of a nucleic acid (also referred to herein as a "polynucleotide") is generally expressed as the number of base pairs (bp) or nucleotide pairs of a double-stranded polynucleotide, or in the case of a single-stranded polynucleotide, as the number of nucleotides (nt). One thousand bp or nt equals one kilobase (kb). A polynucleotide having a length of less than about 40 nucleotides is often referred to as an "oligonucleotide," and can comprise primers used in the manipulation of DNA, such as by polymerase chain reaction (PCR).

[0045] In the context of the present disclosure, the term "amino acid" is used in its broadest sense and is intended to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, as well as a side chain (e.g., R group) specific to each amino acid. Amino acid generally refers to a naturally occurring L a-amino acid or residue. The commonly used one- and three-letter abbreviations for naturally occurring amino acids are used herein: A = Ala; C = Cys; D = Asp; E = Glu; F = Phe; G = Gly; H = His; I = lie; K = Lys; L = Leu; M = Met; N = Asn; P = Pro; Q = Gin; R = Arg; S = Ser; T = Thr; V = Val; W = Trp; and Y = Tyr (Lehninger, A.L., (1975) Biochemistry, 2nd Ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids, and chemically modified amino acids (such as amino acid analogs), naturally occurring amino acids that are not typically incorporated into proteins (such as norleucine), and chemically synthesized compounds having properties characteristic of amino acids known in the art (such as beta-amino acids). For example, analogs or mimetics of phenylalanine or proline that allow for the same conformational restrictions of the peptide compound as the natural Phe or Pro are included within the definition of amino acid. Such analogs and mimetics are referred to herein as "functional equivalents" of the corresponding amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5, p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.

[0046] The terms "polypeptide" and "peptide" are used interchangeably herein to refer to polymers of amino acid residues and variants and synthetic analogs thereof. Thus, these terms apply to amino acid polymers in which one or more amino acid residues are synthetic non-naturally occurring amino acids, such as chemical analogs of corresponding naturally occurring amino acids, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes, which can include, but are not limited to, glycosylation, proteolytic cleavage, lipidation, signal peptide cleavage, propeptide cleavage, phosphorylation, and the like. Peptides can be produced using recombinant technology (e.g., by expressing a recombinant or synthetic polynucleotide). Recombinantly produced peptides are typically substantially free of culture medium, e.g., less than about 20%, more preferably less than about 10%, and most preferably less than about 5% culture medium by volume of the protein preparation.

[0047] The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. A protein can be composed of a single polypeptide, or can comprise multiple polypeptides that assemble to form a multimer. The multimer can be a homo-oligomer or a hetero-oligomer. A protein can be a naturally occurring or wild-type protein, or a modified or non-naturally occurring protein. A protein can differ from a wild-type protein, for example, by the addition, substitution, or deletion of one or more amino acids.

[0048] A "variant" of a protein encompasses a peptide, oligopeptide, polypeptide, protein, and enzyme that has amino acid substitutions, deletions, and / or insertions relative to the unmodified or wild-type protein in question, and that has similar biological and functional activities to the unmodified protein from which it is derived. As used herein, the term "amino acid identity" refers to the extent that sequences are identical on an amino acid-by-amino acid basis over a comparison window. Thus, "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the comparison window, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Val, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys, and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the comparison window (i.e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.

[0049] For all aspects and embodiments of the application, a "variant" has at least 50%, 60%, 70%, 80%, 90%, 95%, or 99% overall sequence identity to the amino acid sequence of the corresponding wild-type protein. The sequence identity can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Thus, a sequence can have only 50% overall sequence identity to a full-length reference sequence, but the sequence of a particular region, domain, or subunit can share 80%, 90%, or up to 99% sequence identity with the reference sequence.

[0050] The term "wild-type" refers to a gene or gene product that is isolated from a naturally occurring source. The wild-type gene is the most commonly observed gene in a population and is therefore arbitrarily designed the "normal" or "wild-type" form of the gene. In contrast, the term "modified," "mutant," or "variant" refers to a gene or gene product that displays a modification of sequence (e.g., substitution, truncation, or insertion), post-translational modification, and / or functional property (e.g., altered characteristic) as compared to the wild-type gene or gene product. Notably, naturally occurring mutants can be isolated; these mutants are identified by the fact that they have an altered characteristic when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For example, a methionine (M) can be substituted with an arginine (R) by replacing the codon for methionine (ATG) with the codon for arginine (CGT) at the relevant position in the polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For example, a non-naturally occurring amino acid can be introduced by including a synthetic aminoacyl-tRNA in the IVTT system used to express the mutant monomer. Alternatively, they can be introduced by expressing the mutant monomer in E. coli that is auxotrophic for a particular amino acid in the presence of a synthetic (i.e., non-naturally occurring) analog of that particular amino acid. If the mutant monomer is produced using partial peptide synthesis, it can also be produced by de novo ligation. A conservative substitution replaces an amino acid with one having similar chemical structure, similar chemical properties, or similar side-chain volume. The introduced amino acid can have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge as the amino acid it replaces. Alternatively, a conservative substitution can introduce an aromatic or aliphatic amino acid in place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well known in the art and can be selected according to the properties of the 20 primary amino acids as defined in Table 1 below. Where the amino acids have similar polarity, this can also be determined with reference to the hydrophilicity scale of the side chains of the amino acids in Table 2.

[0051] Table 1 - Chemical properties of amino acids

[0052]

[0053] Table 2 - Hydrophilicity scale

[0054]

[0055] The mutant or modified protein, monomer or peptide can also be chemically modified in any way and at any site. Preferably the mutant or modified monomer or peptide is chemically modified by linking the molecule to one or more cysteines (cysteine ligation), linking the molecule to one or more lysines, linking the molecule to one or more unnatural amino acids, enzymatic modification of the epitope or modification of the terminus. Suitable methods for performing such modifications are well known in the art. The mutant of the modified protein, monomer or peptide can be chemically modified by linkage of any molecule. For example, the mutant of the modified protein, monomer or peptide can be chemically modified by linkage of a dye or fluorophore.

[0056] Methods of the invention

[0057] The present application provides a method of uniquely labeling RNA molecules in a population of cells. The method achieves this by producing a tri-barcode construct from the RNA molecule in step (f).

[0058] In the context of the present application, uniquely labeled RNA molecule is synonymous with producing a uniquely labeled construct transcribed from the RNA molecule. The method is preferably used to produce a uniquely labeled construct transcribed from an RNA molecule in a population of cells. The uniquely labeled construct transcribed from the RNA molecule is preferably a uniquely labeled construct comprising a DNA sequence transcribed from the RNA molecule. The uniquely labeled construct is a tri-barcode construct.

[0059] The population can comprise any number of cells taking into account the discussion below regarding the number of possible barcode combinations. The population for use in the present application preferably comprises at least about 1.00 x 10 4 cells. The population more preferably comprises at least about 1.30 x 10 4 cells, at least about 2.00 x 10 4 cells, at least about 5.00 x 10 4 cells, at least about 1.00 x 10 5 cells, at least about 1.10 x 10 5 cells, at least about 2.00 x 10 5 cells, at least about 5.00 x 10 5 cells, at least about 6.00 x 10 5 cells, at least about 7.00 x 10 51 cell, at least approximately 8.00 x 10 5 1 cell, at least approximately 9.00 x 10 5 1 cell, at least approximately 1.00 x 10 6 1 cell, at least approximately 2.00 x 10 6 1 cell, at least approximately 5.00 x 10 6 1 cell, at least approximately 1.00 x 10 7 1 cell, at least approximately 2.00 x 10 7 At least approximately 5.00 x 10 7 1.00 x 10 cells, at least approximately 1.00 x 10 8 1 cell, at least approximately 2.00 x 10 8 1 cell, at least approximately 5.00 x 10 8 1 cell, at least approximately 1.00 x 10 9 1 cell, at least approximately 2.00 x 10 9 1 cell, at least approximately 3.00 x 10 9 1 cell, at least approximately 5.00 x 10 9 1 cell, at least approximately 1.00 x 10 10 1 cell, at least approximately 2.00 x 10 10 1 cell, at least approximately 4.00 x 10 10 One cell or at least about 5.00 x 10 10 The population may contain even more cells, such as at least about 100 cells. 11 One cell or at least about 10 12 The more starting cells, the greater the transcriptome range in a single cell, two or more cells. More possible barcode combinations are also needed.

[0060] This group preferably contains approximately 5.00 x 10 10 One or fewer cells. More preferably, this population contains cells such as approximately 4.12 x 10⁻⁶. 10 One or fewer cells, approximately 4.10 x 10 10 One or fewer cells, approximately 4.00 x 10 10 One or fewer cells, approximately 3.50 x 10 10 One or fewer cells, approximately 3.00 x 10 10 One or fewer cells, approximately 2.00 x 10 10 One or fewer cells, approximately 1.00 x 10 10 One or fewer cells, approximately 5.00 x 10 9 One or fewer cells, approximately 3.62 x 10 9one or fewer cells, about 3.60 x 10 9 one or fewer cells, about 3.50 x 10 9 one or fewer cells, about 3.00 x 10 9 one or fewer cells, about 2.00 x 10 9 one or fewer cells, about 1.00 x 10 9 one or fewer cells, about 5.00 x 10 8 one or fewer cells, about 2.00 x 10 8 one or fewer cells, about 1.00 x 10 8 one or fewer cells, about 5.66 x10 7 one or fewer cells, about 5.60 x 10 7 one or fewer cells, about 5.50 x 10 7 one or fewer cells, about 5.00x 10 7 one or fewer cells, about 2.00 x 10 7 one or fewer cells, about 1.00 x 10 7 one or fewer cells, about 5.00 x 10 6 one or fewer cells, about 2.00 x 10 6 one or fewer cells, about 1.00 x 10 6 one or fewer cells, about 8.84 x 10 5 one or fewer cells, about 8.80 x 10 5 one or fewer cells, about 8.50 x 10 5 one or fewer cells, about 8.00 x 10 5 one or fewer cells, about 5.00 x 10 5 one or fewer cells, about 2.00 x 10 5 one or fewer cells, about 1.10 x 10 5 one or fewer cells, about 1.00 x 10 5 one or fewer cells, about 5.00 x 10 4 one or fewer cells, about 2.00 x 10 4 one or fewer cells, about 1.38 x 10 4 one or fewer cells, about 1.30 x10 4 one or fewer cells or about 1.00 x 10 4 one or fewer cells.

[0061] The method can be performed in a 24-well plate. The population preferably comprises about 1.38 x 104 about 1.30 x 10 4 about 1.00 x 10 4 about 1.00 x 10

[0062] The method can be performed in a 48-well plate. The population preferably comprises about 1.10 x 10 5 about 1.00 x 10 5 about 5.00 x 10 4 about 2.00 x 10 4 about 1.00 x 10

[0063] The method can be performed in a 96-well plate. The population preferably comprises about 8.84 x 10 5 about 8.80 x 10 5 about 8.50 x 10 5 about 8.00 x 10 5 about 5.00 x 10 5 about 2.00 x 10 5 about 1.00 x 10

[0064] The method can be performed in a 384-well plate. The population preferably comprises about 5.66 x 10 7 about 5.60 x 10 7 about 5.50 x 10 7 about 5.00 x 10 7 about 2.00 x 10 7 about 1.00 x 10 7 about 5.00 x 10 6 about 2.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10

[0065] The method can be performed in a 1536-well plate. The population preferably comprises about 3.62 x 10 9 about 3.60 x 10 9 about 3.50 x 10 9One or fewer, approximately 3.00 x 10 9 One or fewer, approximately 2.00 x 10 9 One or fewer, approximately 1.00 x 10 9 One or fewer, approximately 5.00 x 10 8 One or fewer, approximately 2.00 x 10 8 One or fewer, or approximately 1.00 x 10 8 One or fewer cells. The population may contain any number of cells used in a 24-well, 48-well, 96-well, or 384-well plate.

[0066] This method can be performed in a 3456-well plate. The group preferably contains approximately 4.12 x 10⁻⁶ wells. 10 One or fewer, approximately 4.10 x 10 10 One or fewer, approximately 4.00 x 10 10 One or fewer, approximately 2.00 x 10 10 One or fewer, approximately 1.00 x 10 10 One or fewer, or approximately 5.00 x 10 9 The population may contain any number of cells used in 24-well, 48-well, 96-well, 384-well, or 1536-well plates.

[0067] The cell can be any type of cell. The cell can be a prokaryotic cell. The cell can be a bacterial cell or an archaea cell. The cell is typically a eukaryotic cell. The cell can be a protozoan, algae, fungus, plant, or animal cell. For example, the cell can be a mammalian cell, such as a human, dog, cat, primate, horse, rat, rat, rodent, cow, mouse, pig, or sheep cell. The cell is preferably a human cell. The cell can be a plant cell, such as a cereal, legume, fruit, or vegetable cell. Examples include, but are not limited to, wheat, barley, oats, rapeseed, corn, soybeans, rice, bananas, apples, tomatoes, potatoes, grapes, tobacco, beans, lentils, sugarcane, cocoa, cotton, tea, or coffee cell.

[0068] The animal cells can be derived from ectoderm, endoderm, or mesoderm. The cells can be stem cells, such as embryonic stem cells, induced pluripotent stem cells, or mesenchymal stem cells; bone cells, such as osteoclasts, osteoblasts, or osteocytes; tendon cells, such as tenoblasts or tenocytes, chondrocytes, synoviocytes, vascular cells; blood cells, such as red blood cells, immune cells, platelets, neutrophils, or basophils; muscle cells, such as skeletal muscle cells, cardiomyocytes, or smooth muscle cells; germ cells, such as sperm, oocytes, duct cells, or epididymal cells; secretory cells, adipocytes, hepatocytes, epithelial cells, odontoblasts, cementoblasts, hormone-secreting cells, barrier cells, exocrine secretory epithelial cells, neural cells, astrocytes, oligodendrocytes, or neurons.

[0069] The immune cells can be neutrophils and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophils and precursors, basophils and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.

[0070] The cells can be wild-type or naturally occurring. The cells can be genetically modified or genetically engineered. For example, the immune cells can be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR). The cells can be genetically modified or genetically engineered using transduction or transfection or any other common technique known to one of skill in the art.

[0071] The cells can be healthy cells or obtained from a healthy donor or source. The cells can be diseased or damaged, associated with a disease or injury, or obtained from a diseased or damaged donor or source.

[0072] The cells in the population can be homogenous. The cells can be the same type of cell, or derived from a single donor or source. The cells in the population can be heterogenous. The cells can be a mixture of different types of cells. The cells can be a mixture of different types of cells from the same donor or source. The cells in the population can be derived from different donors or sources.

[0073] The number of cells used in the method and each round is discussed in more detail below. The method typically involves the loss and / or destruction of at least some of the cells in the population. This means that not all of the cells in the population are uniquely labelled using the method. The number of cells at the end of the method is typically lower than the number of cells in the population at the start of the method. Preferably, about 80% or less of the cells in the population, such as about 75% or less, about 70% or less, about 65% or less, about 60% or less, about 55% or less, about 50% or less, about 45% or less, about 40% or less, about 35% or less, about 30% or less, about 25% or less, or about 20% or less of the cells in the population are present at the end of the method. Even fewer cells can be present at the end of the method.

[0074] The method typically involves at least three rounds, as explained in more detail below. The number of cells at the end of each round is typically lower than the number of cells at the start of each round. The number of cells at the end of the first round is typically lower than the number of cells at the start of the first round (i.e. lower than the number of cells in the population). The number of cells at the end of the second round is typically lower than the number of cells at the start of the second round. The number of cells at the end of the third round is typically lower than the number of cells at the start of the third round. Preferably, about 95% or less of the cells at the start of a round, such as about 90% or less, about 85% or less, about 80% or less, about 75% or less, about 70% or less, about 65% or less, or about 60% or less of the cells at the start of a round are present at the end of the round.

[0075] The method uniquely labels RNA molecules in the cell population. This generally means that at the end of the method, RNA molecules from at least about 80% of the cells are labeled with a unique combination of barcodes. At the end of the method, RNA molecules from at least about 80% of the cells are preferably labeled with a unique combination of at least three barcodes. At the end of the method, RNA molecules from at least about 80% of the cells are preferably labeled with a unique tri-barcode. In other words, at the end of the method, RNA molecules from at least about 80% of the cells are labeled with a combination of barcodes that does not exist at the end of the method for any other cell. At the end of the method, RNA molecules from at least about 80% of the cells are preferably labeled with a combination of at least three barcodes that does not exist at the end of the method for any other cell. At the end of the method, RNA molecules from at least about 80% of the cells are preferably labeled with a tri-barcode that does not exist at the end of the method for any other cell. In this paragraph, "does not exist at the end of the method from any other cell" is interchangeable with "is not detectable at the end of the method from any other cell." At the end of the method, RNA molecules from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99% of the cells are preferably labeled in any of these ways. At the end of the method, RNA molecules from about 100% of the cells are preferably labeled in any of these ways.

[0076] The method uniquely labels the constructs produced in step (f). This generally means that, at the end of the method, the constructs from at least about 80% of the cells are labeled with a unique combination of barcodes. At the end of the method, the constructs from at least about 80% of the cells are preferably labeled with a unique combination of at least three barcodes. At the end of the method, the constructs from at least about 80% of the cells are preferably labeled with a unique tri-barcode. In other words, at the end of the method, the constructs from at least about 80% of the cells are labeled with a combination of barcodes that does not exist at the end of the method for any other cell. At the end of the method, the constructs from at least about 80% of the cells are preferably labeled with a combination of at least three barcodes that does not exist at the end of the method for any other cell. At the end of the method, the constructs from at least about 80% of the cells are preferably labeled with a tri-barcode that does not exist at the end of the method for any other cell. In this paragraph, “does not exist at the end of the method for any other cell” is interchangeable with “is not detectable at the end of the method for any other cell.” At the end of the method, the constructs from at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99% of the cells are preferably labeled in any of these ways. At the end of the method, the constructs from about 100% of the cells are preferably labeled in any of these ways.

[0077] The method uniquely labels the RNA molecules in the population of cells. This generally means that, at the end of the method, the RNA molecules from each cell are labeled with a unique combination of barcodes. At the end of the method, the RNA molecules from each cell are preferably labeled with a unique combination of at least three barcodes. At the end of the method, the RNA molecules from each cell are preferably labeled with a unique tri-barcode. In other words, at the end of the method, the RNA molecules from each cell are labeled with a combination of barcodes that does not exist at the end of the method for any other cell. At the end of the method, the RNA molecules from each cell are preferably labeled with a combination of at least three barcodes that does not exist at the end of the method for any other cell. At the end of the method, the RNA molecules from each cell are preferably labeled with a tri-barcode that does not exist at the end of the method for any other cell. In this paragraph, “does not exist at the end of the method for any other cell” is interchangeable with “is not detectable at the end of the method for any other cell.”

[0078] The method uniquely labels the constructs produced in step (f). This generally means that, at the end of the method, the constructs from each cell are labeled with a unique combination of barcodes. At the end of the method, the constructs from each cell are preferably labeled with a unique combination of at least three barcodes. At the end of the method, the constructs from each cell are preferably labeled with a unique tri-barcode. In other words, at the end of the method, the constructs from each cell are labeled with a combination of barcodes that does not exist for any other cell at the end of the method. At the end of the method, the constructs from each cell are preferably labeled with a combination of at least three barcodes that does not exist for any other cell at the end of the method. At the end of the method, the constructs from each cell are preferably labeled with a tri-barcode that does not exist for any other cell at the end of the method. In this paragraph, “does not exist for any other cell at the end of the method” is interchangeable with “is not detectable for any other cell at the end of the method”.

[0079] At the end of step (f), the tri-barcode for at least 80% of the cells is generally unique. At the end of the method, the tri-barcode for at least about 80% of the cells is preferably not present in any other cell. At the end of the method, the tri-barcode for at least about 85%, at least 90%, at least about 95%, at least about 97%, at least about 98%, or at least about 99% of the cells is preferably unique or not present in any other cell. At the end of step (f), the tri-barcode construct from each cell at the end of the method generally comprises a unique tri-barcode. At the end of the method, the tri-barcode construct from each cell preferably comprises a tri-barcode that does not exist for any other cell at the end of the method. In this paragraph, “does not exist for any other cell at the end of the method” is interchangeable with “is not detectable for any other cell at the end of the method”.

[0080] Methods for achieving unique barcoding of cells will be discussed in more detail below.

[0081] First round

[0082] The first round of the method of the application comprises steps (a) and (b). Step (a) comprises dividing the population of cells into a plurality of first samples. The population can be divided into any number of first samples. The population is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000, or at least about 5000 first samples. The population is preferably divided into about 24, about 48, about 96, about 384, about 1536, or about 3456 first samples.

[0083] The number of cells in each first sample will generally depend on the number of cells in the starting population as described above and the number of first samples. Each first sample will generally comprise at least about 5.00 x 10 2 cells. Each first sample will more preferably comprise at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x 10 4 cells, at least about 2.00 x 10 4 cells, at least about 5.00 x 10 4 cells, at least about 1.00 x 10 5 cells, at least about 1.40 x 10 5 cells, at least about 2.00 x 10 5 cells, at least about 5.00 x 10 5 cells, at least about 1.00 x 10 6 cells, at least about 2.00 x 10 6 cells, at least about 2.30 x 10 6 cells, at least about 5.00 x 10 6 cells, at least about 1.00 x 107 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 8 about 1.00 x 10 9 about 1.00 x 10

[0084] about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 4 about 1.00 x 10 4 about 1.00 x 10 4 about 1.00 x 10 3 about 1.00 x 10 3 about 1.00 x 103 about 5.00 x 10 3 about 2.30 x 10 3 about 2.00 x 10 3 about 1.00 x 10 3 about 5.76 x 10 2 about 5.70 x 10 2 about 5.00 x 10 2 about 5.00 x 10

[0085] The method can be performed in a 24-well plate. Each first sample preferably comprises about 576 cells or fewer, about 570 cells or fewer, about 550 cells or fewer, or about 500 cells or fewer.

[0086] The method can be performed in a 48-well plate. Each first sample preferably comprises about 2304 cells or fewer, about 2300 cells or fewer, or about 2000 cells or fewer. Each first sample can comprise any of the number of cells for a 24-well plate.

[0087] The method can be performed in a 96-well plate. Each first sample preferably comprises about 9216 cells or fewer, about 9200 cells or fewer, or about 9000 cells or fewer. Each first sample can comprise any of the number of cells for a 24-well plate or a 48-well plate.

[0088] The method can be performed in a 384-well plate. Each first sample preferably comprises about 147,456 cells or fewer, about 147,000 cells or fewer, about 1450,000 cells or fewer, or about 140,000 cells or fewer. Each first sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, or a 96-well plate.

[0089] The method can be performed in a 1536-well plate. Each first sample preferably comprises about 2,359,296 cells or fewer, about 2,300,000 cells or fewer, or about 2,000,000 cells or fewer. Each first sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, a 96-well plate, or a 384-well plate.

[0090] The method can be performed in a 3456 well plate. Each first sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each first sample can comprise any of the number of cells for a 24 well plate, a 48 well plate, a 96 well plate, a 384 well plate, or a 1536 well plate.

[0091] The number of cells in each first sample is typically approximately or about the same. The skilled person will understand that the number of cells can vary between samples based on standard techniques for measuring the number of cells and dividing the cells into different samples.

[0092] It is step (d) of the method of the application and comprises reverse transcribing the RNA molecules in the plurality of second samples. This is the “R” in the LRL method of the application. The method of the application preferably does not comprise reverse transcription (RT) prior to step (a). The method of the application preferably comprises no reverse transcription (RT) prior to step (a). The method of the application is preferably performed on RNA molecules that have not been reverse transcribed. The method of the application is preferably used to uniquely label RNA molecules that have not been reverse transcribed in the cell population.

[0093] Step (b) comprises ligating a first adaptor to the 3’ end of the RNA molecules in the plurality of first samples. Any method of ligation can be used according to the application, including those described in the examples. The skilled person understands ligases that can be used in the application, such as T4 DNA ligase and other DNA ligases. Suitable conditions for ligation reactions are known in the art and described in the examples.

[0094] The first adaptor comprises a primer site for reverse transcription (RT). Any RT primer site can be used according to the application, including those described in the examples. The RT primer site is typically a sequence in the first adaptor that specifically hybridises to a part or region of the RT primer used in the second round.

[0095] An RT primer site is said to "specifically hybridize" to a portion or region of an RT primer when it hybridizes to that portion or region with preferential or high affinity, but does not substantially hybridize, does not hybridize, or hybridizes only with low affinity to other polynucleotide sequences, especially other RNA molecules, adaptors or sequences used in the present application. Conditions permitting hybridization are well known in the art (e.g., Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rdedition, Cold Spring Harbour Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., eds., Greene Publishing and Wiley-lnterscience, New York (1995)). Hybridization can be performed under low stringency conditions, for example in the presence of a buffered solution of 30% to 35% formamide, 1 M NaCl and 1% sodium dodecyl sulfate (SDS) at 37°C, followed by 20 washes in 1X (0.1650 M Na+) to 2X (0.33 M Na+) standard sodium citrate (SSC) at 50°C. Hybridization can be performed under medium stringency conditions, for example in the presence of a buffered solution of 40% to 45% formamide, 1 M NaCl and 1% SDS at 37°C, followed by 0.5X (0.0825 M Na+) to 1X (0.1650 M Na+) SSC at 55°C. Hybridization can be performed under high stringency conditions, for example in the presence of a buffered solution of 50% formamide, 1 M NaCl, 1% SDS at 37°C, followed by 0.1X (0.0165 M Na+) SSC at 60°C. An RT primer site will "specifically hybridize" if it hybridizes to a portion or region of an RT primer with a melting temperature (Tm) that is at least 2°C, such as at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, or at least 10°C, higher than its Tm to other polynucleotide sequences. More preferably, an RT primer site hybridizes to a portion or region of an RT primer with a Tm that is at least 2°C, such as at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, at least 10°C, at least 20°C, at least 30°C, or at least 40°C, higher than its Tm to other polynucleotide sequences.Preferably, the RT primer site hybridizes with the portion or region of the RT primer at a Tm that is at least 2°C, such as at least 3°C, at least 4°C, at least 5°C, at least 6°C, at least 7°C, at least 8°C, at least 9°C, at least 10°C, at least 20°C, at least 30°C, or at least 40°C higher than its Tm for a polynucleotide that differs from the portion or region of the RT primer by one or more nucleotides, such as by 1, 2, 3, 4, or 5 or more nucleotides. The RT primer site typically hybridizes with the portion or region of the RT primer at a Tm of at least 90°C, such as at least 92°C or at least 95°C. The Tm can be experimentally measured using known techniques, including using DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available on the internet.

[0096] The RT primer site typically comprises or consists of a sequence that is at least about 80% identical or homologous to the reverse complement of the portion or region of the RT primer. The RT primer site preferably comprises or consists of a sequence that is at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99% identical or homologous to the reverse complement of the portion or region of the RT primer. The RT primer site most preferably comprises or consists of the sequence that is the reverse complement of the portion or region of the RT primer. By complement, it is meant that the RT primer site comprises or consists of a sequence that is 100% identical or homologous to the reverse complement of the portion or region of the RT primer. Identity is typically determined using typical Watson-Crick base pairing.

[0097] Homology can be determined using standard methods in the art. For example, the UWGCG package provides the BESTFIT program which can be used to calculate homology, for example on its default settings (Devereux et al. (1984) Nucleic Acids Research 12, pp. 387-395). The PILEUP and BLAST algorithms can be used to calculate homology or align sequences such as to identify equivalent or corresponding residues (typically on their default settings), as described in Altschul S. F. (1993) J Mol Evol 36:290-300; Altschul, S.F. et al. (1990) J Mol Biol 215:403-10. Software for performing BLAST analyses is publicly available at the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ).

[0098] The RT primer site and / or the portion or region of the RT primer can be any length, such as at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 21 nucleotides, at least about 25 nucleotides, or at least about 25 nucleotides in length. The RT primer site and the portion or region of the RT primer are preferably the same length.

[0099] The first adaptors used in each sample can use the same RT primer site. This allows the method to use a standard set of first adaptors and RT primers, with only the first barcode and the second barcode being different between samples.

[0100] The first adaptors used in each first sample comprise a different reverse complement of the first barcode. The first adaptors comprise a different reverse complement of the first barcode because the reverse complement is reversed transcribed in the second round to produce the first different barcode. The first adaptors used in each first sample comprise a unique reverse complement of the first barcode. The RNA molecules in each first sample are labeled with a different or unique first barcode. The first adaptors used in each first sample comprise a reverse complement of the first barcode that is different from the first barcode used in all other first samples. The RNA molecules in each first sample are labeled with a first barcode that is different from the first barcode used to label the RNA molecules in all other first samples. The first adaptors used in each first sample comprise a reverse complement of the first barcode that is different from any or all of the other barcodes used in the method. The RNA molecules in each first sample are labeled with a first barcode that is different from any or all of the other barcodes used in the method. Barcodes are discussed in more detail below.

[0101] The first adaptor is typically a polynucleotide. The first adaptor can be any type of polynucleotide. A polynucleotide, such as a nucleic acid, is a macromolecule comprising two or more nucleotides. A polynucleotide can be single-stranded or double-stranded. A double-stranded polynucleotide is made from two single-stranded polynucleotides hybridized together. The first adaptor can be a single-stranded polynucleotide or a double-stranded polynucleotide.

[0102] A polynucleotide can comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. A nucleotide typically contains a nucleobase, a sugar, and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines, and more specifically include adenine (A), guanine (G), thymine (T), uracil (U), and cytosine (C).

[0103] The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG), and deoxycytidine (dC).

[0104] The nucleotide is typically a ribonucleotide or a deoxyribonucleotide. The nucleotide typically contains a mono-, di-, or tri-phosphate. The nucleotide can comprise more than three phosphates, such as 4 or 5 phosphates. The phosphates can be linked on the 5' or 3' side of the nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5-hydroxymethylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP), and deoxymethylcytidine monophosphate. The nucleotide is preferably selected from the group consisting of AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP, and dUMP.

[0105] The nucleotide can be abasic (i.e., lack a nucleobase). The nucleotide can also lack both a nucleobase and a sugar (i.e., be a C3 spacer).

[0106] The nucleotides in the polynucleotide can be linked to one another in any manner. The nucleotides are typically linked by their sugars and phosphate groups, as in nucleic acids. The nucleotides can be linked by their nucleobases, as in pyrimidine dimers.

[0107] The polynucleotide can be a nucleic acid, such as a deoxyribonucleic acid (DNA) or a ribonucleic acid (RNA). The polynucleotide can comprise one RNA strand hybridized to one DNA strand. The polynucleotide can be any synthetic nucleic acid known in the art, such as a peptide nucleic acid (PNA), a glycerol nucleic acid (GNA), a threose nucleic acid (TNA), a locked nucleic acid (LNA), a bridged nucleic acid (BNA), or other synthetic polymer with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose units linked together by phosphodiester bonds. The LNA is formed from ribonucleotides with a bridge as discussed above that links the 2' oxygen and the 4' carbon in the ribose moiety.

[0108] The polynucleotide is preferably DNA, RNA, or a DNA or RNA hybrid, most preferably DNA. The DNA / RNA hybrid can contain DNA and RNA on the same strand. Preferably, the DNA / RNA hybrid contains one DNA strand hybridized to an RNA strand.

[0109] The backbone of the polynucleotide can be altered to reduce the likelihood of strand breakage. For example, DNA is known to be more stable than RNA under many conditions. The backbone of the polynucleotide strand can be modified to avoid damage caused by, for example, aggressive chemicals such as free radicals. DNA or RNA containing non-natural or modified bases can be produced by amplifying a natural DNA or RNA polynucleotide using an appropriate polymerase in the presence of modified NTPs.

[0110] The first adaptor is typically synthetic or semi-synthetic. For example, the DNA or RNA can be purely synthetic, synthesized by conventional DNA synthesis methods such as phosphoramidite-based chemistry. Synthetic polynucleotide subunits can be joined together by known means, such as ligation or chemical bonding, to produce longer strands. Internal self-forming structures (e.g., hairpins, quadruplexes) can be designed into the object, for example, by ligating appropriate sequences. Synthetic polynucleotides can be replicated and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, etc.

[0111] The first adaptor can be of any length. For example, the first adaptor can be at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 35, at least about 40, at least about 45, at least about 50, at least about 55, at least about 56 nucleotides, at least about 60, or at least 60 nucleotides or nucleotide pairs in length.

[0112] The first adaptor is preferably double-stranded. The first adaptor is preferably double-stranded DNA. The first adaptor can be referred to as a barcoded cDNA reverse transcription adaptor (barcoded CDRA).

[0113] The first adaptor is preferably double-stranded and comprises an overhang capable of hybridizing or specifically hybridizing to the 3' end of the RNA molecule. The non-overhanging strand of such an adaptor is typically ligated to the RNA molecule. The skilled person is able to design an overhang capable of hybridizing or specifically hybridizing to the 3' end of the RNA molecule. For example, the first adaptor used in each first sample is preferably double-stranded and comprises an overhang comprising all possible combinations of sequences based on nucleotides comprising adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). Such overhangs are capable of hybridizing or specifically hybridizing to the 3' end of the RNA molecule. The first adaptor comprising the overhang can be generated randomly to include combinations of sequences. Step (a) preferably comprises hybridizing or specifically hybridizing the overhang of the double-stranded first adaptor to the 3' end of the RNA molecule in the plurality of first samples, ligating the non-overhanging strand of the first adaptor to the 3' end of the RNA molecule, wherein the first adaptor comprises a primer site for reverse transcription (RT), and wherein the first adaptor used in each first sample comprises the reverse complement of a different first barcode.

[0114] Eukaryotic RNA molecules typically have a poly(A) tail (see Figure 1 ). The first adaptor is preferably double-stranded and comprises an overhang capable of hybridizing or specifically hybridizing to the poly(A) tail of the RNA molecule. The non-overhanging strand of such an adaptor is typically ligated to the RNA molecule. An example of this is shown in Figure 1 . Step (a) preferably comprises hybridizing or specifically hybridizing the overhang of the double-stranded first adaptor to the poly(A) tail of the RNA molecule in the plurality of first samples, ligating the non-overhanging strand of the first adaptor to the poly(A) tail of the RNA molecule, wherein the first adaptor comprises a primer site for reverse transcription (RT), and wherein the first adaptor used in each first sample comprises the reverse complement of a different first barcode.

[0115] Non-eukaryotic RNA molecules typically do not have a poly(A) tail. Step (b) can further comprise polyadenylating the 3' end of the RNA molecule and using a double-stranded first adaptor comprising an overhang capable of hybridizing or specifically hybridizing to the polyadenylated 3' end of the RNA molecule. The non-overhang strand of such an adaptor is typically ligated to the RNA molecule. Step (a) can comprise: polyadenylating the 3' end of the RNA molecule; hybridizing or specifically hybridizing the overhang of a double-stranded first adaptor to the polyadenylated 3' end of the RNA molecule in the plurality of first samples; ligating the non-overhang strand of the first adaptor to the polyadenylated 3' end of the RNA molecule, wherein the first adaptor comprises a primer site for reverse transcription (RT), and wherein the first adaptor used in each first sample comprises the reverse complement of a different first barcode. The RNA molecule can be polyadenylated using any known method, such as using E. coli Poly(A) polymerase (EPAP).

[0116] The overhang can be any length. For example, the overhang can be at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 15, at least about 20, or at least about 25 nucleotides in length. The overhang is preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. Hybridization and specific hybridization are discussed above, and any of these embodiments apply to the overhang to the 3' end, poly(A) tail, or polyadenylated 3' end.

[0117] The overhang preferably comprises one or more of: (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides, and (iii) universal nucleotides. Overhangs comprising (i) and (ii) are capable of specifically hybridizing to a poly(A) tail or polyadenylated 3' end as described above. Nucleotides are defined above. A universal nucleotide is a nucleotide that hybridizes or binds to all nucleotides in a polynucleotide to some extent. A universal nucleotide is preferably a nucleotide that hybridizes or binds to nucleotides comprising nucleosides adenine (A), thymine (T), uracil (U), guanine (G), and cytosine (C) to some extent. A universal nucleotide can hybridize or bind to some nucleotides with greater strength compared to the strength of hybridization or binding to other nucleotides. For example, a universal nucleotide comprising the nucleoside 2'-deoxyinosine (I) will show a preference pairing order of I-C > I-A > I-G, approximately = I-T.

[0118] The universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole, nitroimidazole, 4-nitro pyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4- aminobenzimidazole, or phenyl (C6-aromatic ring). The universal nucleotide more preferably comprises one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'-deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methyl inosine, 4-nitroindole 2'-deoxynucleoside, 4-nitroindole nucleoside, 5-nitroindole 2'-deoxynucleoside, 5-nitroindole nucleoside, 6-nitroindole 2'-deoxynucleoside, 6-nitroindole nucleoside, 3-nitropyrrole 2'-deoxynucleoside, 3-nitropyrrole nucleoside, non-ring sugar analog of hypoxanthine, nitroimidazole 2'-deoxynucleoside, nitroimidazole nucleoside, 4-nitro pyrazole 2'-deoxynucleoside, 4-nitro pyrazole nucleoside, 4-nitrobenzimidazole 2'-deoxynucleoside, 4-nitrobenzimidazole nucleoside, 5-nitroindazole 2'-deoxynucleoside, 5-nitroindazole nucleoside, 4- aminobenzimidazole 2'-deoxynucleoside, 4-aminobenzimidazole nucleoside, phenyl C-nucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'-deoxyxanthosine, 2'-deoxyisoguanosine, K-2'-deoxynucleoside, P-2'-deoxynucleoside, and pyrrolidine. The universal nucleotide more preferably comprises 2'-deoxyinosine. The universal nucleotide more preferably is IMP or dIMP. The universal nucleotide most preferably is dPMP (2'-deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2,6-diaminopurine monophosphate).

[0119] The overhang preferably comprises one or more consecutive repeat units of UTT (in the 5' to 3' direction), such as 2 or more, 3 or more, 4 or more, or 5 or more consecutive repeat units of UTT.

[0120] The first adapter used in each sample can use the same overhang. This allows the method to use a standard set of first adapters, where only the first barcode differs between first samples.

[0121] If the first adapter is double stranded, the reverse complement of the different first barcode is preferably present in the strand of the first adapter that is ligated to the RNA molecule. The reverse complement of the different first barcode is preferably present in the strand opposite to the strand that has the overhang that hybridizes or specifically hybridizes to the 3' end, poly(A) tail, or polyadenylated 3' end of the RNA molecule. If the reverse complement of the different first barcode is ligated to the RNA molecule, the different first barcode will be present in the reverse transcribed cDNA product in the second round (see, e.g., Figure 2 ).

[0122] The nucleotide at the 3' end of the first adaptor is preferably a dideoxycytosine (ddC) or an inverted thymine (inverted dT). The nucleotide at the 3' end of one of the strands of the double-stranded first adaptor is preferably a ddC or an inverted dT. The nucleotide at the 3' end of the strand of the double-stranded first adaptor is preferably a ddC or an inverted dT. These nucleotides prevent DNA polymerase from further extending the cDNA sequence. They also protect the oligonucleotide from 3' exonuclease cleavage. The use of these types of modifications in the adaptor, especially on the strand that is not directly incorporated into the cDNA of the second round of synthesis (see, e.g., Figure 1 and Figure 2 ), prevents the occurrence of unexpected cross-talk and mispriming during amplification of the cDNA.

[0123] An exemplary first adaptor is shown in the example.

[0124] The non-ligated strand of the first adaptor is preferably removed prior to step (d). If the first adaptor is double-stranded and comprises an overhang capable of hybridizing or specifically hybridizing to the 3' end, poly(A) tail, or polyadenylated 3' end of the RNA molecule, the overhang strand of the adaptor is preferably removed prior to step (d). Any suitable method, including digestion, can be used to remove the non-ligated / overhang strand. The strand is preferably removed using USER digestion. USER enzyme mix is a mixture of uracil DNA glycosylase (UDG) and DNA glycosylase-lyase Endonuclease VIII. As explained above, this reduces the possibility of re-barcoding or cross-talk upon pooling.

[0125] The second round

[0126] The second round of the method of the application comprises steps (c) and (d). Step (c) comprises pooling the plurality of first samples and dividing the pool of cells into a plurality of second samples. The pool can be divided into any number of second samples. The pool is preferably divided into any number of second samples discussed above with respect to the first samples. The pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000, or at least about 5000 second samples. Each second sample is preferably divided into about 24, about 48, about 96, about 384, about 1536, or about 3456 second samples. The number of second samples is typically the same as the number of first samples.

[0127] Each second sample typically comprises at least about 5.00 x 10 2 cells. Each second sample more preferably comprises at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x 10 4 cells, at least about 2.00 x 10 4 cells, at least about 5.00 x 10 4 cells, at least about 1.00 x 10 5 cells, at least about 1.40 x 10 5 cells, at least about 2.00 x 10 5 cells, at least about 5.00 x 10 5 cells, at least about 1.00 x 10 6 cells, at least about 2.00 x 10 6 cells, at least about 2.30 x 10 6 cells, at least about 5.00 x 106 about 1.00 x 10 7 about 1.10 x 10 7 about 2.00 x 10 7 about 3.00 x 10 7 about 5.00 x 10 7 about 10 8 about 10 9 about 10

[0128] about 5.00 x 10 7 about 10 7 about 10 7 about 10 7 about 10 7 about 10 7 about 10 7 about 10 6 about 10 6 about 10 6 about 10 6 about 10 6 about 10 5 about 10 5 about 10 5 about 10 5 about 10 5 about 10 5 about 10 5 about 10 4 about 10 4 about 10 4 about 10 3 about 103 about 9.00 x 10 3 about 5.00 x 10 3 about 2.30 x 10 3 about 2.00 x 10 3 about 1.00 x 10 3 about 5.76 x 10 2 about 5.70 x 10 2 about 5.00 x 10 2 about 5.00 x 10

[0129] The method can be performed in a 24-well plate. Each second sample preferably comprises about 576 cells or fewer, about 570 cells or fewer, about 550 cells or fewer, or about 500 cells or fewer.

[0130] The method can be performed in a 48-well plate. Each second sample preferably comprises about 2304 cells or fewer, about 2300 cells or fewer, or about 2000 cells or fewer. Each second sample can comprise any of the number of cells for a 24-well plate.

[0131] The method can be performed in a 96-well plate. Each second sample preferably comprises about 9216 cells or fewer, about 9200 cells or fewer, or about 9000 cells or fewer. Each second sample can comprise any of the number of cells for a 24-well plate or a 48-well plate.

[0132] The method can be performed in a 384-well plate. Each second sample preferably comprises about 147,456 cells or fewer, about 147,000 cells or fewer, about 1450,000 cells or fewer, or about 140,000 cells or fewer. Each second sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, or a 96-well plate.

[0133] The method can be performed in a 1536-well plate. Each second sample preferably comprises about 2,359,296 cells or fewer, about 2,300,000 cells or fewer, or about 2,000,000 cells or fewer. Each second sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, a 96-well plate, or a 384-well plate.

[0134] The method can be performed in a 3456 well plate. Each second sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each second sample can comprise any of the number of cells for a 24 well plate, a 48 well plate, a 96 well plate, a 384 well plate, or a 1536 well plate.

[0135] The number of cells in each second sample is typically approximately or about the same. The number of cells in each second sample is typically approximately or about the same as the number of cells in each first sample. The skilled person will appreciate that the number of cells between samples can vary based on standard techniques for measuring the number of cells and dividing the cells into different samples.

[0136] Step (d) comprises reverse transcribing the RNA molecules in the plurality of second samples. Methods for performing reverse transcription are well known in the art and any suitable conditions can be used. Preferred conditions are described in the examples. Reverse transcription involves the use of a reverse transcriptase to convert RNA into cDNA. Reverse transcriptases are commercially available (e.g. Maxima H minus Reverse Transcriptase (ThermoFisher Scientific, EP0753), Superscript® II Reverse Transcriptase (Invitrogen) and Affinity script (Agilent)). The reverse transcriptase is preferably Maxima H minus Reverse Transcriptase (ThermoFisher Scientific, EP0753). The double stranded construct produced in step (b) typically comprises RNA molecules hybridised to a cDNA strand produced by reverse transcription.

[0137] Step (d) uses the primer site in the first adaptor and a RT primer to form the double stranded construct. The skilled person is able to design suitable RT primers for use in the present application. The RT primer is typically a polynucleotide RT primer and can be any of the polynucleotides discussed above in relation to the first adaptor. The RT primer is typically synthetic or semi-synthetic.

[0138] The RT primer is preferably a single stranded polynucleotide RT primer. The RT primer can be any length. For example, the length of the RT primer is preferably at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 32 nucleotides, at least about 35 nucleotides, or at least about 40 nucleotides. The length of the RT primer is preferably less than about 50 nucleotides, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length.

[0139] Step (d) preferably comprises hybridising the RT primer to the first adaptor. Step (d) preferably comprises hybridising the RT primer to the RT primer site in the first adaptor. Step (d) preferably comprises hybridising the RT primer to the first adaptor or the RT primer site in the first adaptor and using the primer site and the RT primer to reverse transcribe the RNA molecules in the plurality of second samples to form double stranded constructs. The hybridisation is preferably specific hybridisation. As explained above, the RT primer site is typically a sequence in the first adaptor that specifically hybridises to a portion or region of the RT primer used in the second round. The RT primer site and the portion or region of the RT primer are discussed above with reference to the first round. The RT primer can comprise any of the portions or regions discussed above. The RT primer can use the same portion or region that specifically hybridises to the RT primer site. This allows the method to use a standard set of RT primers, with only the second barcode differing between second samples.

[0140] The RT primer used in each second sample comprises a different second barcode. The RT primer used in each second sample comprises a unique second barcode. The RNA molecules in each second sample are labelled with a different or unique second barcode. The RT primer used in each second sample comprises a second barcode that is different to the second barcode used in all other second samples. The RNA molecules in each second sample are labelled with a second barcode that is different to the second barcode used to label the RNA molecules in all other second samples. The RT primer used in each second sample comprises a second barcode that is different to any or all other barcodes used in the method. The RNA molecules in each second sample are labelled with a second barcode that is different to any or all other barcodes used in the method. Barcodes are discussed in more detail below.

[0141] The RT primer preferably comprises a sequence at or near its 5' end that is capable of specifically hybridizing to the second adaptor used in step (f). The sequence preferably specifically hybridizes to the overhang in the second adaptor, as discussed in more detail below. Specific hybridization is defined above. The sequence can be of any length. The length of the sequence is preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides. The length of the sequence is preferably less than about 50 nucleotides, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. The RT primer can use the same sequence at or near its 5' end that is capable of specifically hybridizing to the second adaptor used in step (f). This allows the method to use a standard set of RT primers, with only the second barcode differing between second samples.

[0142] Exemplary RT primers are shown in the examples.

[0143] The RT in step (d) preferably produces strands comprising a different first barcode, a different second barcode, and a sequence complementary to the RNA molecule. The different first barcode is transcribed from the reverse complement in the first adaptor. The different second barcode is transcribed from the reverse complement in the second adaptor. Figure 2 Examples of this are shown in the examples. The components are preferably in the following order 5' to 3' in the strand: the different second barcode, the different first barcode, and the complementary sequence. These strand constructs are double barcoded.

[0144] Step (d) preferably further comprises: (i) hybridizing a strand switch primer (SSP) to the overhang at the 3' end of the strand in the double stranded construct resulting from reverse transcription of the RNA molecule; and (ii) reverse transcribing the SSP to extend the double stranded construct. Examples of this are shown in the examples. Figure 2An example of this is shown in the Examples. The SSP preferably hybridizes specifically to the overhang. The strand in the double-stranded construct produced by reverse transcription is preferably the cDNA strand. Overhangs, hybridization, and specific hybridization are discussed above, and any of these embodiments are applicable to the SSP embodiments. The overhang preferably comprises contiguous cytosine-containing nucleotides, such as contiguous deoxycytidine (dC)-containing nucleotides. Such overhangs can be produced by the addition of non-template dC to the synthesized cDNA by the reverse transcriptase when the cDNA reaches the end of the RNA template. The SSP can be any type of polynucleotide and / or any length discussed above with reference to the RT primers used in the application. The SSP preferably comprises one or more unique molecular identifiers (UMIs) to facilitate characterization. The SSP can comprise any number of UMIs, such as 2 or more, 3 or more, 4 or more, 5 or more, or 10 or more. An example of a UMI is TTVVVVVTT. This UMI structure is optimized for characterization using a nanopore by reducing homopolymers. The SSP facilitates deduplication of PCR replicates. By including one or more UMIs in the SSP, they can be removed from the second adapter used in the third round (unlike the RLL method) to save on production costs and complexity of the second adapter. Any reverse transcription method can be used in step (ii), including any of the methods discussed above.

[0145] An exemplary SSP is shown in the Examples.

[0146] Third round

[0147] The third round of the method of the application comprises steps (e) and (f). Step (e) comprises pooling the plurality of second samples and dividing the pool of cells into a plurality of third samples. The pool can be divided into any number of third samples. The pool is preferably divided into any number of third samples discussed above with respect to the first samples. The pool is preferably divided into at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000, or at least about 5000 third samples. Each third sample is preferably divided into about 24, about 48, about 96, about 384, about 1536, or about 3456 third samples. The number of third samples is typically the same as the number of first and / or second samples. The number of first samples, the number of second samples, and the number of third samples are preferably the same.

[0148] Each third sample typically comprises at least about 5.00 x 10 2 cells. Each third sample more preferably comprises at least about 5.70 x 10 2 cells, at least about 1.00 x 10 3 cells, at least about 2.00 x 10 3 cells, at least about 2.30 x 10 3 cells, at least about 5.00 x 10 3 cells, at least about 9.00 x 10 3 cells, at least about 9.20 x 10 3 cells, at least about 1.00 x 10 4 cells, at least about 2.00 x 10 4 cells, at least about 5.00 x 10 4 cells, at least about 1.00 x 10 5 cells, at least about 1.40 x 10 5 cells, at least about 2.00 x 10 5 cells, at least about 5.00 x 10 5 cells, at least about 1.00 x 10 6 cells, at least about 2.00 x 10 6 cells, at least about 2.30 x 106 about 1.00 x 10 6 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 8 about 1.00 x 10 9 about 1.00 x 10

[0149] about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 7 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 6 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 5 about 1.00 x 10 4 about 1.00 x 10 4 about 1.00 x 10 4 about 1.00 x 103 about 9.20 x 10 3 about 9.00 x 10 3 about 5.00 x 10 3 about 2.30 x 10 3 about 2.00 x 10 3 about 1.00 x 10 3 about 5.76 x10 2 about 5.70 x 10 2 about 5.00 x 10 2 about 2.00 x 10

[0150] The method can be performed in a 24-well plate. Each third sample preferably comprises about 576 cells or fewer, about 570 cells or fewer, about 550 cells or fewer, or about 500 cells or fewer.

[0151] The method can be performed in a 48-well plate. Each third sample preferably comprises about 2304 cells or fewer, about 2300 cells or fewer, or about 2000 cells or fewer. Each third sample can comprise any of the number of cells for a 24-well plate.

[0152] The method can be performed in a 96-well plate. Each third sample preferably comprises about 9216 cells or fewer, about 9200 cells or fewer, or about 9000 cells or fewer. Each third sample can comprise any of the number of cells for a 24-well plate or a 48-well plate.

[0153] The method can be performed in a 384-well plate. Each third sample preferably comprises about 147,456 cells or fewer, about 147,000 cells or fewer, about 1450,000 cells or fewer, or about 140,000 cells or fewer. Each third sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, or a 96-well plate.

[0154] The method can be performed in a 1536-well plate. Each third sample preferably comprises about 2,359,296 cells or fewer, about 2,300,000 cells or fewer, or about 2,000,000 cells or fewer. Each third sample can comprise any of the number of cells for a 24-well plate, a 48-well plate, a 96-well plate, or a 384-well plate.

[0155] The method can be performed in a 3456 well plate. Each third sample preferably comprises about 11,943,936 or fewer, about 11,500,000 or fewer, about 11,000,000 or fewer, or about 10,000,000 or fewer cells. Each third sample can comprise any of the number of cells for a 24 well plate, a 48 well plate, a 96 well plate, a 384 well plate, or a 1536 well plate.

[0156] The number of cells in each third sample is typically approximately or about the same. The number of cells in each third sample is typically approximately about the same as the number of cells in each first sample and / or the number of cells in each second sample. The number of cells in each first sample, each second sample, and each third sample is typically approximately or about the same. The skilled person will appreciate that the number of cells between samples can vary based on standard techniques for measuring the number of cells and dividing the cells into different samples.

[0157] Step (f) comprises ligating a second adaptor to the double stranded construct in the plurality of third samples to form a triple barcoded construct. The triple barcoded construct comprises a DNA sequence transcribed from the RNA molecule. Ligation is discussed above and any of the methods above can be used. The skilled person is able to design a suitable second adaptor for use in the present application. The second adaptor is typically a polynucleotide adaptor and can be any of the polynucleotides discussed above with reference to the first adaptor. The second adaptor is typically synthetic or semi-synthetic.

[0158] The second adaptor is preferably a double stranded polynucleotide adaptor. The second adaptor can be any length. The length of the second adaptor is preferably at least about 10, at least about 15, at least about 20, at least about 25, at least about 30, at least about 32, at least about 35, at least about 40, at least about 45, at least 47, or at least about 50 nucleotides or nucleotide pairs.

[0159] The second adaptor is preferably ligated to the strand in the double stranded construct that comprises a different first barcode, a different second barcode, and a complement sequence. This allows a third different barcode to be added to the double barcoded construct produced in the second round.

[0160] Step (f) preferably comprises hybridizing (preferably specific hybridization) the second adaptor with a double-stranded construct in one of a plurality of third samples and ligating the second adaptor to the double-stranded construct to form a triple barcode construct. The triple barcode construct contains a DNA sequence transcribed from an RNA molecule. Hybridization, specific hybridization, and ligation have been discussed above, and any of these methods may be used with the second adaptor. The second adaptor preferably contains an overhang capable of hybridizing (preferably specific hybridization) with an RT primer. The second adaptor preferably contains an overhang capable of hybridizing (preferably specific hybridization) with a sequence at or near the 5' end of the RT primer. These sequences have been discussed in more detail above with reference to the RT primer. The overhang can be of any length. The length of the overhang is preferably at least about 5 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 12 nucleotides, at least about 15 nucleotides, or at least about 20 nucleotides. The length of the overhang is preferably less than about 50 nucleotides, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides. The overhang is preferably the same length as the sequence at or near the 5' end of the RT primer. The second adaptor can use the same overhang. This allows the method to use a standard set of second adaptors, where only the third barcode differs between third samples.

[0161] If the second connector is double-chained, then the distinct third barcode preferably resides in the non-protruding end chain of the second connector connected to the double-chain construct. Figure 3 An example of this situation is shown in the figure. This method produces a construct as a triple barcode.

[0162] Step (f) is preferably performed in the presence of a 5' phosphorylation inhibitor complementary to the protrusion in the second adaptor. Figure 4 An example of this is illustrated. This reduces erroneous re-barcoding or barcode crosstalk between redundant RT primers after cell pooling. These blockers may contain sequences at or near the 5' end of the RT primer. Figure 3 An example of this situation is shown in the figure.

[0163] The second adaptor preferably contains an additional sequence at its 5' end that facilitates the isolation, amplification, and / or characterization of the triple barcode construct. The additional sequence may contain click chemistry. The additional sequence may contain or facilitate the addition of a sequencing adaptor. Such adaptors are discussed in more detail below. The second adaptor preferably contains the same additional sequence. This allows the method to use a standard set of RT primers, where only the second barcode differs between the second samples.

[0164] The second adaptors used in each third sample comprise a different third barcode. The second adaptors used in each third sample comprise a unique third barcode. The double stranded constructs in each third sample are labelled with a different or unique third barcode. The second adaptors used in each third sample comprise a third barcode that is different to the third barcode used in all other third samples. The double stranded constructs in each third sample are labelled with a third barcode that is different to the third barcode used to label the double stranded constructs in all other third samples. The second adaptors used in each third sample comprise a third barcode that is different to any or all of the other barcodes used in the method. The double stranded constructs in each third sample are labelled with a third barcode that is different to any or all of the other barcodes used in the method. Barcodes are discussed in more detail below.

[0165] Exemplary second adaptors are shown in the examples.

[0166] Barcodes

[0167] The method involves generating triple barcoded constructs. Step (b) comprises using a first different barcode. Step (d) involves using a second different barcode. Step (f) comprises using a third different barcode. The meaning of "different" or "unique" barcodes is discussed above. The barcodes are typically polynucleotide barcodes. The barcodes can be any of those discussed above. The first, second and third different polynucleotides typically all comprise the same type of polynucleotide. Polynucleotide barcodes are well known in the art (Kozarewa, I. et al, (2011), Methods Mol. Biol. 733, pp279-298). A barcode is a specific sequence of polynucleotide that can be characterised and identified. The barcode is preferably a specific sequence of polynucleotide that affects the current flow through the pore in a specific and known way.

[0168] The barcodes can comprise one or more different nucleotide species. For example, T k-mers (i.e. k-mers with a central nucleotide based on thymine, such as TTA, GTC, GTG and CTA) generally have the lowest current state. Modified T nucleotides can be introduced into the modified polynucleotide to further reduce the current state, increasing the overall current range seen as the barcode moves through the pore.

[0169] G k-mers (i.e. k-mers with a central nucleotide based on guanine, such as TGA, GGC, TGT and CGA) tend to be strongly influenced by other nucleotides in the k-mer and so modification of G nucleotides in the modified polynucleotide can help them to have more independent current positions.

[0170] Including three copies of the same nucleotide species (rather than three different species) can facilitate characterization because then only, for example, 3-nucleotide k-mers in the modified polynucleotide need to be mapped. However, such modifications do reduce the information provided by the barcodes.

[0171] One or more abasic nucleotides can be included in the barcode. The use of one or more abasic nucleotides results in a characteristic current spike. This can clearly highlight the location of the one or more nucleotide species in the barcode.

[0172] The nucleotide species in the barcode can comprise a chemical atom or group, such as a propynyl, thio, oxo, methyl, hydroxymethyl, formyl, carboxyl, carbonyl, benzyl, propargyl, or propargylamine group. The chemical group or atom can be or can comprise a fluorescent molecule, biotin, digoxigenin, DNP (dinitrophenol), a photo-labile group, an alkyne, DBCO, azide, a free amino group, a redox dye, a mercury atom, or a selenium atom.

[0173] The barcode can comprise a nucleotide species comprising a halogen atom. The halogen atom can be attached to any position on the different nucleotide species, such as the nucleobase and / or sugar. The halogen atom is preferably fluorine (F), chlorine (CI), bromine (Br), or iodine (I). The halogen atom is most preferably F or I.

[0174] The barcode can be any length. The length of the barcode is preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides. The first, second, and third different barcodes can have the same length or different lengths.

[0175] The present invention is based on the use of unique or different barcodes for each sample. As explained in more detail below, this approach generally requires a large number of unique or different barcodes. The longer the barcode, the more different possible barcodes that can be generated and used in the method of the present invention. The skilled person is able to design barcodes and sufficient numbers of barcodes for use in the method of the present invention.

[0176] Barcode numbers

[0177] The barcode used in each sample is preferably different from any other barcode or all other barcodes used in the method. The barcode used in each first sample, each second sample, and each third sample is preferably different from any other barcode or all other barcodes used in the method. This means that each sample is contacted with a unique or different barcode.

[0178] The number of cells in the population is preferably less than the number of first samples multiplied by the number of second samples multiplied by the number of third samples (i.e. first samples x second samples x third samples). If the number of possible barcode combinations used is greater than the number of cells in the population, then it is very likely that the RNA molecules from each cell or constructs produced from each cell are labelled with a unique combination of three barcodes. If the number of possible barcode combinations used is greater than the number of cells in the population, then it is very likely that the RNA molecules from each cell or constructs produced from each cell are labelled with a unique combination of three barcodes. This allows the RNA molecules to be related to a single cell in the population and the transcriptome of each cell or two or more cells to be measured.

[0179] The number of samples is discussed above. The (i) plurality of first samples, (ii) plurality of second samples, and / or (ii) plurality of third samples preferably comprises at least 24 samples. Preferably, (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least about 24 samples. Preferably, (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii), or (i), (ii) and (iii) preferably comprises at least 24 samples, at least 48 samples, at least 96 samples, at least 384 samples, at least 1536 samples, or at least 3456 samples. The minimum number of unique or different barcodes is preferably at least about 13,824, at least about 110,592, at least about 884,736, at least about 56,623,104, at least about 3,623,878,656, or at least about 41,278,242,816 barcodes. The skilled person is able to design a suitable number of barcodes and number of samples based on the number of cells in the starting population.

[0180] Previous steps

[0181] The cell population is preferably fixed and / or permeabilised prior to step (a). Suitable methods for doing so are known in the art (e.g. Chen et al. (2018). PBMC fixation and processing for Chromium single-cell RNA sequencing. Journal of Translational Medicine, 16(1)). Particular methods are also discussed in the examples.

[0182] Additional steps

[0183] The method preferably further comprises repeating steps (e) to (f) at least once, comprising forming a plurality of fourth or more samples, and using a third or more adaptor comprising a different fourth or more barcode for each fourth or more sample to produce a four-plex or more-plex barcode construct. This repetition can be carried out more than once to produce a five-plex, six-plex, seven-plex or eight-plex or more-plex barcode construct. The skilled person is able to devise a method of uniquely labelling RNA molecules or producing unique barcode constructs comprising sequences transcribed from RNA molecules with any number of barcodes. Any of the embodiments discussed above with reference to the first, second and third rounds are equally applicable to any of the repeated steps.

[0184] The method produces constructs with three or more barcodes. The method can produce constructs with any number of three or more barcodes, such as 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 15 or more or 20 or more barcodes. In the following discussion, all of these constructs will be referred to collectively as “uniquely labelled constructs” or “barcode constructs”.

[0185] Characterisation

[0186] The method preferably further comprises isolating the barcode constructs from the cells. Any method of isolation can be used, including any of the methods discussed below. Any of the adaptors or RT primers can comprise sequences and / or molecules that facilitate isolation.

[0187] The method preferably further comprises amplifying the barcode constructs. Any method of amplification can be used, including polymerase chain reaction (PCR). Any of the adaptors or RT primers can comprise primer sites that allow amplification. This facilitates the preparation of the barcode constructs for amplification.

[0188] The method preferably further comprises isolating the barcode constructs from the cells and amplifying the barcode constructs.

[0189] The method preferably further comprises characterising or sequencing the barcode construct. Any method can be used to sequence or characterise the barcode construct, including next generation sequencing. Preferably the barcode construct is characterised or sequenced using a nanopore. This is discussed in more detail below.

[0190] In preferred embodiments, the application provides a method of characterising the transcriptome of a single cell or two or more cells in a cell population, the method comprising: performing the method of the application on the cell population, and characterising the RNA molecules from the single cell or two or more cells using their unique label. The RNA molecules can be characterised using their barcode. The characterisation preferably comprises sequencing the RNA molecules. This can be achieved by sequencing the barcode constructs produced from the RNA molecules. The application also provides a method of characterising the transcriptome of a single cell or two or more cells in a cell population, the method comprising: producing uniquely labelled constructs transcribed from the RNA molecules in the cell population using the method of the application, and characterising the uniquely labelled constructs from the single cell or two or more cells using their unique label. The constructs can be characterised using their barcode. As explained above, the method of the application can be performed such that the RNA molecules from each cell or the barcode constructs produced from each cell are very likely to be labelled with a unique combination of three barcodes. The characterisation is preferably sequencing the RNA molecules or the barcode constructs, preferably using a nanopore. Any of the embodiments discussed above, especially those relating to the first to third rounds, adaptors or RT primers, are equally applicable to these embodiments.

[0191] The methods can involve characterising the transcriptome of any number of two or more cells, such as 3 or more, 4 or more, 5 or more, 10 or more, 20 or more, 50 or more, 100 or more, 500 or more, 1000 or more, 5000 or more or 10,000 or more cells. The number of two or more cells can be any of the numbers discussed above in relation to the cell population used in the methods of the application.

[0192] Sequencing adaptors

[0193] The barcode constructs produced using the methods of the application are preferably characterised or sequenced. The barcode constructs are preferably modified with sequencing adaptors that facilitate characterisation or sequencing, especially using a nanopore. As explained above, the second adaptor can comprise an additional sequence that comprises or facilitates the addition of a sequencing adaptor.

[0194] Sequencing adaptors generally comprise a polynucleotide strand that is capable of attaching to the end of a target polynucleotide. Target polynucleotides are generally used for characterisation according to the methods disclosed herein, and include first adaptors, RT primers, second adaptors and barcode constructs.

[0195] Sequencing adaptors can be added to both ends of a target polynucleotide. Alternatively, different adaptors can be added to both ends of a target polynucleotide. An adaptor can be added to only one end of a target polynucleotide. Methods of adding adaptors to polynucleotides are known in the art. An adaptor can be attached to a polynucleotide, for example, by ligation, by click chemistry, by labelling, by topoisomerisation or by any other suitable method.

[0196] An adaptor can be synthetic or artificial. Typically, an adaptor comprises a polymer as described herein. An adaptor preferably comprises a polynucleotide. An adaptor can comprise a single stranded polynucleotide strand. An adaptor can comprise a double stranded polynucleotide. A sequencing adaptor can comprise any of the polynucleotides discussed above in relation to first adaptors, and includes DNA, RNA, modified DNA (such as alkaline DNA), RNA, PNA, LNA, BNA and / or PEG. Typically, an adaptor comprises single stranded and / or double stranded DNA or RNA.

[0197] A sequencing adaptor can be a Y-adaptor. Y-adaptors are typically double stranded, and comprise (a) at one end, a region where the two strands hybridise together and (b) at the other end, a region where the two strands are not complementary. The non-complementary portion of the strands forms a overhang. The hybridised stem of the adaptor is typically attached to the 5' end of the first strand of the double stranded polynucleotide and the 3' end of the second strand of the double stranded polynucleotide; or to the 3' end of the first strand of the double stranded polynucleotide and the 5' end of the second strand of the double stranded polynucleotide. The presence of the non-complementary region in Y-adaptors gives them their Y-shape, as the two strands do not typically hybridise to each other, unlike the double stranded portion. The hybridised stem end of a Y-adaptor can also comprise a short overhang, which allows them to specifically hybridise and attach to first adaptors, RT primers, second adaptors and barcode constructs.

[0198] Some of the methods of the application use polynucleotide binding proteins to control the movement of barcode constructs relative to a nanopore. A polynucleotide binding protein can bind to an overhang of an adaptor, such as a Y-adaptor. A polynucleotide binding protein can bind to a double stranded region. A polynucleotide binding protein can bind to a single stranded and / or double stranded region of an adaptor. A first polynucleotide binding protein can bind to a single stranded region of such an adaptor, and a second polynucleotide binding protein can bind to a double stranded region of the adaptor.

[0199] The sequencing adaptor preferably comprises a membrane anchor or a pore anchor. The anchor can be attached to a polynucleotide that is complementary to and thus hybridises to the overhang bound by the polynucleotide binding protein.

[0200] One of the strands in the non-complementary strand of the sequencing adaptor, such as a Y adaptor, can comprise a leader sequence that is able to thread into a nanopore when in contact with the nanopore.

[0201] The leader sequence typically comprises a polymer, such as a polynucleotide, for example DNA or RNA, a modified polynucleotide (such as an abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequence preferably comprises a single strand of DNA, such as a polydT segment. The leader sequence can be of any length, but is typically 10 to 150 nucleotides in length, such as 20 to 120, 30 to 100, 40 to 80 or 50 to 70 nucleotides in length.

[0202] The sequencing adaptor can be a hairpin loop adaptor. A hairpin loop adaptor is an adaptor comprising a single polynucleotide strand, wherein the ends of the polynucleotide strand are able to hybridise to or are hybridised to each other, and wherein the middle segment of the polynucleotide forms a loop. Suitable hairpin loop adaptors can be designed using methods known in the art. Typically, the 3' end of the hairpin loop adaptor is attached to the 5' end of the first strand of the double stranded polynucleotide, and the 5' end of the hairpin loop adaptor is attached to the 3' end of the second strand of the double stranded polynucleotide; or the 5' end of the hairpin loop adaptor is attached to the 3' end of the first strand of the double stranded polynucleotide, and the 3' end of the hairpin loop adaptor is attached to the 5' end of the second strand of the double stranded polynucleotide. As explained in more detail below, the sequencing adaptor can be attached to a target polynucleotide to characterise the target polynucleotide.

[0203] Those skilled in the art will also appreciate that when the adaptor comprises a polynucleotide strand, the sequence of the adaptor is generally not determinative and can be controlled or selected depending on the polynucleotide binding protein and other experimental conditions, such as any polynucleotides to be characterized. Exemplary sequences are provided in the examples merely by way of illustration. For example, the adaptor can comprise a sequence such as one or more of SEQ ID NOs: 21 to 26 or 28 to 33 in WO 2021 / 255476 (incorporated herein by reference in its entirety), or a polynucleotide sequence having at least 20%, such as at least 30%, for example at least 40%, such as at least 50%, for example at least 60%, such as at least 70%, for example at least 80%, for example at least 90%, for example at least 95% sequence similarity or identity to one or more of SEQ ID NOs: 21 to 26 or 28 to 33 in WO 2021 / 255476 (incorporated herein by reference in its entirety). The sequence of the adaptor can generally be varied without negatively impacting the efficacy of the methods of the invention.

[0204] The sequencing adaptor can comprise a loading site for loading the polynucleotide binding protein. The loading site can be, for example, a single-stranded region that can be targeted by the polynucleotide binding protein. The loading site can be a region of the sequencing adaptor to which a foreign polynucleotide strand of the polynucleotide binding protein can bind to transfer the polynucleotide binding protein to a polynucleotide to be evaluated in the methods of the invention.

[0205] The polynucleotide binding protein, if present, can be provided on the sequencing adaptor. WO 2015 / 110813 and WO 2020 / 234612 describe loading polynucleotide binding proteins onto target polynucleotides such as adaptors, and are hereby incorporated by reference in their entirety.

[0206] Spacers

[0207] Any of the polynucleotides described herein, including the first adaptor, the RT primer, the second adaptor, and the barcode construct, can comprise one or more spacers, for example, about one to about 10 spacers, for example, about 1 to about 5 spacers, for example, about 1, 2, 3, 4, or 5 spacers. The spacer can comprise any suitable number of spacer units. The spacer generally provides an energetic barrier that impedes movement of the polynucleotide binding protein. For example, the spacer can impede movement of the polynucleotide binding protein by reducing the protein’s drag, for example, using an abasic spacer. The spacer can physically prevent movement of the protein, for example, by introducing bulky chemical groups to physically impede movement of the polynucleotide binding protein.

[0208] One or more spacers are typically included in the polynucleotide or sequencing adapter to provide a unique signal as they thread or cross the nanopore. One or more spacers can be used to define or separate one or more regions of the polynucleotide; for example to separate an adapter from a target polynucleotide.

[0209] A spacer can comprise a linear molecule, such as a polymer, for example a polypeptide or polyethylene glycol (PEG). Typically, such a spacer has a different structure to the target polynucleotide. For example, if the target polynucleotide is DNA, the spacer or each spacer typically does not comprise DNA. In particular, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the spacer or each spacer preferably comprises a peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or synthetic polymer bearing nucleotide side chains. The spacer can comprise one or more nitroindoles, one or more inosines, one or more azidines, one or more 2-amino purines, one or more 2-6-diamino purines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dT), one or more inverted dideoxythymidines (ddT), one or more dideoxycytidines (ddC), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-methyl RNA bases, one or more iso-doxycytidines (Iso-dC), one or more iso-doxynucleotides (Iso-dG), one or more C3 (OC3H6OPO3) groups, one or more photocleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2-CH(CH3)OPO3] groups, one or more hexandiol groups, one or more Sp9 (iSp9) [(OCH2CH2)3OPO3] groups or one or more Sp18 (iSp18) [(OCH2CH2)6OPO3] groups; or one or more thiol linkages. The spacer can comprise any combination of these groups. Many of these groups are commercially available from (Integrated DNA Technologies®). For example, C3, iSp9 and iSp18 spacers are all available from (Integrated DNA Technologies®). The spacer can comprise any number of the above-mentioned groups as spacer units.

[0210] A spacer can comprise one or more chemical groups, e.g., one or more chemical side groups. The one or more chemical groups can be attached to one or more nucleobases in a sequencing adaptor. The one or more chemical groups can be attached to the backbone of a sequencing adaptor. There can be any number of suitable chemical groups, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenol (DNP), digoxin and / or anti-digoxin, and dibenzylcyclooctyne groups.

[0211] A spacer can comprise one or more abasic nucleotides (i.e., nucleotides that lack a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more abasic nucleotides. In an abasic nucleotide, the nucleobase can be replaced with -H (idSp) or -OH. An abasic spacer can be inserted into a target polynucleotide by removing the nucleobase from one or more adjacent nucleotides. For example, a polynucleotide can be modified to include 3-methyladenine, 7-methylguanine, 1,N6-ethenoadenine inosine, or hypoxanthine, and a human alkyladenine DNA glycosylase (hAAG) can be used to remove the nucleobase from these nucleotides. Alternatively, a polynucleotide can be modified to include uracil, and a uracil-DNA glycosylase (UDG) is used to remove the nucleobase. One or more spacers preferably do not comprise any abasic nucleotides.

[0212] A suitable spacer can be designed or selected according to the nature of the polynucleotide or sequencing adaptor, the polynucleotide binding protein, and the conditions under which the method is performed.

[0213] Tag

[0214] Any of the polynucleotides used in the present invention (including the first adaptor, the RT primer, the second adaptor, the barcode construct, and / or the sequencing adaptor) can comprise a tag or tether. For example, a polynucleotide can bind to a tag on a nanopore, e.g., via its adaptor, and be released at some point, e.g., during nanopore characterization of the polynucleotide. Strong non-covalent bonds (e.g., biotin / avidin) are still reversible and can be used in some embodiments of the methods described herein.

[0215] A pore tag and sequencing adaptor pair can be configured such that the binding site on the polynucleotide (e.g., provided by the anchor or the leader sequence of the adaptor or by a capture sequence within the duplex stem of the adaptor) and the tag on the nanopore are of sufficient strength or affinity of binding to maintain the linkage between the nanopore and the polynucleotide until an applied force is placed thereon to release the bound polynucleotide from the nanopore.

[0216] The tag or tether is preferably not charged. This can ensure that the tag or tether is not pulled into the nanopore under the influence of a potential difference.

[0217] One or more molecules that attract or bind to the polynucleotide or adaptor can be attached to the detector (e.g., pore). Any molecule that hybridizes to the adaptor and / or target polynucleotide can be used. The molecules attached to the pore can be selected from the group consisting of PNA tags, PEG linkers, short oligonucleotides, positively charged amino acids, and aptamers. Making such molecules and attaching them to the pores to which they are attached are known in the art. For example, in Howarka et al. (2001) Nature Biotech. 19: 636-639 and WO 2010 / 086620, and pores comprising PEG attached within the cavity of the pore are disclosed in Howarka et al. (2000) J. Am. Chem. Soc. 122(11): 2411-2416.

[0218] Short oligonucleotides attached to the detector (e.g., nanopore) can be used to enhance capture of the target polynucleotide in the methods described herein, the oligonucleotide comprising a sequence complementary to a sequence in the leader sequence or another single-stranded sequence in the adaptor.

[0219] The tag or tether can comprise or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) can be about 10 to 30 nucleotides in length, or about 10 to 20 nucleotides in length. The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) used in the tag or tether can have at least one end (e.g., 3'- or 5'-end) modified to conjugate with other modifications or to conjugate with a solid matrix surface, including, for example, a bead. An end-modifier can add a reactive functional group that can be used for conjugation. Examples of functional groups that can be added include, but are not limited to, amino, carboxyl, thiol, maleimide, aminooxy, and any combination thereof. The functional groups can be combined with spacers of different lengths (e.g., C3, C9, C12, Spacer 9, and 18) to increase the physical distance of the functional group from the end of the oligonucleotide sequence.

[0220] The tag or tether can comprise or be a morpholino oligonucleotide. The morpholino oligonucleotide can be about 10 to 30 nucleotides in length or about 10 to 20 nucleotides in length. The morpholino oligonucleotide can be modified or unmodified. For example, the morpholino oligonucleotide can be modified on the 3' and / or 5' end of the oligonucleotide. Examples of modifications on the 3' and / or 5' end of the morpholino oligonucleotide include, but are not limited to, 3' affinity tags and functional groups for chemical ligation (including, for example, 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl disulfide, and any combination thereof); 5' end modifications (including, for example, 5'-primary amine and / or 5'-dabcyl), click chemistry modifications (including, for example, 3'-azide, 3'-alkyne, 5'-azide, 5'-alkyne), and any combination thereof.

[0221] The tag or tether can further comprise a polymeric linker, for example, to facilitate coupling to a detector (e.g., a nanopore). Exemplary polymeric linkers include, but are not limited to, polyethylene glycol (PEG). The polymeric linker can have a molecular weight of about 500 Da to about 10 kDa (including the end values), or about 1 kDa to about 5 kDa (including the end values). The polymeric linker (e.g., PEG) can be functionalized with different functional groups including, for example, but not limited to, maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combination thereof. The tag or tether can further comprise a 1 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether can further comprise a 2 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether can further comprise a 3 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether can further comprise a 5 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.

[0222] Other examples of tags or tethers include, but are not limited to, His tags, biotin or streptavidin, antibodies that bind to an analyte, aptamers that bind to an analyte, analyte binding domains (such as DNA binding domains) (including, for example, peptide leucine zippers such as leucine zippers, single-stranded DNA binding proteins (SSB)), and any combination thereof.

[0223] The tags or tethers can be attached to the outer surface of the nanopore (e.g., on the cis side of the membrane) using any method known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine bonds), one or more primary amines (such as lysine), one or more unnatural amino acids, one or more histidines (His tags), one or more biotin or streptavidin, one or more antibody-based tags, one or more enzymatic modifications of epitopes (including, for example, acetyltransferases), and any combination thereof. Suitable methods for making such modifications are well known in the art. Suitable unnatural amino acids include, but are not limited to, 4-azido-L-phenylalanine (Faz) and any of the amino acids numbered 1-71 in Liu C.C. and Schultz P.G., Annu. Rev. Biochem., 2010, 79, 413-444. Figure 1 any of the amino acids numbered 1-71 in Liu C.C. and Schultz P.G., Annu. Rev. Biochem., 2010, 79, 413-444.

[0224] In the case where one or more tags or tethers are connected to the nanopore via cysteine linkages, one or more cysteines can be introduced into one or more monomers forming the nanopore via substitution. The nanopore can be chemically modified by attachment of (i) a maleimide, including dibromomaleimide, such as: 4-phenazinylmaleimide, 1. N-(2-hydroxyethyl)maleimide, N-cyclohexylmaleimide, 1.3-maleimidopropionic acid, 1.1-4-aminophenyl-1H-pyrrole,2,5,diketone, 1.1-4-hydroxyphenyl-1H-pyrrole,2,5,diketone, N-ethylmaleimide, N-methoxycarbonylmaleimide, N-tert-butylmaleimide, N-(2-aminoethyl)maleimide, 3-maleimidyl-PROXYL, N-(4-chlorophenyl)maleimide, 1-[4-(dimethylamino)-3,5-dinitrophenyl]-1H-pyrrole-2,5-diketone, N-[4-(2-benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(1-naphthyl)-maleimide, N-(2,4-dimethylphenyl)maleimide, N-(2,4-difluorophenyl)maleimide, N-(3-chloro-p-tolyl)-maleimide, 1-(2-amino-ethyl)-pyrrole-2,5-diketone hydrochloride, 1-cyclopentyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-diketone, 1-(3-aminopropyl)-2,5-dihydro-1H-pyrrole-2,5-diketone hydrochloride, 3-methyl-1-[2-oxo-2-(piperazin-1-yl)ethyl]-2,5-dihydro-1H-pyrrole-2,5-diketone hydrochloride, 1-benzyl-2,5-dihydro-1H-pyrrole-2,5-diketone, 3-methyl-1-(3,3,3-trifluoropropyl)-2,5-dihydro-1H-pyrrole-2,5-diketone, 1-[4-(methylamino)cyclohexyl]-2,5-dihydro-1H-pyrrole-2,5-diketone trifluoroacetate, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, 1-benzyl-3-methyl-2,5-dihydro-1H-pyrrole-2,5-diketone, 1-(2-fluorophenyl)-3-methyl-2,5-dihydro-1H-pyrrole-2,5-diketone, N-(4-phenoxyphenyl)maleimide, N-(4-nitrophenyl)maleimide;(ii) iodoacetamides such as: 3-(2-iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2- iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2- trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4- (aminosulfonyl)phenyl)-2-iodoacetamide, N-(l,3-benzothiazol-2-yl)-2- iodoacetamide, N-(2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4- chlorophenyl)-2-iodoacetamide; (iii) bromoacetamides such as N-(4- (acetylamino)phenyl)-2-bromoacetamide, N-(2-acetylphenyl)-2- bromoacetamide, 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3- (trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide, 2-bromo-N-(4-fluorophenyl)-3-methylbutanamide, N-benzyl-2-bromo-N- phenylpropanamide, N-(2-bromo-butyryl)-4-chloro-benzenesulfonamide, 2-bromo-N-methyl-N-phenylacetamide, 2-bromo-N-phenethyl-acetamide, 2- adamant- 1 -yl-2-bromo-N-cyclohexyl-acetamide, 2-bromo-N-(2- methylphenyl)butanamide, acetotoluidide; (iv) disulfides such as: aldehyde thiol-2, aldehyde thiol-4, isopropyl disulfide, l-(isobutyldisulfanyl)-2- methylpropane, dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2- pyridyldisulfanyl)propanoic acid, 3-(2-pyridyldisulfanyl)propanoic acid hydrazide, 3-(2-pyridyldisulfanyl)propanoic acid N-succinimidyl ester, am6amPDP1-βCD; and (v) thiols such as: 4-phenylthiazole-2-thiol, Purpald, 5,6,7,8-tetrahydro-quinazoline-2-thiol.

[0225] The tag or tether can be connected to the nanopore directly or through one or more linkers. The tag or tether can be connected to the nanopore using a hybrid linker as described in WO 2010 / 086602, which is incorporated herein by reference in its entirety. Alternatively, a peptide linker can be used. A peptide linker is an amino acid sequence. The length, flexibility and hydrophilicity of the peptide linker are generally designed such that it does not interfere with the function of the monomer and the pore. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and / or glycine amino acids. More preferred flexible linkers include (SG)i, (SG)2, (SG)3, (SG)4, (SG)5 and (SG)8, where S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P) 12 where P is proline. where P is proline.

[0226] Suitable hole tags are also described in WO 2018 / 100370, which describes a non- hairpin method for characterizing double-stranded polynucleotides and is incorporated herein by reference in its entirety.

[0227] Anchor

[0228] Any of the polynucleotides used in the present invention (including the first adaptor, the RT primer, the second adaptor, the barcode construct, and / or the sequencing adaptor) can comprise a membrane anchor. The anchor generally facilitates the characterization of the target polynucleotide according to the methods disclosed herein. For example, the membrane anchor can facilitate the positioning of the selected polynucleotide around the nanopore.

[0229] The anchor can be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into a membrane. The hydrophobic anchor is preferably a lipid, a fatty acid, a sterol, a carbon nanotube, a polypeptide, a protein, or an amino acid, such as cholesterol, palmitate, or tocopherol. The anchor can comprise a thiol, a biotin, or a surfactant.

[0230] The anchor can be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or fusion proteins), Ni-NTA (for binding to polyhistidine or polyhistidine tagged proteins), or a peptide (such as an antigen).

[0231] The anchor preferably comprises a linker, or 2, 3, 4, or more linkers. Preferred linkers include, but are not limited to, polymers such as polynucleotides, polyethylene glycol (PEG), polysaccharides, and polypeptides. These linkers can be linear, branched, or cyclic. For example, the linker can be a cyclic polynucleotide. The adaptor can hybridize to a complementary sequence on the cyclic polynucleotide linker. The anchor or anchors or linker or linkers can comprise a component that can be cleaved or broken down, such as a restriction site or a photo-labile group. The linker can be functionalized with a maleimide group to link to a cysteine residue in a protein. Suitable linkers are described in WO 2010 / 086602, which is incorporated herein by reference in its entirety.

[0232] The anchor is preferably a cholesterol or a fatty acyl chain. For example, any fatty acyl chain of 6 to 30 carbon atoms in length can be used, such as hexadecanoic acid.

[0233] Examples of suitable anchors and methods of linking the anchors to adaptors are disclosed in WO 2012 / 164270 and WO 2015 / 150786, which are incorporated herein by reference in their entirety.

[0234] The anchor can consist of or comprise a hydrophobic modification to the polynucleotide or sequencing adaptor. The hydrophobic modification can comprise a modified phosphate group comprised in the polynucleotide or polynucleotide anchor. The hydrophobic modification can for example comprise a phosphorothioate, such as a charge neutralized alkyl phosphorothioate (PPT) described in Jones et al., J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are hereby incorporated by reference. Suitable alkyl groups include for example Ci to C6alkyl groups, for example methyl, ethyl, propyl, butyl, pentyl and hexyl. Incorporation of a charge neutralized alkyl phosphorothioate into a polynucleotide allows the polynucleotide anchor to be anchored to a hydrophobic region, such as a lipid bilayer. 10 alkyl groups, such as C2 to C6 alkyl groups, for example methyl, ethyl, propyl, butyl, pentyl and hexyl. Incorporation of a charge neutralized alkyl phosphorothioate into a polynucleotide allows the polynucleotide anchor to be anchored to a hydrophobic region, such as a lipid bilayer.

[0235] Biotin enrichment

[0236] Any of the polynucleotides (including the first adaptor, RT primer, second adaptor, barcode construct and / or sequencing adaptor) preferably comprise biotin. Biotin can be used to isolate the barcode construct. Suitable methods for biotin-based enrichment are known in the art. For example, a surface (such as a bead) comprising avidin and / or streptavidin can be used to isolate a barcode construct comprising biotin.

[0237] Characterization methods

[0238] The methods of the application preferably further comprise characterizing or sequencing the barcode construct. This allows the RNA molecule to be characterized or sequenced. As explained in more detail above, the methods of the application preferably further comprise characterizing or sequencing the barcode construct using the sequencing adaptor. The application also provides a method of characterizing the transcriptome of a single cell or two or more cells in a cell population, comprising: characterizing an RNA molecule from a single cell or characterizing a uniquely labeled construct from a single cell or two or more cells using its unique label. The RNA molecule or construct can be characterized using its barcode code. These methods preferably comprise using a sequencing adaptor.

[0239] Any characterization method can be used. The method preferably uses next generation sequencing (NGS).

[0240] The barcode construct is preferably moved relative to a detector, such as a nanopore. The detector can be selected from the group consisting of (i) a zero mode waveguide, (ii) a field effect transistor, optionally a nanowire field effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube; and (v) a nanopore. Preferably, the detector is a nanopore.

[0241] The barcode construct can be characterised in any suitable way in the methods of the application. The barcode construct is preferably characterised by detecting an ionic current or optical signal as it moves relative to a nanopore. This is described in more detail herein. This approach is applicable to these and other methods of characterising polynucleotides.

[0242] In another non-limiting example, the barcode construct is characterised by detecting a by-product of a polynucleotide processing reaction, such as a sequencing-by-synthesis reaction. The method can thus involve detecting a product of the sequential addition of a nucleotide(s) to the barcode construct by an enzyme, such as a polymerase. The product can be a change in one or more properties of the enzyme, such as a change in the conformation of the enzyme. Such a method can thus comprise subjecting an enzyme, such as a polymerase or reverse transcriptase, to a barcode construct as a template under conditions such that the template-dependent incorporation of a nucleotide base into a growing oligonucleotide chain in response to a template nucleic acid base encountered in sequence and / or a natural or analogue base specified by the template for incorporation (i.e. an incorporation event) causes a change in the conformation of the enzyme, detecting the change in the conformation of the enzyme in response to such an incorporation event, and thereby detecting the sequence of the template. In such a method, the barcode construct can be moved in accordance with the methods of the application. Such a method can involve detecting and / or measuring the incorporation event using methods known to those skilled in the art, such as the methods described in US 2017 / 0044605.

[0243] In another embodiment, the by-product can be labelled such that a phosphate label species is released upon addition of a nucleotide to a synthetic nucleic acid strand complementary to the template barcode construct, and detected, for example, using a detector as described herein. A barcode construct characterised in this way can be moved in accordance with the methods herein. A suitable label can be an optical label detected using a nanopore or zero-mode waveguide or by Raman spectroscopy or other detector. A suitable label can be a non-optical label detected using a nanopore or other detector.

[0244] In another method, the nucleoside phosphate (nucleotide) is not labelled and a natural by-product species is detected upon addition of a nucleotide to a synthetic nucleic acid strand complementary to the barcode construct. A suitable detector can be an ion-sensitive field effect transistor or other detector.

[0245] These and other detection methods are applicable to the methods described herein. Any suitable measurement can be made using the detector as the barcode construct moves relative to the detector.

[0246] Nanopore characterisation

[0247] The barcode construct is preferably characterised using a nanopore.

[0248] The method preferably comprises: (i) contacting the barcode construct with a nanopore such that the barcode construct moves relative to the nanopore; and (ii) taking one or more measurements as the barcode construct moves relative to the nanopore, wherein the measurement(s) is indicative of one or more characteristics of the barcode construct, and therefrom characterising the barcode construct. The one or more characteristics is preferably selected from the group consisting of (i) the length of the barcode construct, (ii) the identity of the barcode construct, (iii) the sequence of the barcode construct, (iv) the secondary structure of the barcode construct, and (v) whether the barcode construct is modified. The barcode construct can be modified by methylation, oxidation, damage, with one or more proteins, or with one or more labels, tags or spacers. The one or more characteristics of the barcode construct is preferably measured by electrical measurements and / or optical measurements. The electrical measurements are preferably current measurements, impedance measurements, tunnelling measurements or field effect transistor (FET) measurements.

[0249] The method more preferably comprises: (i) contacting the barcode construct with a nanopore such that the barcode construct moves through the nanopore; and (ii) measuring the current flowing through the nanopore as the barcode construct moves through the nanopore, wherein the current is indicative of one or more characteristics of the barcode construct, and therefrom characterising the barcode construct. The one or more characteristics can be any of those described above.

[0250] The movement of the barcode construct relative to the nanopore or through the nanopore is preferably controlled using a polynucleotide binding protein. The use of such proteins in nanopore sequencing is known. Examples of suitable proteins are discussed in more detail below.

[0251] Any suitable nanopore can be used. The nanopore is preferably a transmembrane pore. A transmembrane pore is a structure that is through a membrane to some extent. The transmembrane pore allows the flow of hydrated ions across or within the membrane driven by an applied potential. The transmembrane pore is typically through the whole of the membrane such that hydrated ions can flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to be through the membrane. The transmembrane pore can be closed at one end. For example, the pore can be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions can flow.

[0252] The nanopore typically has a first opening and a second opening. The first opening is typically the cis opening and the second opening is typically the trans opening. However, the first opening can be the trans opening and the second opening can be the cis opening. Any polynucleotide binding protein used in the methods of the application is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore towards the first opening of the nanopore.

[0253] Any transmembrane pore can be used in the methods of the application. The pore can be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores, and solid state pores. The pore can be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.

[0254] The nanopore is preferably a transmembrane protein pore. A transmembrane protein pore is a polypeptide or collection of polypeptides that allows the flow of hydrated ions, such as polynucleotides, from one side of a membrane to the other. In the methods of the application, the transmembrane protein pore is capable of forming a pore that allows the flow of hydrated ions from one side of a membrane to the other driven by an applied electric potential. The transmembrane protein pore preferably allows the flow of polynucleotides from one side of a membrane, such as a triblock copolymer membrane, to the other. The transmembrane protein pore allows the movement of polynucleotides through the pore.

[0255] The nanopore can be a transmembrane protein pore that is a monomer or an oligomer. The pore is preferably composed of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pore is preferably a hexameric, heptameric, octameric, or nonameric pore. The pore can be a homooligomer or a heterooligomer.

[0256] The transmembrane protein pore can comprise a barrel or channel through which ions can flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane beta-barrel or channel or a transmembrane alpha-helix bundle or channel.

[0257] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near the constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine, or histidine, or an aromatic amino acid, such as tyrosine or tryptophan. These amino acids typically facilitate interaction between the pore and a nucleotide, polynucleotide, or nucleic acid.

[0258] The nanopore can be a transmembrane protein pore derived from a beta-barrel pore or an alpha-helix bundle pore. Beta-barrel pores comprise a barrel or channel formed by beta-strands. Suitable beta-barrel pores include, but are not limited to, beta-toxins such as alpha-hemolysin, anthrax toxin and leukocidin, as well as outer membrane proteins / pore proteins of bacteria such as Mycobacterium smegmatis pore proteins (Msp) (e.g. MspA, MspB, MspC or MspD), CsgG, outer membrane pore protein F (OmpF), outer membrane pore protein G (OmpG), outer membrane phospholipase A and Neisseria autotransporter lipoprotein (NalP) and other pores such as cytolysin. Alpha-helix bundle pores comprise a barrel or channel formed by alpha-helices. Suitable alpha-helix bundle pores include, but are not limited to, inner membrane proteins and alpha outer membrane proteins such as WZA and ClyA toxin.

[0259] The nanopore can be a transmembrane pore derived from or based on Msp, alpha-hemolysin (a-HL), cytolysin, CsgG, ClyA, Sp1 or hemolysin fragaceatoxin C (FraC).

[0260] The nanopore can be a transmembrane protein pore derived from, for example, CsgG from Escherichia coli Str. K-12 Subsp. MC4100. Such pores are oligomeric and typically comprise 7, 8, 9 or 10 monomers derived from CsgG. The pore can be a homooligomeric pore derived from CsgG comprising the same monomer. Alternatively, the pore can be a heterooligomeric pore derived from CsgG comprising at least one monomer that is different from the other monomers. Examples of suitable pores derived from CsgG are disclosed in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318 and WO 2019 / 002893 (all documents are incorporated by reference in their entirety).

[0261] The nanopore can be a transmembrane pore derived from cytolysin. Examples of suitable pores derived from cytolysin are disclosed in WO 2013 / 153359, which is incorporated by reference in its entirety.

[0262] The nanopore can be a transmembrane pore derived from or based on alpha-hemolysin (a-HL). The wild-type alpha-hemolysin pore is formed from 7 identical monomers or subunits (i.e. it is heptameric). The alpha-hemolysin pore can be alpha-hemolysin-NN or a variant thereof. The variant preferably comprises N residues at positions E111 and K147.

[0263] The nanopore can be a transmembrane protein pore derived from Msp (e.g. derived from MspA). Examples of suitable pores derived from MspA are disclosed in WO 2012 / 107778, which is incorporated herein in its entirety by reference.

[0264] The nanopore can be a transmembrane pore derived from or based on ClyA.

[0265] membrane

[0266] The detector or nanopore is typically present in a membrane. Any suitable membrane can be used.

[0267] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules such as phospholipids which have both hydrophilic and lipophilic properties. The amphiphilic molecules can be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form monolayers are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomeric subunits polymerise together to produce a single polymer chain. Block copolymers generally have properties contributed by each monomeric subunit. However, block copolymers can have unique properties not possessed by the polymers formed from the individual subunits. Block copolymers can be designed such that one of the monomeric subunits is hydrophobic (i.e. lipophilic) while the other subunits are hydrophilic when in an aqueous medium. In this case, the block copolymers can possess amphiphilic properties and can form structures which mimic biological membranes. Block copolymers can be diblock (consisting of two monomeric subunits) but can also be constructed from more than two monomeric subunits to form more complex arrangements which behave as amphiphiles. The copolymers can be triblock, tetrablock or pentablock copolymers. The membrane can be a triblock copolymer membrane.

[0268] Archaeal bipolar tetraether lipids are naturally occurring lipids which are constructed such that the lipids form monolayer membranes. These lipids are generally found in extremophiles, thermophiles, halophiles and acidophiles which survive in harsh biological environments. Their stability is thought to be derived from the fusogenic nature of the final bilayer. It is simple to construct a block copolymer material which mimics these biological entities by producing a triblock polymer with the general motif hydrophilic-hydrophobic-hydrophilic. This material can form a monolayer which behaves similar to a lipid bilayer and encompasses a range of phases of behaviour from vesicles to lamellar membranes. The membranes formed from these triblock copolymers retain several advantages over biological lipid membranes. Because the triblock copolymers are synthetic, the exact construction can be carefully controlled to provide the correct chain length and properties required to form the membrane and to interact with pores and other proteins.

[0269] The block copolymer can also be composed of subunits that are not lipid submaterials; for example, the hydrophobic polymer can be made of siloxane or other non-carbon-hydrocarbon based monomers. The hydrophilic subsegment of the block copolymer can also possess low protein binding properties, which allows for the creation of a membrane that is highly resistant when exposed to raw biological samples. The head group unit can also be derived from non-classical lipid head groups.

[0270] The triblock copolymer membrane also has increased mechanical and environmental stability compared to biological lipid membranes, such as much higher operating temperature or pH ranges. The synthetic nature of the block copolymer provides a platform for customizing polymer-based membranes for a wide range of applications.

[0271] The membrane can be one of the membranes disclosed in International Application No. WO 2014 / 064443 or WO 2014 / 064444 (both of which are incorporated by reference herein in their entireties).

[0272] The amphiphilic molecule can be chemically modified or functionalized to facilitate coupling of the polynucleotide. The amphiphilic layer can be a single layer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer can be curved. The amphiphilic layer can be supported.

[0273] The amphiphilic membrane is typically naturally mobile, essentially acting as a two- dimensional fluid with a lipid diffusion rate of about 10 -8 cm s -1 This means that pores and coupled polynucleotides can generally move within the amphiphilic membrane.

[0274] The membrane can be a lipid bilayer. Lipid bilayers are models of cell membranes and are used as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro studies of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer can be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, planar lipid bilayers, supported bilayers, or liposomes. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, WO 2009 / 077734, and WO 2006 / 100484 (which are incorporated by reference herein in their entireties).

[0275] Methods for forming lipid bilayers are known in the art. Lipid bilayers are typically formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).

[0276] A lipid bilayer can be formed as described in WO 2009 / 077734, which is incorporated herein in its entirety by reference. In this method, the lipid bilayer is formed from dry lipids. The lipid bilayer can be formed across an opening, as described in WO 2009 / 077734.

[0277] The membrane can comprise a solid state layer. The solid state layer can be formed from organic and inorganic materials, including but not limited to microelectronic materials, insulating materials such as Si3N4, A12O3 and SiO, organic and inorganic polymers such as polyamides, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber and glass. The solid state layer can be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647, which is incorporated herein in its entirety by reference. If the membrane comprises a solid state layer, then pores are typically present in an amphiphilic membrane or layer, which is contained within the solid state layer, for example within a hole, pore, gap, channel, trench or interstice within the solid state layer. The skilled person can prepare suitable solid state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857, which are incorporated herein in their entirety by reference. Any of the amphiphilic membranes or layers discussed above can be used.

[0278] The methods disclosed herein are typically performed using (i) an artificial amphiphilic layer comprising pores, (ii) a separate naturally occurring lipid bilayer comprising pores, or (iii) a cell into which pores have been inserted. The methods are typically performed using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer can comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pores. Suitable apparatus and conditions are discussed below. The methods of the application are typically performed in vitro.

[0279] Polynucleotide binding protein

[0280] As will be appreciated by the skilled person, any suitable polynucleotide binding protein can be used in the methods and products of the application. The polynucleotide binding protein can be any protein that is capable of binding to a polynucleotide and controlling its movement relative to a detector (e.g. a nanopore).

[0281] In more detail, the polynucleotide binding protein (such as a helicase) can typically control the movement of DNA in at least two active modes of operation (when provided with all the necessary components for facilitating movement, e.g. ATP and Mg 2+ and one non-active mode of operation (when not provided with the necessary components for facilitating movement; or when the polynucleotide binding protein is modified to prevent the active mode).

[0282] Polynucleotide binding proteins can move along a polynucleotide, such as DNA, in either the 5'-3' direction or the 3'-5' direction when provided with all the necessary components for facilitating movement. Many polynucleotide binding proteins process polynucleotides, such as DNA, in the 5'-3' direction. Polynucleotide binding proteins that control the movement of polynucleotides in this way are generally suitable for use in the methods of the application.

[0283] However, when a polynucleotide binding protein is not provided with the necessary components for facilitating movement, or is modified to prevent it from actively controlling the movement of a polynucleotide relative to a nanopore, it can still passively control the movement of a polynucleotide relative to a nanopore. For example, a polynucleotide binding protein can bind to a polynucleotide and act as a brake to slow the movement of the polynucleotide as it is pulled into a pore by an applied field (e.g. by the first force in the methods of the application). In "non-active" mode, it is generally irrelevant whether the DNA is captured 3' or 5' down (i.e. moves through the nanopore in the 5'-3' direction or in the 3'-5' direction) as the applied force provides the motive force for the movement of the polynucleotide through the nanopore. However, in such embodiments, the polynucleotide binding protein can still control the movement of the polynucleotide relative to the nanopore, for example by acting as a brake. When in non-active mode, the control of the movement of the polynucleotide by the polynucleotide binding protein can be described in a number of ways, including ratcheting, slipping and braking. Generally, the methods of the application do not comprise the use of a polynucleotide binding protein that operates in passive mode. However, when a polynucleotide binding protein is used, it can be a polynucleotide binding protein that operates in passive mode.

[0284] Some methods of the application can comprise the use of a polynucleotide binding protein as a pause moiety to impede the movement of a polynucleotide strand through a nanopore. The polynucleotide binding protein can be a protein that binds to a polynucleotide but does not have polynucleotide processing ability, i.e. it is not a polynucleotide processing enzyme.

[0285] A polynucleotide processing enzyme is a polypeptide that is capable of interacting with a polynucleotide. The enzyme can modify the polynucleotide by cleaving the polynucleotide to form individual nucleotides or shorter chains of nucleotides, such as di- or tri-nucleotides. The enzyme can modify the polynucleotide by directing or moving it to a particular location. A polynucleotide binding protein as used herein can be or can be derived from a polynucleotide processing enzyme. A polynucleotide binding protein can be or can be derived from a polynucleotide processing enzyme.

[0286] The polynucleotide binding protein can be derived from a member of any of the groups in the Enzyme Classification (EC) groups: 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30, and 3.1.31.

[0287] Typically, the polynucleotide binding protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.

[0288] The polynucleotide binding protein can be modified to prevent the polynucleotide binding protein from dissociating from the polynucleotide. Thus, the target polynucleotide is preferably not dissociated from the polynucleotide binding protein.

[0289] As used herein, the term "dissociation" refers to the unbinding of the polynucleotide binding protein from the target polynucleotide. Thus, the polynucleotide binding protein can be modified to prevent it from unbinding from the target polynucleotide, e.g., unbinding into the reaction medium. It is important to distinguish between potential "dissociation" of the polynucleotide binding protein from "unbinding" of the polynucleotide binding protein from the target polynucleotide. As used herein, "unbinding" refers to the transient release of the target polynucleotide to the active site of the polynucleotide binding protein (described in more detail herein), but does not imply dissociation. Thus, for example, the polynucleotide binding protein can be modified to prevent the polynucleotide binding protein from dissociating from the polynucleotide, but not from unbinding from the polynucleotide. When unbound, the polynucleotide binding protein is still conjugated to the target polynucleotide. For example, the polynucleotide binding protein can still be conjugated to the target polynucleotide (i.e., it can be prevented from dissociating from the target polynucleotide) because it is topologically closed around the target polynucleotide. The polynucleotide binding site can still be free to bind or unbind the target polynucleotide, such that the polynucleotide binding protein can bind or unbind the target polynucleotide while the polynucleotide binding protein is still conjugated to the target polynucleotide. When the polynucleotide binding protein is unbound from the target polynucleotide, it can be able to move over (e.g., along) the target polynucleotide under an applied force and can be able to rebind to the target polynucleotide. When conjugated on the target polynucleotide but unbound from the target polynucleotide, the polynucleotide binding protein cannot dissociate from the target polynucleotide.

[0290] A polynucleotide binding protein can be modified in any suitable way to prevent dissociation. For example, a polynucleotide binding protein can be loaded onto a polynucleotide and then modified to prevent it from dissociating from the polynucleotide. Alternatively, a polynucleotide binding protein can be modified to prevent it from dissociating from a polynucleotide before being loaded onto the polynucleotide. Modification of a polynucleotide binding protein and / or a polynucleotide binding protein to prevent it from dissociating from a polynucleotide can be achieved using methods known in the art, such as the methods discussed in WO 2014 / 013260 (hereby incorporated by reference in its entirety) and with particular reference to the description of modifying a polynucleotide binding protein such as a helicase to prevent it from dissociating from a polynucleotide strand. For example, a polynucleotide binding protein can be modified by treatment with tetramethylazodicarbonamide (TMAD). Various other closure moieties are described in WO 2021 / 255476 (which is incorporated herein by reference in its entirety).

[0291] For example, a polynucleotide binding protein and / or a polynucleotide binding protein can have a polynucleotide unbinding opening, e.g., a cavity, crevice, or gap through which a polynucleotide strand can pass when the polynucleotide binding protein dissociates from the strand. A polynucleotide unbinding opening can be an opening through which a polynucleotide can pass when the polynucleotide binding protein dissociates from the polynucleotide. A polynucleotide unbinding opening of a given polynucleotide binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. X-ray crystal structures can be obtained in the presence and / or absence of a polynucleotide substrate. The position of a polynucleotide unbinding opening in a given polynucleotide binding protein can be inferred or confirmed by molecular modeling using standard packages known in the art. A polynucleotide unbinding opening can be transiently created by movement of one or more portions, e.g., one or more domains, of a polynucleotide binding protein.

[0292] A polynucleotide binding protein can be modified by closing a polynucleotide unbinding opening. A polynucleotide unbinding opening can be closed with a closure moiety. Thus, closing a polynucleotide unbinding opening can prevent a polynucleotide binding protein from dissociating from a polynucleotide. For example, a polynucleotide binding protein can be modified by covalently closing a polynucleotide unbinding opening. However, as explained above, closing a polynucleotide unbinding opening does not necessarily prevent a target polynucleotide from unbinding from a polynucleotide binding site of a polynucleotide binding protein. A preferred protein for addressing in this way is a helicase.

[0293] The polynucleotide binding protein can be modified with a closure moiety for (i) topologically closing the polynucleotide binding site of the polynucleotide binding protein around the target polynucleotide and (ii) facilitating unbinding of the target polynucleotide from the polynucleotide binding site of the polynucleotide binding protein and / or delaying rebinding of the target polynucleotide to the polynucleotide binding site of the polynucleotide binding protein. The polynucleotide binding protein can be modified in any suitable way to facilitate ligation of such a closure moiety.

[0294] The closure moiety can comprise a bifunctional cross-linking moiety. The closure moiety can comprise a bifunctional cross-linking agent. The bifunctional cross-linking agent can ligate at two points on the polynucleotide binding protein and close the polynucleotide unbinding opening of the polynucleotide binding protein, thereby preventing the polynucleotide from dissociating from the polynucleotide binding protein while allowing unbinding of the polynucleotide from the polynucleotide binding site of the polynucleotide binding protein.

[0295] The closure moiety can be ligated at any suitable position on the polynucleotide binding protein. For example, the closure moiety can cross-link two amino acid residues of the polynucleotide binding protein. Typically, at least one of the amino acids cross-linked by the closure moiety is a cysteine or a non-natural amino acid. The cysteine or non-natural amino acid can be introduced into the polynucleotide binding protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide binding protein. Methods for introducing non-natural amino acids are well known in the art and include, for example, native chemical ligation with synthetic polypeptide chains comprising such non-natural amino acids. Methods for introducing cysteines into polynucleotide binding proteins are likewise within the ability of one skilled in the art, for example using techniques disclosed in the references such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 4th Ed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0296] The closure moiety can have a length of about 1 A to about 100 A. The length of the closure moiety can be calculated according to static bond lengths or more preferably using molecular dynamics simulations. The length may, for example, be about 2 A to about 80 A, such as about 5 A to about 50 A, for example about 8 A to about 30 A, such as about 10 A to about 25 A or about 20 A, for example about 1 A, 2 A, 3 A, 4 A, 5 A, 6 A, 7 A, 8 A, 9 A, 10 A, 11 A, 12 A, 13 A, 14 A, 15 A, 16 A, 17 A, 18 A, or 19 A.

[0297] Polynucleotide binding proteins suitable for closure using a closure moiety as described above are discussed in more detail herein. The polynucleotide binding protein is preferably a helicase, for example a Dda helicase as described herein.

[0298] The polynucleotide binding protein can be or can be derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III from E. coli, RecJ from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.

[0299] The polynucleotide binding protein can be a polymerase. The polymerase can be PyroPhage® 3173 DNA polymerase (which is commercially available from Lucigen®), SD polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or variants thereof. Modified forms of Phi29 polymerase that can be used in the present application are disclosed in US Patent No. 5,576,204.

[0300] The polynucleotide binding protein can be a topoisomerase. In one embodiment, the topoisomerase is a member of any of the partial classification (EC) groups: 5.99.1.2 and 5.99.1.3. The topoisomerase can be a reverse transcriptase, which is an enzyme capable of catalyzing the formation of cDNA from an RNA template. Such topoisomerases are commercially available from, for example, New England Biolabs® and Invitrogen®.

[0301] The polynucleotide binding protein is preferably a helicase. Any suitable helicase can be used in accordance with the methods of the present application. For example, the enzyme or each enzyme used in accordance with the present disclosure can be independently selected from a Hel308 helicase, a RecD helicase, a TraI helicase, a TrwC helicase, an XPD helicase, and a Dda helicase or a variant thereof. Monomeric helicases can comprise several domains linked together. For example, TraI helicases and the TraI subgroup of helicases can contain two RecD helicase domains, a relaxase domain, and a C-terminal domain. These domains generally form a monomeric helicase that is able to function without forming oligomers. Specific examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pif1, and TraI. These helicases generally act on single stranded DNA. Examples of helicases that can move along both strands of double stranded DNA include FtfK and the hexameric helicase complex, or multi-subunit complexes such as RecBCD. The polynucleotide binding protein is preferably a Dda (DNA-dependent ATPase) helicase.

[0302] Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.

[0303] The helicase can be a Trwc Cba or a variant thereof, a Hel308 Mbu or a variant thereof, or a Dda or a variant thereof. The variant can differ from the native sequence in any of the ways discussed herein. An exemplary variant of Dda comprises E94C / A360C. A further exemplary variant of Dda comprises E94C / A360C, and then is (AM1)G1G2 (i.e., a deletion of M1, and then an addition of G1 and G2).

[0304] General methods

[0305] As noted above, the methods of the present application can be operated using any suitable detector, and thus can use any suitable apparatus for detecting polynucleotides.

[0306] The methods of the application can be performed using any device suitable for nanopore sensing. For example, the device can comprise a chamber comprising an aqueous solution and a barrier dividing the chamber into two sections. The barrier typically has an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.

[0307] The methods can be performed using a device described in WO 2008 / 102120, WO 2010 / 122293 or WO 00 / 28312, which are incorporated herein by reference in their entirety. Briefly, binding of a molecule (e.g., a target polynucleotide) in the channel of a pore will have an effect on the open channel ion flow through the pore, which is the essence of “molecular sensing” of the pore channel. Changes in open channel ion flow can be measured by changes in current using suitable measurement techniques. The extent of the reduction in ion flow measured by a reduction in current is related to the size of the obstacle within or near the pore. Thus, binding of a molecule of interest (e.g., a target polynucleotide) in or near the pore provides a detectable and measurable event, forming the basis of a “biosensor”. Detecting the presence of a biological molecule can be applied to personalized medicine development, medicine, diagnostics, life science research, environmental monitoring, and the security and / or defense industries.

[0308] When used to characterize a polynucleotide, the presence, absence, or one or more characteristics of a target polynucleotide are determined. The method can be used to determine the presence, absence, or one or more characteristics of at least one target polynucleotide. The method can involve determining the presence, absence, or one or more characteristics of two or more target polynucleotides. The method can comprise determining the presence, absence, or one or more characteristics of any number of target polynucleotides, such as 2, 5, 10, 15, 20, 30, 40, 50, 100, or more target polynucleotides. Any number of characteristics of one or more target polynucleotides can be determined, such as 1, 2, 3, 4, 5, 10, or more characteristics. Characteristics suitable for detection in the methods provided herein include the identity or sequence of a polynucleotide, the length of a polynucleotide, whether a polynucleotide is modified, etc. In some embodiments, the methods of the application are methods of sequencing a barcode construct. In some embodiments, the sequence of a barcode construct can be determined in real time by aligning real-time signals or base calls to a known reference. Exemplary methods of determining polynucleotide sequences are described in WO 2016 / 059427, which is incorporated herein by reference in its entirety.

[0309] When used to characterise a polynucleotide, the method can involve measuring the ionic current flowing through the pore, typically by measuring the current. Alternatively, the flow of ions through the pore can be measured optically, such as disclosed in Heron et al: J. Am. Chem. Soc. Vol. 131, No. 5, 2009. Thus, the apparatus can also comprise circuitry capable of applying an electrical potential and measuring the electrical signal across the membrane and pore. The characterisation method can be performed using a patch clamp or voltage clamp. The characterisation method preferably involves the use of a voltage clamp.

[0310] The method can involve measuring an optical signal, as described in Chen et al, Nature Communications (2018) 9: 1733, the entire contents of which are hereby incorporated by reference. For example, single molecule surface enhanced Raman spectroscopy (SERS) can be locally enabled using a nanopore such as an optically engineered nanopore structure (e.g. plasmonic nanoslit) to allow characterisation of a polynucleotide by direct Raman spectroscopic detection.

[0311] The method can be performed on a silicon-based pore array, wherein each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more pores.

[0312] The method can involve measuring the current flowing through the pore. The method is typically performed with an applied voltage across the membrane and pore. The voltage used is typically +2 V to -2 V, typically -400 mV to +400 mV. The voltage used is preferably within a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20 mV and 0 mV, and an upper limit independently selected from +10 mV, +20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more preferably within the range 100 mV to 240 mV, and most preferably within the range 120 mV to 220 mV. By using an increased applied potential, the discrimination between different nucleotides can be increased through the pore.

[0313] The method is typically performed in the presence of any charge carrier, such as a metal salt, e.g., an alkali metal salt; a halide salt, e.g., a chloride salt such as an alkali metal chloride salt. The charge carrier can include an ionic liquid or an organic salt, such as tetramethylammonium chloride, trimethylphenylammonium chloride, phenyltrimethylammonium chloride, or 1-ethyl-3-methylimidazolium chloride. In the exemplary apparatus discussed above, the salt is present in an aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI), or cesium chloride (CsCI) are typically used. KCI is preferred. The salt can be an alkaline earth metal salt, such as calcium chloride (CaCI2). The salt concentration can be at saturation. The salt concentration can be 3 M or less, and is typically 0.1 M to 2.5 M, 0.3 M to 1.9 M, 0.5 M to 1.8 M, 0.7 M to 1.7 M, 0.9 M to 1.6 M, or 1 M to 1.4 M. The salt concentration is preferably 150 mM to 1 M. The method is preferably performed using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M, or at least 3.0 M. High salt concentrations provide a high signal-to-noise ratio and allow for the identification of binding / no binding of current against the background of normal current fluctuations.

[0314] The method is typically performed in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in an aqueous solution in the chamber. Any suitable buffer can be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCl buffer. The method is typically performed at a pH of 4.0 to 12.0, 4.5 to 10.0, 5.0 to 9.0, 5.5 to 8.8, 6.0 to 8.7, or 7.0 to 8.8, or 7.5 to 8.5. The pH used is preferably about 7.5.

[0315] The method can be performed at 0 °C to 100 °C, 15 °C to 95 °C, 16 °C to 90 °C, 17 °C to 85 °C, 18 °C to 80 °C, 19 °C to 70 °C, or 20 °C to 60 °C. The method is typically performed at room temperature. The method is optionally performed at a temperature that supports the function of the enzyme, such as about 37 °C.

[0316] Any of the proteins described herein, such as the protein pore, can be prepared synthetically or by recombinant means. For example, the pore can be synthesized by in vitro translation and transcription (IVTT). The amino acid sequence of the pore can be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When the protein is produced by synthetic means, such amino acids can be introduced during production. The pore can also be altered after synthetic or recombinant production.

[0317] Any of the proteins described herein, such as the protein pores, can be produced using standard methods known in the art. Polynucleotide sequences encoding the pores or constructs can be derived and replicated using standard methods in the art. Polynucleotide sequences encoding the pores or constructs can be expressed in bacterial host cells using standard techniques in the art. Pores can be produced in cells by in situ expression of the polypeptides from recombinant expression vectors. The expression vectors optionally carry an inducible promoter to control expression of the polypeptides. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rdedition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0318] Pores can be produced in large scale from organisms producing the proteins after purification by any protein liquid chromatography system or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA system, Bio-Cad system, Bio-Rad BioLogic system, and Gilson HPLC system.

[0319] Kits and systems

[0320] The present application also provides a kit for uniquely labeling RNA molecules in a population of cells. The kit is preferably used to produce uniquely labeled constructs transcribed from the RNA molecules in the population of cells. The uniquely labeled constructs transcribed from the RNA molecules are preferably uniquely labeled constructs comprising DNA sequences transcribed from the RNA molecules.

[0321] The kit comprises two or more first adapters, wherein the two or more first adapters comprise a primer site for RT, and wherein each of the two or more first adapters comprises the reverse complement of a different first barcode. The kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000, or at least about 5000 first adapters. The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536, or at least about 3456 first adapters.

[0322] Any of the embodiments discussed above with reference to the first adapters used in the methods of the application are equally applicable to the kit. The kit can comprise any number of first adapters. The two or more first adapters are preferably double-stranded and comprise an overhang capable of hybridizing to the 3' end, poly(A) tail, or polyadenylated 3' end of an RNA molecule. The nucleotide at the 3' end of one of the two strands of the double-stranded first adapter is preferably ddC or inverted dT.

[0323] The kit preferably further comprises: (a) two or more RT primers capable of hybridizing to two or more first adaptors, and each primer comprises a different second barcode; and / or (b) two or more second adaptors, each primer comprising a different third barcode. The kit preferably comprises at least about 20, at least about 24, at least about 25, at least about 30, at least about 40, at least about 48, at least about 50, at least about 60, at least about 70, at least about 80, at least about 90, at least about 96, at least about 100, at least about 200, at least about 300, at least about 350, at least about 380, at least about 384, at least about 400, at least about 500, at least about 1000, at least about 1500, at least about 1536, at least about 2000, at least about 3000, at least about 3400, at least about 3456, at least about 4000, or at least about 5000 first adaptors. The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536, or at least about 3456 second adaptors and / or RT primers. The kit preferably comprises the same number of first adaptors, RT primers, and second adaptors. Any of the embodiments discussed above with reference to the RT primers and second adaptors used in the methods of the application are equally applicable to the kit.

[0324] Any of the embodiments discussed above with reference to the barcodes used in the methods of the application are equally applicable to the kit. Each barcode in the kit is preferably different from any other barcode or all other barcodes in the kit.

[0325] The kit preferably comprises at least about 24, at least about 48, at least about 96, at least about 384, at least about 1536, or at least about 3456 first adaptors, second adaptors, and RT primers.

[0326] The application also provides a system for performing the methods of the application. The system is for uniquely labeling RNA molecules in a population of cells. The system is preferably for producing uniquely labeled constructs transcribed from RNA molecules in a population of cells. The uniquely labeled constructs transcribed from RNA molecules are preferably uniquely labeled constructs comprising DNA sequences transcribed from RNA molecules.

[0327] The system comprises: (a) two or more first adapters, wherein the two or more first adapters comprise a primer site for RT, and wherein each of the two or more first adapters comprises the reverse complement of a different first barcode; (b) two or more RT primers capable of hybridizing to the two or more first adapters, and each primer comprises a different second barcode; (c) two or more second adapters, each primer comprises a different third barcode; and (d) a nanopore. Any of the embodiments discussed above in relation to the methods of the application are equally applicable to the systems of the application. The system can comprise any number of first adapters, RT primers, and second adapters discussed above in relation to the kits of the application. The system preferably comprises the same number of first adapters, RT primers, and second adapters.

[0328] The nanopore is preferably present in a membrane. Suitable membranes are discussed above. The system can comprise any of the membranes disclosed above, such as an amphiphilic layer, a triblock copolymer membrane, or a solid state membrane. The membrane is typically part of an array of membranes, wherein each membrane preferably comprises a nanopore. The array can be any of the arrays described in WO 2018 / 060740, which is incorporated herein by reference in its entirety.

[0329] The system is preferably adapted to apply a voltage across the membrane and to take one or more electrical measurements. Suitable adaptations are discussed in WO 2018 / 060740, which is incorporated herein by reference in its entirety.

[0330] The kit or system preferably further comprises one or more sequencing adapters. The one or more sequencing adapters can be any of the sequencing adapters discussed above in relation to the methods of the application.

[0331] The kit or system can further comprise a polynucleotide binding protein. The kit or system can further comprise a microparticle. Any of the embodiments discussed above in relation to the methods of the application are equally applicable to the systems of the application.

[0332] The kit or system can additionally comprise one or more other reagents or instruments that enable any of the embodiments described above. Such reagents or instruments include one or more of the following: a suitable buffer (aqueous solution), a device to obtain a sample from a subject (such as a container or instrument comprising a needle), a device to amplify and / or express the polynucleotide, a membrane or voltage or patch clamp apparatus as defined above. The reagents can be present in the kit or system in a dry state, such that a fluid sample is used to resuspend the reagents. The kit or system can also optionally comprise instructions that enable the kit or system to be used in the methods described herein or details as to which organisms the methods can be used for. The kit or system can comprise a magnet or electromagnet. The kit or system can optionally comprise nucleotides.

[0333] The following examples illustrate the application. It is to be understood that while the application has been discussed in terms of particular embodiments, particular constructions, and materials and / or molecules in accordance with the application, various changes and modifications could be made in form and detail without departing from the scope and spirit of the application. The following examples are provided to better illustrate particular embodiments and should not be taken as limiting the application. The application is only limited by the claims.

[0334] Example 1

[0335] This example describes one exemplary embodiment of the method of the application. This is referred to as the ligation + RT + ligation (LRL) method.

[0336] The LRL method applies a first round of barcoding by ligating a cDNA reverse transcription adaptor (CDRA) barcodes to isolate a sample of cells Figure 1 ). Following ligation, USER enzyme can be used to digest the bottom strand and reveal a primer site for reverse transcription. Ligation and subsequent USER digestion requires a total of 45 minutes incubation at 37°C.

[0337] The cells are pooled and then split out for a second round of barcoding by barcoding RT primers that anneal to the top strand of the first adaptor Figure 2 In the LRL method, a strand switch primer (SSP) can be used that contains an optimised UMI sequence with reduced homopolymers through a TTVVVVT structure to facilitate detection by our nanopore Figure 2 The SSP facilitates deduplication of PCR replicates.

[0338] Following reverse transcription, the cells are pooled and then split out for a third round of barcoding by ligation of barcoded adaptors Figure 3 .

[0339] In the LRL method, the first round of connections are separated from the third round of connections by a second round of RT steps. This minimizes the potential for barcode cross-talk between connections. In any case, a third round of adapter blocker has been developed that is fully complementary to the adapter overhang. The blocker itself is 5' phosphorylated and thus should permanently attach to the excess adapter once annealed in place Figure 4 ).

[0340] The fully barcoded molecules are then enriched by biotin-streptavidin pull down and amplified with PCR primers. These primers can use click chemistry moieties for fast ligation chemistry, thus speeding up library preparation. In addition, universal primers for the barcodes can be used for potential fourth round of barcode encoding.

[0341] Example 2

[0342] Oligonucleotides used

[0343] Oligonucleotides were resuspended in ribonuclease-free TE buffer (10 mM Tris.HCl, 1 mM EDTA pH 8.0, final) to 100 mM.

[0344]

[0345]

[0346]

[0347] Methanol fixation of cells

[0348] Based on Chen, J., Cheung, F., Shi, R., Zhou, H., Lu, W., Candia, J., Kotliarov, Y., Stagliano, K. R., and Tsang, J. S. (2018). PBMC fixation and processing for Chromium single-cell RNA sequencing. Journal of Translational Medicine, 16(1). https: / / doi.org / 10.1186 / s12967-018-1578-4.

[0349] 1. The process of cell fixation and subsequent rehydration results in approximately 50% cell loss. Proceed with the number of cells appropriate for the scale of experiment desired. For example, if 1E6 cells are desired, start with 2E6 cells for methanol fixation.

[0350] 2. Centrifuge cells at 300 g for 5 min at 4°C.

[0351] 3. Discard supernatant and resuspend in 1 mL ice cold 1X DPBS (Dulbecco's Phosphate Buffered Saline, Gibco, 14080089) using wide bore pipette tips.

[0352] 4. Transfer cells to centrifuge tubes appropriate for the number of cells being input:

[0353]

[0354] 5. 300 g, 4°C for 5 min.

[0355] 6. Discard supernatant and resuspend in 1 mL ice cold 1X DPBS (Gibco, 14190144) using wide bore pipette tips.

[0356] 7. Remove supernatant, resuspend with ice cold 1X DPBS (Gibco, 14190144) to achieve approximately 5E3 cells / μL. For example, 1E6 cells require 200 μL 1X DPBS (Gibco, 14190144).

[0357] 8. Transfer cell suspension to fume hood.

[0358] 9. Take 4x volume of ice cold 100% methanol relative to the volume of resuspended cells and add dropwise, gently continuously swirling to achieve a final of 80% methanol.

[0359] 10. Cool at -20°C for at least 30 minutes.

[0360] 11. Store at -80°C or at -20°C for up to 6 weeks. Alternatively, proceed immediately to the following steps.

[0361] Preparation of barcode oligonucleotides

[0362] The following steps describe an experiment at lx, 50x or lOOx scale, which is relative to 10 K, 500 K or 1 M cells of combinatorially barcoded input material, respectively. Continue with the desired scale of experiment.

[0363] 12. Prepare oligonucleotide pools:

[0364]

[0365]

[0366] 13. Take CDRA adapters and round 3 adapters, heat to 95°C for 5 min, cool at -0.1°C per second to anneal the strands. Store at 4°C.

[0367] Rehydration

[0368] 14. Prepare rehydration buffer according to the number of cells:

[0369]

[0370] 15. Place fixed cells on ice for 5 min.

[0371] 16. Centrifuge at 1000 g for 5 min at 4°C.

[0372] 17. Remove supernatant and resuspend in 1000 μL rehydration buffer.

[0373] 18. Centrifuge at 1000 g for 5 min at 4°C.

[0374] 19. Remove supernatant. Assuming 50% of the cells have been lost, resuspend the cell pellet in rehydration buffer to achieve 1250 cells / μL. For example, 1E6 cells input requires 400 μL rehydration buffer.

[0375] CDRA ligation (first round of barcode encoding)

[0376] 20. Prepare CDRA ligation mix as follows:

[0377]

[0378]

[0379] 21. Prepare a new plate with 4 μL / well of CDRA barcodes on ice.

[0380] 22. Add CDRA ligation mix to each of the CDRA adapters:

[0381]

[0382] 23. Incubate at 37°C for 30 min.

[0383] 24. Now add the following components to expose the RT primer site:

[0384]

[0385] 25. Incubate at 37°C for 15 min.

[0386] 26. Pool and centrifuge at 1000 g for 5 min at 4°C.

[0387] 27. Remove supernatant and resuspend in 1000 μΐ^resuspension buffer.

[0388]

[0389] In situ reverse transcription (second round of barcode encoding)

[0390] 28. Prepare RT mix:

[0391]

[0392] 29. On ice, prepare a new plate with 4 μΐ^ / well of the RT oligo mix.

[0393] 30. Place the RT barcode plate on a chilled rack and In order Add the following to each well:

[0394]

[0395] 31. Seal the plate and then briefly centrifuge.

[0396] 32. Incubate at 42°C for 90 min and then immediately cool to 4°C.

[0397] 33. Briefly centrifuge the plate and then place on a chilled rack.

[0398] 34. Pool the wells into a single 1.5 mL microcentrifuge tube on ice.

[0399] 35. Add 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction to a final 0.1% (v / v) ECOSURF EH9.

[0400] 36. Centrifuge at 1000 g for 5 min at 4°C.

[0401] 37. Discard the supernatant and resuspend the cells in ice-cold IX NEB buffer r3.1 (NEB, B6003S) to 500 cells / μL.

[0402] Ligation of third barcode (third round of barcode encoding)

[0403] 38. On ice, prepare a new plate with 10 μΐ^ / well of the 3rd round ligation barcode.

[0404] 39. On ice, add the following components to the resuspended cells to make the 3rd round ligation mix:

[0405]

[0406] 40. Add 40 μL of the mixture to each well containing 10 μL of the ligation barcodes, mix by gentle pipetting. Seal the plate.

[0407]

[0408] 41. Incubate at 37°C for 30 min.

[0409] 42. Prepare the third round of blocking solution:

[0410]

[0411] 43. Add 20 μL of the third round of blocking solution to each well:

[0412]

[0413] 44. Incubate on ice for 5 min.

[0414] 45. Pool the cells into a single tube.

[0415] Dissolution

[0416] 46. Add 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to the pooled reaction to achieve a final 0.1% (v / v / ) ECOSURF EH9.

[0417] 47. Centrifuge the cells at 1000 g for 5 min at 4°C.

[0418] 48. Prepare the wash buffer:

[0419]

[0420] 49. Carefully remove the supernatant. Wash the cells in 1 mL of the wash buffer without disturbing the pellet.

[0421] 50. Centrifuge at 1000 g for 5 min at 4°C.

[0422] 51. Remove the supernatant without disturbing the pellet. Resuspend the cells in 30 μL of the wash buffer.

[0423] 52. Mix 5 μL of the cell suspension with 5 μL of trypan blue to count the cells. Expect a few thousand cells in total. If necessary, the cells can be split into sub-libraries.

[0424] 53. Add 30 μL lysis buffer to the remaining 25 μL cell suspension.

[0425]

[0426] 54. Incubate for 1 hour at 55°C.

[0427] 55. If desired, freeze the lysate at -20°C.

[0428] cDNA purification

[0429] 56. Take 5 μL of the lysate and dilute to 25 μL with nuclease-free H2O.

[0430] 57. Purify using NEB monarch PCR purification kit (NEB, T1030S) using a 5: 1 buffer: sample ratio as described by the manufacturer.

[0431] 58. Elute in 52 μL nuclease-free H2O. Quantify the eluate by Qubit fluorometer (Invitrogen).

[0432] Pull-down

[0433] Preparation of beads

[0434] 59. Prepare 4 mL 2X wash / binding buffer (10 mM Tris.HCl pH 7.5, 2 M NaCl, 1 mM EDTA).

[0435] 60. Use 3.5 mL 2X wash / binding buffer and add 3.5 mL nuclease-free H2O to make 7 mL IX wash / binding buffer (5 mM Tris.HCl pH 7.5, 1 M NaCl, 0.5 mM EDTA).

[0436] 61. Take 25 μL stock MyOne Cl streptavidin beads (10 μg beads / μL in PBS pH 7.4, 0.1% BSA, 0.02% sodium azide, 65001 Invitrogen) ensuring the stock is fully resuspended.

[0437] 62. Wash the beads 3 times with 1 mL IX wash / binding buffer, vortex for 5 sec and then collect on a magnet for 2 minutes between each wash.

[0438] 63. Resuspend the beads in 50 μL (twice the original volume of beads) of 2X wash / binding buffer to achieve a final bead concentration of 5 μg / μL of beads.

[0439] Sample binding

[0440] 64. Take 50 μL of biotinylated cDNA and combine with 50 μL of 5 μg / μL prepared beads (total of 250 μg beads) to achieve a final buffer concentration of 1 M NaCl optimal for binding.

[0441] 65. Incubate at room temperature with gentle rotation for 20 minutes.

[0442] 66. Wash beads 3 times with 1 mL IX wash / binding buffer. Vortex gently for 5 seconds and collect on a magnet for 3 minutes between each wash. Discard supernatant after each wash. Be careful not to aspirate any beads.

[0443] 67. Wash beads once in 200 μL 10 mM Tris.HCl pH 7.5, vortex for 5 seconds, spin down sample briefly and collect on a magnet for 3 minutes. Discard supernatant.

[0444] 68. Resuspend pellet in 80 μL nuclease-free H2O, vortex for 5 seconds, and then spin down to collect amplicon-bead conjugate.

[0445] Post-pull-down PCR

[0446] 69. Set up the following PCR reaction, divided into 2 aliquots of 100 μL each:

[0447]

[0448] Run in a thermocycler with the following parameters:

[0449]

[0450] SPRI clean-up

[0451] 70. Pool PCR reactions.

[0452] 71. Perform a 0.8X SPRI clean-up using AMPure XP reagent (Beckman Coulter, A63882), wash pellet twice with 70% EtOH, without resuspension.

[0453] 72. Elute in 60 μL nuclease-free H2O and quantify using a Qubit fluorometer (Invitrogen).

[0454] Sequencing library preparation

[0455] 73. Sequencing libraries were prepared using the Oxford Nanopore Technologies SQK-LSK114 Ligation Sequencing Kit. The manufacturer’s protocol was followed and 200 fmol of amplified cDNA was used as input. Sequencing was performed on a PromethION with a FLO-PRO114M flow cell.

[0456] Example 3 - RNA Integrity

[0457] Integrity during methanol fixation

[0458] To test if methanol fixation is suitable for preserving RNA integrity, cell culture stocks of GM12878 (Coriel Institute, GM12878) were fixed in methanol for testing.

[0459] 2E7 cells of GM12878 were washed in 1 mL ice-cold IX DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated and the pellet was washed and pelleted again.

[0460] The supernatant was aspirated and the pellet was resuspended to approximately 5E3 cells / µL in ice-cold IX DPBS.

[0461] The cell suspension was then mixed with 4x volume of ice-cold 100% methanol to achieve a final concentration of 80% methanol. The suspension was chilled at -20°C for 30 minutes.

[0462] The methanol-fixed cell suspension was then split into 4 equal aliquots of 5E6 fixed cells, incubated on ice for 5 minutes and then pelleted at 1000 g.

[0463] The supernatant was aspirated and the cell pellet was resuspended in rehydration buffer of 3X SSC (Invitrogen, AM9770) or IX DPBS (Gibco, 14080089) supplemented with 1 mM DTT (Sigma, 43816), 200 ng / µL recombinant albumin (NEB, B9200s) and 0.2 U / µL SUPERnaseIn RNase inhibitor (Invitrogen, AM2694) to a final volume of 1 mL.

[0464] The cell suspension was pelleted again at 1000 g.

[0465] The supernatant was aspirated and the cell pellet (assuming 50% cell loss) was resuspended in 2000 μΐ, of rehydration in the same composition as before.

[0466] The resuspended cells were then pelleted at 1000 g and the supernatant was aspirated.

[0467] Total RNA was then extracted from the cell pellet following the method described by Workman et al. (2018) immediately after extraction based on TRIzol reagent (Invitrogen, 15596026). Briefly, 400 μΐ, of TRIzol reagent was added to each sample and incubated at room temperature for 5 minutes. 80 μΐ, of chloroform was then added and the sample was vortexed and then incubated at room temperature for 5 minutes. The sample was vortexed and then pelleted at 2000 g at 4°C. The supernatant was transferred to a 5 mL falcon and mixed with an equal volume of isopropanol. The tube was inverted to mix and incubated at room temperature for 15 minutes and then centrifuged at 2000 g at 4°C for 20 minutes. The supernatant was aspirated and the pellet was washed with 400 μΐ, of ice-cold 70% ethanol and then further centrifuged at 2000 g at 4°C for 5 minutes. The supernatant was aspirated and the pellet was resuspended in 10 μΐ, of TE buffer.

[0468] The RNA integrity was then assessed using an Agilent 2100 Bioanalyzer (Agilent, G2939BA) with RNA Nano 6000 kit and chip (Agilent, 5067-1511).

[0469] The results are shown in Figure 5 .

[0470]

[0471] Integrity during incubation

[0472] To test whether the RNA integrity remains stable during the 37°C incubation step required for CDRA adaptor ligation and USER digestion in the first round of combinatorial barcoding in the LRL method, cell culture stocks of GM12878 (Coriel Institute, GM12878) were fixed in methanol for testing.

[0473] 2E7 cells of GM12878 were washed in 1 mL of ice-cold 1X DPBS (Gibco, 14080089) and pelleted at 300 g. The supernatant was aspirated, the pellet was washed and pelleted again.

[0474] Aspirate the supernatant and then resuspend the pellet to approximately 5E3 cells / μL in ice-cold 1X DPBS.

[0475] The cell suspension is then mixed with 4x volume of ice-cold 100% methanol to achieve a final concentration of 80% methanol. The suspension is chilled at -20°C for 30 minutes.

[0476] A portion of the methanol-fixed cell suspension is then divided into two equal aliquots of 5E6 fixed cells and incubated on ice for 5 minutes and then pelleted at 1000 g.

[0477] Aspirate the supernatant and resuspend the cell pellet in 3X SSC (Invitrogen, AM9770) rehydration buffer supplemented with 1 mM DTT (Sigma, 43816), 200 ng / μL recombinant albumin (NEB, B9200s), and 0.2 U / μL SUPERnaseIn RNase inhibitor (Invitrogen, AM2694) to a final volume of 1 mL.

[0478] The cell suspension is again pelleted at 1000 g.

[0479] Aspirate the supernatant and then resuspend the cell suspension (assuming 50% cell loss) in 2000 μL of rehydration with the same composition as before.

[0480] At this point, one aliquot of cells is immediately extracted as follows for an unincubated control. The second aliquot is incubated at 37°C for 45 minutes, which is required for CDRA ligation and USER digestion of the LRL method.

[0481] The rehydrated cells are then pelleted at 1000 g and the supernatant aspirated.

[0482] Total RNA was then extracted from cell pellets immediately after extraction based on TRIzol reagent (Invitrogen, 15596026) following the method described by Workman et al. (2018). Briefly, 400 μΐ of TRIzol reagent was added to each sample and incubated for 5 min at room temperature. Then 80 μΐ of chloroform was added and the sample was vortexed and then incubated for 5 min at room temperature. The sample was vortexed and then pelleted at 2000 g at 4°C. The supernatant was transferred to a 5 mL falcon and mixed with an equal volume of isopropanol. The tube was inverted to mix and incubated for 15 min at room temperature and then centrifuged at 2000 g for 20 min at 4°C. The supernatant was aspirated and the pellet was washed with 400 μΐ of ice-cold 70% ethanol and further centrifuged at 2000 g for 5 min at 4°C. The supernatant was aspirated and the pellet was resuspended in 10 μΐ of TE buffer.

[0483] RNA integrity was then assessed using an Agilent 2100 Bioanalyzer (Agilent, G2939BA) with RNA Nano 6000 kit and chip (Agilent, 5067-1511).

[0484] Results are shown in Figure 6 .

[0485]

[0486] Example 4 - Barcode Crosstalk

[0487] To assess RNA manipulation at the molecular level, we used a specific synthetic RNA analyte derived from the enolase II gene of S. cerevisiae (ENO2, YHR174W SGDID: S000001217, chrVIII: 451327..452640). The ENO2 gene was amplified from S. cerevisiae genomic DNA using primers with a gene-specific portion complementary to the first and last 22 nt of the coding sequence, and a 5' flanking portion of 27 nt of exogenous sequence for downstream applications (Table 1).

[0488]

[0489] Table , YHR174W primer; enolase II gene-specific portion (green) and 5' flanking (blue).

[0490] The amplicon product was then further amplified using an upstream primer flanking a 5' T7 RNA polymerase promoter sequence and a downstream primer flanking a 5' poly(T) tail to obtain an in vitro transcription template (see ).

[0491]

[0492] Table , IVT template primer; T7 RNAP promoter sequence (red) and 3' poly(T) tail (yellow). The transcription start site is underlined.

[0493] The in vitro transcription template is then used in an in vitro transcription reaction with T7 RNAP to obtain the RNA control strand (RCS) test analyte (SEQ ID NO: 18), as follows:

[0494] >RCS

[0495] GGGAGA

[0496] First round of barcoding

[0497] When cells are pooled together, barcodes in the first round of barcode encoding can potentially be shared between cells. To investigate this, we performed an in vitro barcode competition assay using RCS test analyte.

[0498] Given that 1 cell produces approximately 1.25 pg of polyadenylated RNA, and that each singleplex workflow requires 10 K cells, 12.5 ng of RCS test analyte was used as a template for the in vitro barcode competition assay.

[0499] Five reactions (each consisting of 12.5 ng of RCS synthetic analyte) were resuspended in rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng / µL recombinant albumin (NEB, B9200S), and 0.2 U / µL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694) for a final volume of 16 µL to mimic the input of 10 K rehydrated cells.

[0500] Two reactions were then mixed with CDRA BC01 (2.27 µM of SEQ ID NO: 1 top strand hybridized to 2.27 µM of SEQ ID NO: 3 bottom strand), two reactions were mixed with CDRA BC02 (2.27 µM of SEQ ID NO: 2 top strand to 2.27 µM of SEQ ID NO: 3 bottom strand), one reaction was mixed with nuclease-free water (as a control for no CDRA adapter ligation) and 0.625 U / µL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L), and 25 U / µL T4 ligase (NEB, M0202L) for a final volume of 20 µL and incubated at 37°C for 30 minutes. Meanwhile, separate competitive barcode CDRA adapter ligation reactions were assembled as above, without RNA template, and incubated on ice.

[0501] Following ligation, each reaction was supplemented with 0.25 U / µL lambda exonuclease (NEB, M0262S) and 0.05 U / µL USER mix (NEB, M5505S) for a final volume of 22 µL and incubated at 37°C for 15 minutes, followed by incubation on ice for 5 minutes.

[0502] To simulate potential competition between CDRA barcodes during pooling, RNA template CDRA ligation reactions were mixed with equal volume reactions with no competing CDRA adaptors (no competition), or with competing CDRA adaptors (competition). The pooled reactions were then incubated on ice for 5 minutes.

[0503] To simulate cell pooling, centrifugation, and resuspension, each reaction was subjected to a 0.8X SPRI clean up with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 8 pL of IX NEB Buffer r3.1 (NEB, B6003S) supplemented with 0.5 U / pL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694).

[0504] Each reaction was then supplemented with 25 pM RT primer BC03 (SEQ ID NO: 4), 25 pM strand transfer primer (SEQ ID NO: 6), IX RT Buffer (ThermoFisher Scientific, EP0753), 0.5 U / pL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U / pL Maxima H minus Reverse Transcriptase (ThermoFisher Scientific, EP0753) for a final volume of 20 pL and incubated at 42°C for 90 minutes, followed by 5 minutes on ice.

[0505] Each reaction was then supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to a final concentration of 0.1% (v / v) ECOSURF EH9 in the pooled reaction. To simulate cell pooling, centrifugation, and resuspension, each reaction was subjected to a 0.8X SPRI clean up with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 20 pL of IX NEB Buffer r3.1 (NEB, B6003S).

[0506] Each reaction was then supplemented with ligation barcodes BC05 (2.16 mM SEQ ID NO: 9 top strand and 2.4 mM SEQ ID NO: 7 bottom strand hybridised), IX T4 ligation buffer (NEB, M0202L), 20 U / pL T4 ligase (NEB, M0202L) to a final volume of 50 pL and incubated at 37 °C for 30 minutes.

[0507] Each reaction was then supplemented with 3.24 pM blocking oligo (SEQ ID NO: 11) and 31.25 mM EDTA to a final volume of 70 pL and incubated on ice for 5 minutes.

[0508] To simulate cell lysis, each reaction was supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) and then subjected to 1.5X SPRI clean up with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol and resuspended in 10 pL nuclease free water.

[0509] Samples were amplified using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) and hot start LongAmp Taq (NEB, M0533S) for over 16x cycles. Amplicons were subjected to 0.5X SPRI clean up with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol and resuspended in 50 pL nuclease free.

[0510] 200 fmol of amplified cDNA was prepared for sequencing using the Oxford Nanopore Technologies SQK-LSK114 Ligation Sequencing Kit and each sample was sequenced on a PromethION with a FLO-PRO114M flow cell. Reads were aligned to the reference for RCS using MinKNOW software.

[0511] Viewed Figure 7With the results in FIG. 6B, there was only background misalignment detection of the competing barcodes (<4%) without the addition of competing CDRA adapters. However, with the addition of competing barcodes, there was still misalignment detection of the same competing barcodes. These data suggest that there is no apparent cross-talk between CDRA barcodes, likely due to the USER enzyme mix digestion, which destroys the bottom strand of the adapter required to plate the 3’ end of the RNA poly(A) tail to the 5’ end of the top strand of the CDRA adapter.

[0512] With the addition of competing barcodes before sample pooling, there was 11.4% detection of competing background, higher than the background detection in the other samples. However, the PCR yield of this reaction was 8.5-fold lower than the other samples, suggesting that there is still little cross-talk of adapters upon sample pooling.

[0513] In summary, these results show that the design and application of the first round of combinatorial barcoding of CDRA adapters in the LRL method is robust to barcoding cross-talk.

[0514] Second round of barcoding

[0515] When cells are pooled together, there is a possibility that barcodes in the second round of barcoding are shared between cells. To investigate this, we performed an in vitro barcode competition assay using RCS test analyte.

[0516] Assuming 1 cell produces approximately 1.25 pg of polyadenylated RNA, and each singleplex workflow requires 10 K cells, we used 12.5 ng of RCS test analyte as a template for the in vitro barcode competition assay.

[0517] Four reactions (each consisting of 12.5 ng of RCS synthetic analyte) were resuspended in rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng / µL recombinant albumin (NEB, B9200S), and 0.2 U / µL SUPERnaseIn RNase inhibitor (Invitrogen, AM2694) for a final volume of 16 µL to mimic the input of 10 K rehydrated cells.

[0518] The four reactions were then mixed with CDRA BC01 (2.27 mM top strand of SEQ ID NO: 1 and 2.27 mM bottom strand of SEQ ID NO: 3), 0.625 U / miL SUPERnaseIn ribonuclease inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L), and 25 U / miL T4 ligase (NEB, M0202L) for a final volume of 20 pL and incubated at 37 °C for 30 minutes.

[0519] Following ligation, each reaction was supplemented with 0.25 U / miL lambda exonuclease (NEB, M0262S) and 0.05 U / miL USER mix (NEB, M5505S) for a final volume of 22 pL and incubated at 37 °C for 15 minutes, followed by incubation on ice for 5 minutes.

[0520] To simulate cell pooling, centrifugation, and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in IX NEB Buffer r3.1 (NEB, B6003S) supplemented with 0.5 U / miL SUPERnaseIn ribonuclease inhibitor (Invitrogen, AM2694) for a final volume of 8 pL.

[0521] The reactions were then supplemented with 25 pM RT primer BC03 or BC04 (SEQ ID NO: 4 or SEQ ID NO: 5, respectively), 25 pM strand transfer primer (SEQ ID NO: 6), IX RT buffer (ThermoFisher Scientific, EP0753), 0.5 U / miL SUPERnaseIn ribonuclease inhibitor (Invitrogen, AM2694), 500 pM dNTPs (NEB, N0447S), 10 U / miL Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753) for a final volume of 20 pL and incubated at 42 °C for 90 minutes, followed by incubation on ice for 5 minutes. Simultaneously, a separate competitive barcode RT primer reaction was assembled as above, without RNA template, and incubated on ice.

[0522] Each reaction was then supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to a final of 0.1% (v / v) ECOSURF EH9 pooled reaction. To simulate potential competition between RT primers during pooling, RNA template RT reactions were mixed with an equal volume reaction without competing RT primers (no competition), or with an equal volume reaction with competing RT primers (competition). Pooled reactions were then incubated on ice for 5 minutes.

[0523] To simulate cell pooling, centrifugation, and resuspension, each reaction was 0.8X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 20 μL 1X NEB Buffer r3.1 (NEB, B6003S).

[0524] Each reaction was then supplemented with ligation barcode BC05 (2.16 μM SEQ ID NO: 9 top strand hybridized to 2.4 μM SEQ ID NO: 7 bottom strand), IX T4 ligation buffer (NEB, M0202L), 20 U / μL T4 ligase (NEB, M0202L) to a final of 50 μL, and incubated at 37°C for 30 minutes.

[0525] Each reaction was then supplemented with 3.24 μM blocking oligo (SEQ ID NO: 11) and 31.25 mM EDTA to a final of 70 μL, and incubated on ice for 5 minutes.

[0526] To simulate cell lysis, each reaction was supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006), then 1.5X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 10 μL nuclease-free water.

[0527] Samples were amplified for over 16x cycles using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) and hot-start LongAmp Taq (NEB, M0533S). Amplicons were 0.5X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 50 μL nuclease-free.

[0528] 200 fmol of amplified cDNA was prepared for sequencing using the Oxford Nanopore Technologies SQK-LSK114 Ligation Sequencing Kit and sequenced on a PromethION with a FLO-PRO114M flow cell. Reads were aligned to the reference for the RCS using MinKNOW software.

[0529] Looking at the results in FIG. 6A, there is only background mis-alignment detection (<2%) for the competing barcodes without the addition of competing RT primers. However, with the addition of competing barcodes, there is still mis-alignment detection (<2%) for the same competing barcodes. These data suggest that there is no significant cross-talk between the RT barcodes, likely because the reverse transcriptase has very low activity when incubated on ice and is unable to extend any cross primed templates. The PCR yield for the samples is consistent, suggesting that there is no significant difference in cDNA synthesis with or without RT competition. Figure 8 In summary, these results show that the design and application of the RT barcodes (specifically complementary to the top strand of the CDRA adapter) applied in the second round of combinatorial barcode encoding in the LRL method is robust to barcode cross-talk.

[0530]

[0531] Third round of barcoding When cells are pooled together, there is a possibility that barcodes in the third round of barcode encoding are shared between cells. The application of blocking oligos can reduce this cross-talk. To investigate this, we performed an in vitro barcode competition assay using the RCS test analyte with no blocking oligos, the blocking oligos used in the LRL method, or alternative blocking oligos.

[0532]

[0533] Assuming 1 cell produces approximately 1.25 pg of polyadenylated RNA and 10 K cells are needed per singleplex workflow, 12.5 ng of the RCS test analyte was used as a template for the in vitro barcode competition assay.

[0534] ​Nine reactions (each consisting of 12.5 ng RCS synthetic analyte) were resuspended in rehydration buffer of 3X SSC (Invitrogen, AM9770) supplemented with 1 mM DTT (Sigma, 43816), 200 ng / µL recombinant albumin (NEB, B9200S), and 0.2 U / µL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694) for a final volume of 16 µL to mimic an input of 10 K rehydrated cells.

[0535] Four reactions were then mixed with CDRA BC01 (2.27 µM SEQ ID NO: 1 top strand hybridized to 2.27 µM SEQ ID NO: 3 bottom strand), 0.625 U / µL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694), IX T4 ligation buffer (NEB, M0202L), and 25 U / µL T4 Ligase (NEB, M0202L) for a final volume of 20 µL and incubated at 37°C for 30 minutes.

[0536] Following ligation, each reaction was supplemented with 0.25 U / µL Lambda Exonuclease (NEB, M0262S) and 0.05 U / µL USER Mix (NEB, M5505S) for a final volume of 22 µL and incubated at 37°C for 15 minutes, followed by an incubation on ice for 5 minutes.

[0537] To mimic cell pooling, centrifugation, and resuspension, each reaction was subjected to a 0.8X SPRI clean with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in IX NEB Buffer r3.1 (NEB, B6003S) supplemented with 0.5 U / µL SUPERnaseIn RNase Inhibitor (Invitrogen, AM2694) for a final volume of 8 µL.

[0538] Each reaction was then supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to a final 0.1% (v / v) ECOSURF EH9 pooled reaction and incubated on ice for 5 minutes.

[0539] Each reaction was then supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) to a final 0.1% (v / v) ECOSURF EH9 pooled reaction and incubated on ice for 5 minutes.

[0540] To simulate cell pooling, centrifugation, and resuspension, each reaction was 0.8X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 20 μL of IX NEB Buffer r3.1 (NEB, B6003S).

[0541] Each reaction was then supplemented with no third round ligation barcode, BC05 (2.16 μM SEQ ID NO: 9 top strand hybridized to 2.4 μM SEQ ID NO: 7 bottom strand) or BC06 (2.16 μM SEQ ID NO: 9 top strand hybridized to 2.4 μM SEQ ID NO: 8 bottom strand), IX T4 ligation buffer (NEB, M0202L), 20 U / μL T4 ligase (NEB, M0202L) to a final of 50 μL and incubated at 37°C for 30 minutes.

[0542] Each reaction was then supplemented with no blocking oligo, 3.24 μM standard blocking oligo (SEQ ID NO: 10) or 3.24 μM alternative blocking oligo (SEQ ID NO: 11) and 31.25 mM EDTA to a final of 70 μL and incubated on ice for 5 minutes.

[0543] To simulate potential competition between 3rdround adapters during pooling, cDNA template RT reactions were mixed with an equal volume reaction without competing 3rdround barcodes (no competition) or with an equal volume reaction with competing 3rdround barcodes (competition). The pooled reactions were then incubated on ice for 5 minutes.

[0544] To simulate cell lysis, each reaction was supplemented with 0.01 volume of 10% (v / v) ECOSURF EH9 (Sigma, STS0006) and then 1.5X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 10 μL nuclease-free water.

[0545] Samples were amplified for over 16x cycles using universal primers (SEQ ID NO: 12 and SEQ ID NO: 13) and hot-start LongAmp Taq (NEB, M0533S). Amplicons were 0.5X SPRI cleaned with AMPure XP reagent (Beckman Coulter, A63882), washed with 70% ethanol, and resuspended in 50 μL nuclease-free.

[0546] 200 fmol of amplified cDNA were prepared for sequencing using the Oxford Nanopore Technologies SQK-LSK114 Ligation Sequencing Kit and each sample was sequenced on a PromethION with FLO-PRO114M flow cells. Reads were aligned to the RCS reference using MinKNOW software.

[0547] Looking at the results in Figure 9 without a competing 3rdround ligation barcode, only background misalignment detection of the competing barcode (<0.2%). However, if the blocking oligo was not used with the addition of the competing barcode, detection of the competing barcode increased significantly (10.7%). Likewise, if the 3rdround ligation barcode was not present in the ligation prior to pooling the competition, there was still significant incorporation of the competing barcode (51.1%) and a corresponding higher amplicon yield from the PCR. This indicates that there is a possibility of barcode cross-talk during the pooling of the 3rdround ligation reactions, even in the presence of 30 mM EDTA.

[0548] From Figure 10With the standard blocking oligo using the RLL method, there is still some degree of barcode cross-talk with the competitor barcode (3.4%). If the third round of ligation barcodes are not present in the ligation prior to pooling for competition, there is still incorporation of the competitor barcode (27.0%), but the PCR yield is 4.4-fold lower than if the third round of ligation barcodes were present in the ligation prior to pooling.

[0549] In Figure 11 With the alternative blocking oligo using the LRL method, there is still some degree of barcode cross-talk with the competitor barcode (4.0%). If the third round of ligation barcodes are not present in the ligation prior to pooling for competition, there is still incorporation of the competitor barcode (43.3.0%), but the PCR yield is 13.3-fold lower than if the third round of ligation barcodes were present in the ligation prior to pooling.

[0550] Focusing only on the conditions where there were no third round of ligation barcodes prior to pooling for competition, Figure 12 most clearly demonstrates the effectiveness of using a blocking oligo prior to pooling of samples. Both the standard blocking oligo and the alternative blocking oligo result in a reduction of incorporation of the competitor barcode with a decrease in PCR yield, but the alternative blocking oligo results in the greatest reduction of competitor barcode amplicons.

[0551] In summary, these results show that the design and application of a third round of ligation barcodes can cause cross-talk between barcodes at the time of sample pooling. However, the use of a blocking oligo complementary to the top strand of the third round of ligation barcodes, particularly the alternative blocking oligo used in the LRL method, can help to reduce potential cross-talk.

Claims

1. A method for uniquely labeling RNA molecules in a cell population, the method comprising: (a) Divide the cell population into multiple first samples; (b) A first adaptor is attached to the 3' end of an RNA molecule in one of the plurality of first samples, wherein the first adaptor contains a primer site for reverse transcription (RT), and wherein the first adaptor used in each of the first samples contains a different first barcode of inverse complement; (c) Collect the plurality of first samples and divide the pool into a plurality of second samples; (d) The RNA molecules in the plurality of second samples were reverse transcribed using the primer sites and RT primers to form a double-stranded construct, wherein the RT primers used in each second sample contained a different second barcode; (e) Collect the plurality of second samples and divide the pool into a plurality of third samples; as well as (f) Connecting the second connector to the double-stranded construct in the plurality of third samples to form a triple barcode construct, wherein the second connector used in each third sample contains a different third barcode.

2. The method of claim 1, wherein the first adaptor is double-stranded and includes a protruding end capable of hybridizing with the 3' end of the RNA molecule.

3. The method of claim 2, wherein the protruding end comprises one or more of the following: (i) a thymine-containing nucleotide, (ii) a uracil-containing nucleotide, and (iii) a universal nucleotide.

4. The method according to claim 2 or 3, wherein the inverse complement of the different first barcode is present in the strand of the first adaptor attached to the RNA molecule.

5. The method according to any one of claims 2 to 4, wherein the nucleotide at the 3' end of one of the two strands of the first double-stranded linker is dideoxycytosine (ddC) or reverse dT.

6. The method according to any one of claims 2 to 5, wherein the non-connecting chain of the first connector is removed before step (d).

7. The method according to any one of the preceding claims, wherein step (d) comprises hybridizing the RT primer with the first adaptor.

8. The method of claim 7, wherein the RT in step (d) generates a strand comprising the different first barcode, the different second barcode, and a sequence complementary to the RNA molecule.

9. The method of claim 8, wherein the components are arranged in the following order in chains 5' to 3': different second barcodes, different first barcodes, and complementary sequences.

10. The method of claim 8 or 9, wherein the second connector is connected to a chain comprising the different first barcodes, the different second barcodes, and the complementary sequence.

11. The method of claim 10, wherein the second adaptor comprises a protruding end capable of hybridizing with the RT primer.

12. The method of claim 11, wherein the distinct third barcode is present in the non-protruding end chain of the second connector connected to the double-chain construct.

13. The method of claim 11 or 12, wherein step (f) is performed in the presence of a 5' phosphorylation inhibitor complementary to the protrusion.

14. The method according to any one of the preceding claims, wherein the barcode used in each sample is different from any other barcode or all other barcodes used in the method.

15. The method according to any one of the preceding claims, wherein the number of cells in the population is less than the number of the first sample multiplied by the number of the second sample multiplied by the number of the third sample.

16. The method according to any one of the preceding claims, wherein (i) the plurality of first samples, (ii) the plurality of second samples and / or (ii) the plurality of third samples comprise at least about 24 samples.

17. The method according to any one of the preceding claims, wherein the cell population is fixed and / or permeabilized prior to step (a).

18. The method according to any one of the preceding claims, wherein the method further comprises repeating steps (e) to (f) at least once, including forming a plurality of fourth or more samples and using a third or more connector to produce a quadruple or more barcode construct, the third or more connector comprising a different fourth or more barcode for each fourth or more sample.

19. The method according to any one of the preceding claims, wherein the method further comprises isolating the barcode construct from the cell and / or amplifying the barcode construct.

20. The method according to any one of the preceding claims, wherein the method further comprises characterizing or sequencing the barcode construct, preferably using nanopores.

21. A method for characterizing the transcriptome of a single cell or two or more cells in a cell population, the method comprising: (a) Applying the method according to any one of the preceding claims to the cell population; as well as (b) Characterize RNA molecules from the single cell or two or more cells using their unique markers.

22. The method of claim 21, wherein the characterization comprises sequencing the RNA molecule, preferably using a nanopore.

23. A kit for uniquely labeling RNA molecules in a cell population, the kit comprising two or more first adaptors, wherein the two or more first adaptors contain primer sites for RT, and wherein each of the two or more first adaptors contains a different first barcode inverse complement.

24. The kit of claim 23, wherein the two or more first adaptors are double-stranded and include overhangs capable of hybridizing with the 3' end of the RNA molecule.

25. The kit of claim 24, wherein the protrusion comprises one or more of the following: (i) a thymine-containing nucleotide, (ii) a uracil-containing nucleotide, and (iii) a universal nucleotide.

26. The kit according to claim 24 or 25, wherein the inverse complement of the different first barcode is present in the strand of the two or more first adaptors connected to the RNA molecule.

27. The kit according to any one of claims 24 to 26, wherein the nucleotide at the 3' end of one of the two strands of the two or more first adapters of the double helix is ​​dideoxycytosine (ddC) or reverse dT.

28. The kit according to any one of claims 23 to 27, wherein the kit further comprises: (a) two or more RT primers capable of hybridizing with the two or more first adapters and each containing a different second barcode; and / or (b) two or more second adapters, each containing a different third barcode.

29. The kit according to any one of claims 23 to 28, wherein (a) each barcode in the kit is different from any other barcode or all other barcodes in the kit, and / or (b) the kit comprises at least 96 first adaptors, at least 96 RT primers and / or at least 96 second adaptors.

Citation Information

Patent Citations

  • Biomolecular sensors and methods

    US20170044605A1

  • phi 29 DNA polymerase

    US5576204A

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Deliver of molecules to a li id bila

    WO2006100484A2

  • Lipid bilayer sensor system

    WO2008102120A1