Adaptors and kits for RNA molecules labelling and characterising

RNA-specific adaptors with RNA barcodes and polynucleotide binding proteins facilitate efficient RNA sequencing by stabilizing RNA and controlling its movement, addressing the limitations of DNA-based methods and enhancing sequencing efficiency and cost-effectiveness.

WO2026057755A1PCT designated stage Publication Date: 2026-03-19OXFORD NANOPORE TECH LTD
View PDF 66 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Current RNA sequencing methods rely on DNA adaptors and barcodes, which require additional software or handling for demultiplexing and are not optimized for RNA sequencing, limiting cost-effectiveness and efficiency, especially in low-read and low-complexity applications.

Method used

Development of RNA-specific adaptors with RNA barcodes that stabilize RNA molecules through hybridization to a second polynucleotide strand, allowing direct sequencing without additional software or analysis, and incorporating polynucleotide binding proteins to control RNA movement during sequencing.

Benefits of technology

Enables cost-effective sequencing of multiple RNA samples on a single run, providing high-quality data for applications like targeted RNAseq, shallow human transcriptomics, and virus identification without the need for additional software or handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000065_0001
    Figure IMGF000065_0001
  • Figure IMGF000066_0001
    Figure IMGF000066_0001
  • Figure IMGF000067_0001
    Figure IMGF000067_0001
Patent Text Reader

Abstract

The invention relates to adaptors and kits for uniquely labelling and characterising RNA molecules. The invention also provides methods using such adaptors and kits.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] ADAPTORS AND KITS

[0002] TECHNICAL FIELD

[0003] The invention relates to adaptors and kits for uniquely labelling and characterising RNA molecules. The invention also provides methods using such adaptors and kits.

[0004] INTRODUCTION

[0005] Biological pores (and other nanopores) have great potential as direct, electrical biosensors for polymers and a variety of small molecules. In particular, recent focus has been given to nanopores as a potential DNA sequencing technology. When a potential is applied across a nanopore, there is a change in the current flow when an analyte, such as a nucleotide, resides transiently in the barrel for a certain period of time. Nanopore detection of the nucleotide gives a current change of known signature and duration. In the strand sequencing method, a single polynucleotide strand is passed through the pore and the identities of the nucleotides are derived. Strand sequencing can involve the use of a molecular brake to control the movement of the polynucleotide through the pore.

[0006] Advances in next generation sequencing (NGS) have allowed RNA molecules to be studied and sequenced effectively. For instance, Smith et al., Molecular barcoding of native RNAs using nanopore sequencing and deep learning, Genome research vol. 30,9 (2020): 1345- 1353 (doi: 10.1101 / gr.260836.120) and Toorn et al., Demultiplexing and barcode-specific adaptive sampling for nanopore direct RNA sequencing bioRxiv 2024.07.22.604276 (doi: htti : Joi.orq / 10.1101 / 2024.07.22.604276) disclose methods for barcoding and demultiplexing direct RNA sequencing using nanopores. However, these methods only use DNA adaptors (also known as RTAs) and DNA barcodes.

[0007] SUMMARY OF THE INVENTION

[0008] The inventors have developed novel adaptors for uniquely labelling and characterising, such as sequencing, RNA molecules. The adaptors comprise an RNA strand comprising a RNA barcode. This means the RNA molecules are uniquely labelled with an RNA barcode.

[0009] It is surprising that RNA barcodes can be used for unique labelling because of the instability of RNA. Without wishing to be bound by theory, the inventors believe the RNA barcodes of the invention can be stabilised by hybridisation to a second polynucleotide strand, such as second DNA strand, in the double stranded polynucleotide adaptors of the invention.

[0010] The adaptors of the invention have several key advantages. Currently, there are no manufacturer-provided protocols for molecular barcoding of direct RNA sequencing data sets, which greatly improve the cost-effectiveness of certain RNA sequencing applications by combining multiple samples on the same consumable flow cell. Multiplexing would greatly benefit sequencing designs in which the required number of reads per sample is low or transcriptomes are of low complexity, such as in wtro-transcribed RIMA sequences, ribosomal RNAs, or viral RNA genomes.

[0011] Prior art systems used for sequencing RNA are typically optimised to basecall in RNA. This means additional specialized software or training (e.g., DeePlexiCon) or even additional handling offline post-sequencing is required to sequence DNA barcodes. The adaptors of the invention have RNA barcodes and this means barcode sequencing and any demultiplexing is solely reliant on basecalling of the RNA barcodes. No additional software or analysis is required.

[0012] The primary use case of the invention is enabling researchers to sequence multiple samples on the same sequencing run. Thus, some of potential applications for the invention are, amongst others, targeted RNAseq (16S, tRNA, etc.), shallow human transcriptomics, small transcriptomes, mRNA quality control (QC), virus identification, and circular RNA (circRNA) analysis.

[0013] The invention provides a double stranded polynucleotide adaptor for uniquely labelling a RNA molecule, the double stranded adaptor comprising an overhang capable of hybridising to the RNA molecule and a barcode, wherein the adaptor comprises a first RNA strand comprising the barcode and a second strand comprising the overhang and a non-RNA polynucleotide.

[0014] The invention also provides a population of double stranded polynucleotide adaptors for uniquely labelling a plurality of RNA molecules, the population comprising two or more double stranded polynucleotide adaptors according to any one of the preceding claims, wherein each double stranded polynucleotide adaptor comprises a different barcode.

[0015] The inventors have also developed novel sequencing adaptors for sequencing RNA molecules. These sequencing adaptors comprise a RNA polynucleotide upstream 5' to 3' of a loading site for a polynucleotide binding protein. This means the polynucleotide binding protein being used to control the movement of the target RNA molecule during characterisation or sequencing interacts with the RNA polynucleotide in the sequencing adaptor before it encounters and controls the movement of the target RNA molecule and, if present, the RNA barcode. This ensures the polynucleotide binding protein controls the movement of the target RNA molecule and any RNA barcode in a consistent manner and provides good quality data that may be used to characterise or sequence the target RNA molecule and, if present, the RNA barcode.

[0016] The invention also provides a polynucleotide sequencing adaptor comprising a RNA polynucleotide upstream 5' to 3' of a loading site for a polynucleotide binding protein.

[0017] The invention also provides: a kit for uniquely labelling a RIMA molecule, the kit comprising a double stranded polynucleotide adaptor of the invention and a polynucleotide sequencing adaptor of the invention; a kit for uniquely labelling a plurality of RNA molecules, the kit comprising a population of double stranded polynucleotide adaptors of the invention and a polynucleotide sequencing adaptor of the invention; a method of uniquely labelling a RNA molecule, the method comprising (a) hybridising the RNA molecule to a double stranded polynucleotide adaptor of the invention, and (b) ligating the hybridised double stranded polynucleotide adaptor to the RNA molecule; a method of uniquely labelling a plurality of RNA molecules, the method comprising (a) hybridising the RNA molecules to a population of double stranded polynucleotide adaptors of the invention, and (b) ligating the hybridised double stranded polynucleotide adaptors to the RNA molecules; and a method of characterising a RNA molecule or a plurality of RNA molecules, the method comprising (i) uniquely labelling the RNA molecule(s) using a method of the invention and (ii) characterising the RNA molecule(s) and the barcode(s).

[0018] DESCRIPTION OF THE FIGURES

[0019] It is to be understood that Figures are for the purpose of illustrating particular embodiments of the invention only and are not intended to be limiting.

[0020] Figure 1 : A sequencing adaptor diagram comprising an RNA polynucleotide upstream 5' to 3' [hashed section 1] of a loading site for a polynucleotide binding protein [2].

[0021] Figure 2: A RNA Barcoding adaptor diagram comprising an RNA top strand [1] containing a barcode sequence [hashed section 2] and a DNA bottom strand [3].

[0022] Figure 3: Chromatographic traces that show the typical profile of an RNA barcoding adaptor and its constituent oligonucleotides [1 and 2]. The shift in retention time for the annealed RNA Barcoding adaptor [3] demonstrates that the constituent oligonucleotides are annealed.

[0023] Figure 4: Denaturing gel electrophoresis showing successful annealing of the DNA constituent strand of the RNA Barcoding adaptor [1] and ligation [higher molecular weight species 4] of the RNA constituent strand of the RNA barcoding adaptor [2] to an RNA strand Figure 5: Denaturing gel electrophoresis showing successful annealing of the DNA constituent strand of the RIMA Barcoding adaptor [3] and ligation [higher molecular weight species 4] of the RNA constituent strand of the RNA barcoding adaptor [3] to the sequencing adaptor [1].

[0024] Figure 6: tRNA Barcoding adaptor illustration comprising an RNA top strand [1] which contains the barcode sequence [hashed area 2] , an RNA bottom strand [3] and a DNA bottom strand [4] which acts as a reverse transcription primer.

[0025] Figure 7: The method steps used in Example 5. The first step consists in ligating a barcode / RT adapter (tRNA RTB) that differs from the standard RTB adapter by the overhang (UGGN) and by featuring two bottom strands, one RNA part with an overhang and one DNA RT primer. In a second step the RT enzyme enters in the gap between the 2 bottom strands and makes the reverse complement of the top strand unravelling the tRNA cloverleaf structure. This results in a tRNA / cDNA double-stranded hybrid molecule. The third step consists in ligating the sequencing adapter (RLA). Several RT enzymes were tested in comparison with a mock condition.

[0026] Figure 8: Synthetic tRNA (75 nucleotides) containing only canonical nucleotides was phosphorylated by PNK treatment (lane 1). The tRNA-specific RTB adapter was formed by annealing equimolar amounts of three strands, one long RNA top strand, and the RNA and DNA bottom strands (lane 2). The phosphorylated tRNA and annealed tRNA-RTB adapter were ligated as described in examples X and Y (lane 3). The ligation product was used for reverse transcription to form a tRNA / cDNA hybrid (lane 4). The reverse transcribed material was treated with RNAseH, which selectively digests the RNA in RNA / DNA hybrids (lane 5).

[0027] Figure 9: Reference anchored overlaid signal traces of reads aligned over the 5' and 3' adapters and the tRNA reference for 3 synthetic Yeast tRNAs (His, Tyr and Trp). For each tRNA, we randomly selected 50 reads spanning the entire tRNA and at least half of each of the RTB top and bottom. Reads were processed as described above, then their raw ionic signal traces were realigned to the references using ONT software.

[0028] DESCRIPTION OF THE SEQUENCE LISTING

[0029] SEQ ID NOs: 1-62 are shown and described below.

[0030] DETAILED DESCRIPTION

[0031] It is to be understood that different applications of the disclosed products and methods may be tailored to the specific needs in the art. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments of the invention only and is not intended to be limiting. All publications, patents and patent applications cited herein, whether supra or infra, are hereby incorporated by reference in their entirety. All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.

[0032] The invention, both as to organization and method of operation, together with features and advantages thereof, may best be understood by reference to the following detailed description when read in conjunction with the accompanying Figures. The aspects and advantages of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. Reference throughout this specification to "embodiment(s)" means that a particular feature, structure, or characteristic described in connection with the embodiment(s) is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment", "in another embodiment" or "in a preferred embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may do so. Similarly, it should be appreciated that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment.

[0033] Definitions

[0034] Where an indefinite or definite article is used when referring to a singular noun e.g., "a” or "an", "the", this includes a plural of that noun unless something else is specifically stated. Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. The term "comprising" is interchangeable with "consisting of" or "consisting essentially of".

[0035] Furthermore, the terms first, second, third and the like in the description and in the claims, are used for distinguishing between similar elements and not necessarily for describing a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances and that the embodiments of the invention described herein are capable of operation in other sequences than described or illustrated herein. The following terms or definitions are provided solely to aid in the understanding of the invention. Unless specifically defined herein, all terms used herein have the same meaning as they would to one skilled in the art of the invention. Practitioners are particularly directed to Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016), for definitions and terms of the art. The definitions provided herein should not be construed to have a scope less than understood by a person of ordinary skill in the art.

[0036] "About" as used herein when referring to a measurable value such as an amount, percentage, a temporal duration, and the like, is meant to encompass variations of ± 20 % or ± 10 %, more preferably ± 5 %, even more preferably ± 1 %, and still more preferably ± 0.1 % from the specified value, as such variations are appropriate to perform the disclosed methods. Any embodiment containing the term "about" includes the same feature without the term. For example, about 10 nucleotides or more includes 10 nucleotides or more.

[0037] "Polynucleotide" as used herein refers to a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, this term includes double- and single-stranded DNA, and RNA. The term "polynucleotide" as used herein, is a single or double stranded covalently linked sequence of nucleotides in which the 3' and 5' ends on each nucleotide are joined by phosphodiester bonds. The polynucleotide may be made up of deoxyribonucleotide bases or ribonucleotide bases. Polynucleotides may be manufactured synthetically in vitro or isolated from natural sources. Polynucleotides may further include modified DNA or RNA, for example DNA or RNA that has been methylated, or RNA that has been subject to post- translational modification, for example 5'-capping with 7-methylguanosine, 3'-processing such as cleavage and polyadenylation, and splicing. Polynucleotides may also include synthetic nucleic acids (XNA), such as hexitol nucleic acid (HNA), cyclohexene nucleic acid (CeNA), threose nucleic acid (TNA), glycerol nucleic acid (GNA), locked nucleic acid (LNA), peptide nucleic acid (PNA) and (poly)ADPribose modified DNA. Sizes of polynucleotides are typically expressed as the number of base pairs (bp) or nucleotide pairs for double stranded polynucleotides, or in the case of single stranded polynucleotides as the number of nucleotides (nt). One thousand bp or nt equal a kilobase (kb). Polynucleotides of less than around 40 nucleotides in length are typically called "oligonucleotides" and may comprise primers for use in manipulation of DNA such as via polymerase chain reaction (PCR). The term "polynucleotide" is interchangeable with "polynucleotide sequence", "nucleotide sequence", "DNA sequence", "nucleic acid", or "nucleic acid molecule(s)".

[0038] The term "amino acid" in the context of the present disclosure is used in its broadest sense and is meant to include organic compounds containing amine (NH2) and carboxyl (COOH) functional groups, along with a side chain (e.g., a R group) specific to each amino acid. The amino acids typically refer to naturally occurring L o-amino acids or residues. The commonly used one and three letter abbreviations for naturally occurring amino acids are used herein: A=Ala; C=Cys; D=Asp; E=Glu; F=Phe; G=Gly; H=His; I=Ile; K=Lys; L=Leu; M = Met; N=Asn; P=Pro; Q=Gln; R=Arg; S=Ser; T=Thr; V=Val; W=Trp; and Y=Tyr (Lehninger, A. L., (1975) Biochemistry, 2d ed., pp. 71-92, Worth Publishers, New York). The general term "amino acid" further includes D-amino acids, retro-inverso amino acids as well as chemically modified amino acids such as amino acid analogues, naturally occurring amino acids that are not usually incorporated into proteins such as norleucine, and chemically synthesised compounds having properties known in the art to be characteristic of an amino acid, such as P-amino acids. For example, analogues or mimetics of phenylalanine or proline, which allow the same conformational restriction of the peptide compounds as do natural Phe or Pro, are included within the definition of amino acid. Such analogues and mimetics are referred to herein as "functional equivalents" of the respective amino acid. Other examples of amino acids are listed by Roberts and Vellaccio, The Peptides: Analysis, Synthesis, Biology, Gross and Meiehofer, eds., Vol. 5 p. 341, Academic Press, Inc., N.Y. 1983, which is incorporated herein by reference.

[0039] The terms "polypeptide", and "peptide" are interchangeably used herein to refer to a polymer of amino acid residues and to variants and synthetic analogues of the same. Thus, these terms apply to amino acid polymers in which one or more amino acid residues is a synthetic non-naturally occurring amino acid, such as a chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. Polypeptides can also undergo maturation or post-translational modification processes that may include, but are not limited to glycosylation, proteolytic cleavage, lipidization, signal peptide cleavage, propeptide cleavage, phosphorylation, and such like. A peptide can be made using recombinant techniques, e.g., through the expression of a recombinant or synthetic polynucleotide. A recombinantly produced peptide it typically substantially free of culture medium, e.g., culture medium represents less than about 20 %, more preferably less than about 10 %, and most preferably less than about 5 % of the volume of the protein preparation.

[0040] The term "protein" is used to describe a folded polypeptide having a secondary or tertiary structure. The protein may be composed of a single polypeptide or may comprise multiple polypepties that are assembled to form a multimer. The multimer may be a homooligomer, or a heterooligmer. The protein may be a naturally occurring or wild type protein, or a modified, or non-naturally, occurring protein. The protein may, for example, differ from a wild type protein by the addition, substitution, or deletion of one or more amino acids.

[0041] A "variant" of a protein encompasses peptides, oligopeptides, polypeptides, proteins, and enzymes having amino acid substitutions, deletions and / or insertions relative to the unmodified or wild-type protein in question and having similar biological and functional activity as the unmodified protein from which they are derived. The term "amino acid identity" as used herein refers to the extent that sequences are identical on an amino acid- by-amino acid basis over a window of comparison. Thus, a "percentage of sequence identity" is calculated by comparing two optimally aligned sequences over the window of comparison, determining the number of positions at which the identical amino acid residue (e.g., Ala, Pro, Ser, Thr, Gly, Vai, Leu, He, Phe, Tyr, Trp, Lys, Arg, His, Asp, Glu, Asn, Gin, Cys and Met) occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison ( / .e., the window size), and multiplying the result by 100 to yield the percentage of sequence identity.

[0042] For all aspects and embodiments of the invention, a "variant" has at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98% or at least about 99% complete sequence identity or homology to the amino acid sequence of the corresponding wild-type protein. Sequence identity or homology can also be to a fragment or portion of the full-length polynucleotide or polypeptide. Hence, a sequence may have only at least about 40% overall sequence identity or homology with a full-length reference sequence, but a sequence of a particular region, domain or subunit could share at least about 80%, at least about 90%, or as much as at least about 99% sequence identity or homology (or any of the %s discussed above) with the reference sequence.

[0043] Standard methods in the art may be used to determine identity and / or homology. For example, the UWGCG Package provides the BESTFIT program which can be used to calculate homology, for example used on its default settings (Devereux et al (1984) Nucleic Acids Research 12, p387-395). The PILEUP and BLAST algorithms can be used to calculate homology or line up sequences (such as identifying equivalent residues or corresponding sequences (typically on their default settings)), for example as described in Altschul S. F. (1993) J Mol Evol 36:290-300; Altschul, S.F et al (1990) J Mol Biol 215:403-10. Software for performing BLAST analyses is publicly available through the National Center for Biotechnology Information (http: / / www.ncbi.nlm.nih.gov / ). In all instances herein, identity or homology is typically measured over the entire length of the reference sequence.

[0044] The term "wild-type" refers to a gene or gene product isolated from a naturally occurring source. A wild-type gene is that which is most frequently observed in a population and is thus arbitrarily designed the "normal" or "wild-type" form of the gene. In contrast, the term "modified", "mutant" or "variant" refers to a gene or gene product that displays modifications in sequence (e.g., substitutions, truncations, or insertions), post-translational modifications and / or functional properties (e.g., altered characteristics) when compared to the wild-type gene or gene product. It is noted that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics when compared to the wild-type gene or gene product. Methods for introducing or substituting naturally occurring amino acids are well known in the art. For instance, methionine (M) may be substituted with arginine (R) by replacing the codon for methionine (ATG) with a codon for arginine (CGT) at the relevant position in a polynucleotide encoding the mutant monomer. Methods for introducing or substituting non-naturally occurring amino acids are also well known in the art. For instance, non-naturally occurring amino acids may be introduced by including synthetic aminoacyl-tRNAs in the IVTT system used to express the mutant monomer. Alternatively, they may be introduced by expressing the mutant monomer in E. coli that are auxotrophic for specific amino acids in the presence of synthetic ( / .e., non-naturally occurring) analogues of those specific amino acids. They may also be produced by naked ligation if the mutant monomer is produced using partial peptide synthesis. Conservative substitutions replace amino acids with other amino acids of similar chemical structure, similar chemical properties, or similar side-chain volume. The amino acids introduced may have similar polarity, hydrophilicity, hydrophobicity, basicity, acidity, neutrality, or charge to the amino acids they replace. Alternatively, the conservative substitution may introduce another amino acid that is aromatic or aliphatic in the place of a pre-existing aromatic or aliphatic amino acid. Conservative amino acid changes are well- known in the art.

[0045] A mutant or variant protein can also be chemically modified in any way and at any site. A mutant or variant protein is preferably chemically modified by attachment of a molecule to one or more cysteines (cysteine linkage), attachment of a molecule to one or more lysines, attachment of a molecule to one or more non-natural amino acids, enzyme modification of an epitope or modification of a terminus. Suitable methods for carrying out such modifications are well-known in the art. The mutant or variant protein may be chemically modified by the attachment of any molecule. For instance, the mutant or variant protein may be chemically modified by attachment of a dye or a fluorophore.

[0046] As explained in more detail below, the adaptor of the invention may be used to uniquely label a single RNA molecule or may be used to uniquely label a plurality of RNA molecules that are the same or different. The specific barcode in the adaptor is responsible for the unique labelling. The term "unique label", "uniquely labelling" or "uniquely labelled" is therefore interchangeable with "label", "labelling" or "labelled", respectively.

[0047] Adaptors of the invention

[0048] The invention provides a double stranded polynucleotide adaptor. The adaptor is for uniquely labelling a RNA molecule. In the context of the invention, a RNA molecule is uniquely labelled when it is labelled with a barcode. The barcode is a RNA barcode. All instances of "barcode" or "barcodes" herein are interchangeable with "RIMA barcode" or "RNA barcodes", respectively.

[0049] Target RNA

[0050] The RNA molecule being uniquely labelled may also be known as the target RNA molecule. RNA is a macromolecule comprising two or more ribonucleotides. The target RNA molecule may be eukaryotic or prokaryotic RNA. The RNA molecule may comprise any combination of any ribonucleotides. The ribonucleotides can be naturally occurring or artificial. One or more ribonucleotides in the RNA molecule can be oxidized or methylated. One or more ribonucleotides in the RNA molecule may be damaged. For instance, the molecules may comprise a pyrimidine dimer, such as a uracil dimer. Such dimers are typically associated with damage by ultraviolet light and are the primary cause of skin melanomas. One or more ribonucleotides in the RNA molecule may be modified, for instance with a label or a tag. Suitable labels include, but are not limited to, fluorescent molecules (such as Cy3 or AlexaFluor®555), radioisotopes, e.g.125I,35S, enzymes, antibodies, antigens, polynucleotides and ligands such as biotin. Suitable tags are discussed below.

[0051] A ribonucleotide typically contains a nucleobase, a ribose sugar and at least one phosphate group. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine, guanine, thymine, uracil and cytosine. The nucleotide typically contains a monophosphate, diphosphate or triphosphate. Phosphates may be attached on the 5' or 3' side of a nucleotide.

[0052] Ribonucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), cytidine monophosphate (CMP), 5-methylcytidine monophosphate, 5-methylcytidine diphosphate, 5-methylcytidine triphosphate, 5-hydroxymethylcytidine monophosphate, 5- hydroxy methylcytidine diphosphate and 5-hydroxymethylcytidine triphosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP and UMP.

[0053] A ribonucleotide may be abasic (i.e. lack a nucleobase). A ribonucleotide may also lack a nucleobase and a sugar (i.e. is a C3 spacer).

[0054] The ribonucleotides in the RNA molecule may be attached to each other in any manner. The ribonucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The ribonucleotides may be connected via their nucleobases as in pyrimidine dimers.

[0055] RNA is an extremely diverse molecule. The RNA molecule may be any naturally occurring or synthetic ribonucleotide molecule, e.g., RNA, messenger RNA (mRNA), Ribosomal RNA (rRNA), Heterogenous nuclear RNA (hnRNA), Transfer RNA (tRNA), Transfer-messenger RNA (tmRNA), Micro RNA (miRNA), Small nuclear RNA (snRNA), Small nucleolar RNA (snoRNA), Signal recognition particle (SRP RNA), SmY RNA, Small Cajal body-specific RNA (scaRNA), Guide RNA (gRNA), Spliced Leader RNA (SL RNA), Antisense RNA (asRNA), Long noncoding RNA (IncRNA), Piwi-interacting RNA (piRNA), Small interfering RNA (siRNA), Trans-acting siRNA (tasiRNA), Repeat associated siRNA (rasiRNA), Y RNA, viral RNA or chromosomal RNA, all of which where appropriate may be single, double or triple stranded.

[0056] The RNA molecule is preferably messenger RNA (mRNA). The mRNA may be an alternate splice variant. Altered amounts (or levels) of mRNA and / or alternate mRNA splice variants may be associated with diseases or conditions.

[0057] Alternatively the RNA molecule is microRNA (or miRNA). One group of RNAs which are difficult to detect in low concentrations are micro-ribonucleic acids (micro-RNA or miRNAs). miRNAs are highly stable RNA oligomers, which can regulate protein production post- transcriptionally. They act by one of two mechanisms. In plants, miRNAs have been shown to act chiefly by directing the cleavage of messenger RNA, whereas in animals, gene regulation by miRNAs typically involves hybridisation of miRNAs to the 3' UTRs of messenger RNAs, which hinders translation (Lee et al., Cell 75, 843-54 (1993); Wightman et al., Cell 75, 855-62 (1993); and Esquela-Kerscher et al., Cancer 6, 259-69 (2006)). miRNAs frequently bind to their targets with imperfect complementarity. They have been predicted to bind to as many as 200 or more gene targets each and to regulate more than a third of all human genes (Lewis et al., Cell 120, 15-20 (2005)).

[0058] Suitable miRNAs are well known in the art. For instance, suitable miRNAs are stored on publically available databases (Jiang Q., Wang Y., Hao Y., Juan L., Teng M., Zhang X., Li M., Wang G., Liu Y., (2009) miR2Disease: a manually curated database for microRNA deregulation in human disease. Nucleic Acids Res.). The expression level of certain microRNAs is known to change in tumours, giving different tumour types characteristic patterns of microRNA expression (Rosenfeld, N. et al., Nature Biotechnology 26, 462-9 (2008)). In addition, miRNA profiles have been shown to be able to reveal the stage of tumour development with greater accuracy than messenger RNA profiles (Lu et al., Nature 435, 834-8 (2005) and Barshack et al., The International Journal of Biochemistry & Cell Biology 42, 1355-62 (2010)). These findings, together with the high stability of miRNAs, and the ability to detect circulating miRNAs in serum and plasma (Wang et al., Biochemical and Biophysical Research Communications 394, 184-8 (2010); Gilad et al., PloS One 3, e3148 (2008); and Keller et al., Nature Methods 8, 841-3 (2011)), have led to a considerable amount of interest in the potential use of microRNAs as cancer biomarkers. For treatment to be effective, cancers need to be classified accurately and treated differently, but the efficacy of tumour morphology evaluation as a means of classification is compromised by the fact that many different types of cancer share morphological features. miRNAs offer a potentially more reliable and less invasive solution. mRNAs and miRNAs may be used to diagnose or prognose diseases or conditions as discussed in more detail below.

[0059] The adaptor may be used to uniquely label a single RNA molecule. Multiple copies of the same adaptor may be used to uniquely label a plurality of RNA molecules. The RNA molecules may be the same. The RNA molecules may be different. Any number of RNA molecules can be uniquely labelled. For instance, the adaptor may be used to uniquely label about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 20, about 30, about 50, about 100, about 500 or about 1000 or more RNA molecules. The adaptor may be used to label all of the RNA molecules in a cell or a sample of cells. The use of a population of different adaptors of the invention with different barcodes to uniquely label a plurality of RNA molecules is discussed in more detail below.

[0060] The RNA molecule can be any length. For example, the RNA molecule can be at least about 10, at least about 50, at least about 100, at least about 150, at least about 200, at least about 250, at least about 300, at least about 400 or at least about 500 ribonucleotides in length. The RNA can be about 1000 or more ribonucleotides, about 5000 or more ribonucleotides in length or about 100000 or more ribonucleotides in length.

[0061] The RNA molecule is typically present in or derived from any suitable sample. The invention is typically carried out on a sample that is known to contain or suspected to contain the RNA molecule.

[0062] The sample may be a biological sample. The invention may be carried out in vitro on a sample obtained from or extracted from any organism or microorganism. The organism or microorganism is typically archaeal, prokaryotic or eukaryotic and typically belongs to one of the five kingdoms: plantae, animalia, fungi, monera and protista.

[0063] The sample is preferably a fluid sample. The sample typically comprises a body fluid of the patient. The sample may be urine, lymph, saliva, mucus or amniotic fluid but is preferably blood, plasma or serum. Typically, the sample is human in origin, but alternatively it may be from another mammal animal such as from commercially farmed animals such as horses, cattle, sheep or pigs or may alternatively be pets such as cats or dogs. Alternatively a sample of plant origin is typically obtained from a commercial crop, such as a cereal, legume, fruit or vegetable, for example wheat, barley, oats, canola, maize, soya, rice, bananas, apples, tomatoes, potatoes, grapes, tobacco, beans, lentils, sugar cane, cocoa or cotton.

[0064] The RNA molecule may be derived from one or more cells. The cells may be any type of cells. The cells may be a prokaryotic cells. The cells may be bacterial or archaeal. The cells are typically eukaryotic cells. The cells may be a protozoan, algal, fungal, plant or animal cells. The animal cells may be derived from the ectoderm, endoderm, or mesoderm. The cells may be stem cells, such as embryonic stem cells, induced pluripotent stem cells or mesenchymal stem cells, bone cells, such as osteoclasts, osteoblasts or osteocytes, tendon cells, such as tenoblasts or tenocytes, chondrocytes, synovial cells, vascular cells, blood cells, such as red blood cells, immune cells, platelet, neutrophils or basophils, muscle cells, such as skeletal muscle cells, cardiac muscle cells or smooth muscle cells, reproductive cells, such as sperm, oocytes, duct cells or epididymal cells, secretory cells, adipocytes, liver lipocytes, epithelial cells, odontoblasts, cementoblasts, hormone-secreting cells, barrier cells, exocrine secretory epithelial cells, nerve cells, astrocytes, oligodendrocytes, or neurons.

[0065] The immune cells may be neutrophil granulocyte and precursors, such as myeloblasts, promyelocytes, myelocytes, or metamyelocytes, eosinophil granulocyte and precursors, basophil granulocyte and precursors, mast cells, leukocytes, lymphocytes, helper T cells, regulatory T cells, cytotoxic T cells, natural killer T cells, B cells, macrophages, dendritic cells, plasma cells, neutrophils, or monocytes.

[0066] The cells may be wild-type or naturally occurring. The cells may be genetically modified or genetically engineered. For instance, the immune cells may be genetically engineered to express a recombinant chimeric antigen receptor (CAR) or T cell receptor (TCR). The cells may be genetically modified or genetically engineered using transduction or transfection or any other common techniques known to those skilled in the art.

[0067] The cells may be healthy cells or obtained from a healthy donor or source. The cells may be diseased or damaged, associated with a disease or damage or obtained from a diseased or damaged donor or source.

[0068] The sample may be a non-biological sample. The non-biological sample is preferably a fluid sample. Examples of a non-biological sample include surgical fluids, water such as drinking water, sea water or river water, and reagents for laboratory tests.

[0069] The sample is typically processed prior to being assayed, for example by centrifugation or by passage through a membrane that filters out unwanted molecules or cells, such as red blood cells. The sample may be measured immediately upon being taken. The sample may also be typically stored prior to assay, preferably below -70°C. The RNA molecule is typically extracted from the sample before it is used in the method of the invention. RNA extraction kits are commercially available from, for instance, New England Biolabs® and Invitrogen®.

[0070] Double stranded polynucleotide

[0071] The adaptor of the invention comprises or consists of a double stranded polynucleotide. The adaptor comprises a first strand and a second strand. The first strand is a first RNA strand. The second strand a second polynucleotide strand. The first strand and / or the second strand may be any length. The first strand and / or the second strand may be from about 10 to about 150 nucleotides in length, such as from about 20 to about 140 nucleotides in length, from about 30 to about 120 nucleotides in length, from about 40 to about 100 nucleotides in length, or from about 50 to about 80 nucleotides in length. The first strand and / or the second strand may be at least about 10 nucleotides in length, such at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70 or at least about 80 nucleotides in length.

[0072] At least a part of the first strand is hybridised to at least a part of the second strand. A skilled person is capable of hybridising two polynucleotide strands. Any part of the first strand may hybridise to at least a part of the second strand. At least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 72%, at least about 80%, at least about 90% or at least about 95% of the first strand may hybridise to at least a part of the second strand. At least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 72%, at least about 80%, at least about 90% or at least about 95% of the second strand may hybridise to at least a part of the first strand. These %s are typically calculated based on numbers of nucleotides. The same number of nucleotides in the first and second strands typically hybridise. From about 20 to about 100 nucleotides, such as from about 30 to about 80 nucleotides or from about 40 to about 60 nucleotides, in the first and second strands may hybridise. At least about 20, at least about 30, at least about 40, at least about 50 or at least about 60 nucleotides in the first and second strands may hybridise.

[0073] At least part of the first strand preferably specifically hybridises to at least part of the second strand. Strands or parts thereof "specifically hybridise" when they hybridise with preferential or high affinity each other but do not substantially hybridise, does not hybridise, or hybridises with only low affinity to other polynucleotide sequences, especially other RIMA molecule(s), adaptors or sequences used in the invention. Conditions that permit the hybridisation are well-known in the art (for example, Sambrook et al., 2001, Molecular Cloning: a laboratory manual, 3rd edition, Cold Spring Harbour Laboratory Press; and Current Protocols in Molecular Biology, Chapter 2, Ausubel et al., Eds., Greene Publishing and Wiley-lnterscience, New York (1995)). Hybridisation can be carried out under low stringency conditions, for example in the presence of a buffered solution of 30 to 35% formamide, 1 M NaCI and 1 % SDS (sodium dodecyl sulfate) at 37 °C followed by a 20 wash in from IX (0.1650 M Na+) to 2X (0.33 M Na+) SSC (standard sodium citrate) at 50 °C. Hybridisation can be carried out under moderate stringency conditions, for example in the presence of a buffer solution of 40 to 45% formamide, 1 M NaCI, and 1 % SDS at 37 °C, followed by a wash in from 0.5X (0.0825 M Na+) to IX (0.1650 M Na+) SSC at 55 °C. Hybridisation can be carried out under high stringency conditions, for example in the presence of a buffered solution of 50% formamide, 1 M NaCI, 1% SDS at 37 °C, followed by a wash in 0.1X (0.0165 M Na+) SSC at 60 °C. The strands or parts thereof "specifically hybridise" if they hybridise with a melting temperature (Tm) that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C or at least 10 °C, greater than its Tm for other polynucleotide sequences. More preferably, the strands or parts thereof hybridise with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for other polynucleotide sequences. Preferably, the strands or parts thereof hybridise with a Tm that is at least 2 °C, such as at least 3 °C, at least 4 °C, at least 5 °C, at least 6 °C, at least 7 °C, at least 8 °C, at least 9 °C, at least 10 °C, at least 20 °C, at least 30 °C or at least 40 °C, greater than its Tm for a polynucleotide which differs from the corresponding strand or part thereof by one or more nucleotides, such as by 1, 2, 3, 4 or 5 or more nucleotides. The strands or parts thereof typically hybridise with a Tm of at least 90 °C, such as at least 92 °C or at least 95 °C. Tm can be measured experimentally using known techniques, including the use of DNA microarrays, or can be calculated using publicly available Tm calculators, such as those available over the internet.

[0074] Each strand or part thereof typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of the other strand or part thereof. Each strand or part thereof preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the other strand or part. Each strand or part thereof most preferably comprises or consists of a sequence which is the reverse complement the other strand or part. By complementary, it is meant each strand or part thereof comprises or consists of a sequence 100% identical or homologous to the other strand or part. Complementarity is typically determined using canonical Watson-Crick base pairing.

[0075] Standard methods in the art may be used to determine identity and / or homology. Suitable methods are described above.

[0076] First strand

[0077] The first strand is a first RIMA strand. RNA is defined above. The first RNA strand may be any of the RNAs defined above. The first RNA strand may comprise or consist of any of the ribonucleotides discussed above.

[0078] The first RNA strand comprises the barcode. The barcode is a RNA barcode. Polynucleotide barcodes are well-known in the art (Kozarewa, I. et al., (2011), Methods Mol. Biol. 733, p279-298). A barcode is a specific sequence of nucleotide, in this instance ribonucleotides, that can be characterised and identified. A barcode is preferably a specific sequence of nucleotides, in this instance ribonucleotides, that affects the current flowing through the pore in a specific and known manner. Method for generating barcodes are known in the art (e.g., Johnson, M.S., Venkataram, S. & Kryazhimskiy, S. Best Practices in Designing, Sequencing, and Identifying Random DNA Barcodes. J Mol Evol 91, 263-280 (2023). https: / / doi.org / 10.1007 / s00239-022-10083-z).

[0079] The barcode typically comprises a specific sequence of ribonucleotides. The ribonucleotides may be any of those defined above. One or more abasic nucleotides may be included in the barcode. Using one or more abasic nucleotides results in characteristic current spikes. This allows the clear highlighting of the positions of the one or more nucleotide species in the barcode.

[0080] The ribonucleotide species in the barcode may be modified to comprise a chemical atom or group such as a propynyl group, a thiol group, an oxo group, a methyl group, a hydroxymethyl group, a formyl group, a carboxy group, a carbonyl group, a benzyl group, a propargyl group or a propargylamine group. The chemical group or atom may be or may comprise a fluorescent molecule, biotin, digoxigenin, DNP (dinitrophenol), a photo-labile group, an alkyne, DBCO, azide, free amino group, a redox dye, a mercury atom, or a selenium atom.

[0081] The RNA barcode may comprise a nucleotide species comprising a halogen atom. The halogen atom may be attached to any position on the different nucleotide species, such as the nucleobase and / or the sugar. The halogen atom is preferably fluorine (F), chlorine (Cl), bromine (Br) or iodine (I). The halogen atom is most preferably F or I.

[0082] The barcode may be any length. The barcode is preferably at least about 5 nucleotides, at least about 6 nucleotides, at least about 7 nucleotides, at least about 8 nucleotides, at least about 9 nucleotides, at least about 10 nucleotides, at least about 11 nucleotides, at least about 12 nucleotides, at least about 13 nucleotides, at least about 14 nucleotides, at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 35 nucleotides, at least about 40 nucleotides, at least about 45 nucleotides, or at least about 50 nucleotides in length. The longer the barcode the greater the number of different barcode combinations that can be defined and created.

[0083] The barcode preferably comprises one or more of (a) up to about 40 nucleotides in length, such as about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 21, about

[0084] 22, about 23, about 24, about 25, about 26, about 27, about 28, about 29, about 30, about

[0085] 31, about 32, about 33, about 34, about 35, about 36, about 37, about 38, about 39 or about 40 nucleotides in length, (b) no homopolymers longer than about 3 nucleotides, and (c) no secondary structure. The barcode may comprise (a), (b), (c), (a) and (b), (a) and (c), (b) and (c) or (a), (b) and (c).

[0086] The barcode preferably comprises a sequence selected from the sequences underlined in Table 1.

[0087] Second strand

[0088] The second strand is a second polynucleotide strand. A polynucleotide, such as a nucleic acid, is a macromolecule comprising two or more nucleotides. In the context of the second strand, the polynucleotide is typically single-stranded.

[0089] A polynucleotide may comprise any combination of any nucleotides. The nucleotides can be naturally occurring or artificial. A nucleotide typically contains a nucleobase, a sugar and at least one phosphate group. The nucleobase and sugar form a nucleoside. The nucleobase is typically heterocyclic. Nucleobases include, but are not limited to, purines and pyrimidines and more specifically adenine (A), guanine (G), thymine (T), uracil (U) and cytosine (C).

[0090] The sugar is typically a pentose sugar. Nucleotide sugars include, but are not limited to, ribose and deoxyribose. The sugar is preferably a deoxyribose. The polynucleotide preferably comprises the following nucleosides: deoxyadenosine (dA), deoxyuridine (dU) and / or thymidine (dT), deoxyguanosine (dG) and deoxycytidine (dC).

[0091] The nucleotide is typically a ribonucleotide or deoxyribonucleotide. The nucleotide typically contains a monophosphate, diphosphate, or triphosphate. The nucleotide may comprise more than three phosphates, such as 4 or 5 phosphates. Phosphates may be attached on the 5' or 3' side of a nucleotide. Nucleotides include, but are not limited to, adenosine monophosphate (AMP), guanosine monophosphate (GMP), thymidine monophosphate (TMP), uridine monophosphate (UMP), 5-methylcytidine monophosphate, 5- hydroxy methylcytidine monophosphate, cytidine monophosphate (CMP), cyclic adenosine monophosphate (cAMP), cyclic guanosine monophosphate (cGMP), deoxyadenosine monophosphate (dAMP), deoxyguanosine monophosphate (dGMP), deoxythymidine monophosphate (dTMP), deoxyuridine monophosphate (dUMP), deoxycytidine monophosphate (dCMP) and deoxymethylcytidine monophosphate. The nucleotides are preferably selected from AMP, TMP, GMP, CMP, UMP, dAMP, dTMP, dGMP, dCMP and dUMP.

[0092] A nucleotide may be abasic ( / .e., lack a nucleobase). A nucleotide may also lack a nucleobase and a sugar ( / .e., is a C3 spacer).

[0093] The nucleotides in the polynucleotide may be attached to each other in any manner. The nucleotides are typically attached by their sugar and phosphate groups as in nucleic acids. The nucleotides may be connected via their nucleobases as in pyrimidine dimers. The polynucleotide can be a nucleic acid, such as deoxyribonucleic acid (DNA) or ribonucleic acid (RIMA). The polynucleotide can comprise one strand of RNA hybridized to one strand of DNA. The polynucleotide may be any synthetic nucleic acid known in the art, such as peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA), bridged nucleic acid (BNA) or other synthetic polymers with nucleotide side chains. The PNA backbone is composed of repeating N-(2-aminoethyl)-glycine units linked by peptide bonds. The GNA backbone is composed of repeating glycol units linked by phosphodiester bonds. The TNA backbone is composed of repeating threose sugars linked together by phosphodiester bonds. LNA is formed from ribonucleotides as discussed above having an extra bridge connecting the 2' oxygen and 4' carbon in the ribose moiety.

[0094] The polynucleotide is preferably DNA or a DNA or RNA hybrid, most preferably DNA. A DNA / RNA hybrid may comprise DNA and RNA on the same strand.

[0095] The backbone of the polynucleotide can be altered to reduce the possibility of strand scission. For example, DNA is known to be more stable than RNA under many conditions. The backbone of the polynucleotide strand can be modified to avoid damage caused by e.g., harsh chemicals such as free radicals. DNA or RNA that contains unnatural or modified bases can be produced by amplifying natural DNA or RNA polynucleotides in the presence of modified NTPs using an appropriate polymerase.

[0096] The second strand or second polynucleotide comprises or consists of a non-RNA polynucleotide. The non-RNA polynucleotide must comprise at least one nucleotide which is not a ribonucleotide, i.e., which is not from RNA. Ribonucleotides are described above. The non-RNA polynucleotide may additionally comprise a ribonucleotide or RNA but it must also comprise or include at least one non-RNA nucleotide, i.e., a nucleotide that is not RNA. Typically, the non-RNA polynucleotide which comprises RNA comprises less than about 20 ribonucleotides, such as about 1, about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, or about 19 ribonucleotides. The non-RNA polynucleotide may therefore be a "hybrid" polynucleotide comprising, for example, RNA and another polynucleotide, such as DNA or a DNA analogue. The non-RNA may also include DNA spacers etc. The skilled person will be aware that any of the attachment methods described as suitable for making a modified RNA construct of the invention are equally suitable for making the non-RNA polynucleotide, wherein two or more types of nucleic acid sequence may be combined.

[0097] The non-RNA polynucleotide may comprise or consist of the overhang. The non-RNA polynucleotide may be a separate part of the second strand or second polynucleotide strand from the overhang. The second strand or second polynucleotide strand is preferably a second DNA strand. The second DNA strand may comprise deoxynucleotides selected from dAMP, dTMP, dGMP, dCMP and dUMP. The second DNA strand may be modified using any of the methods described herein.

[0098] The second strand, second polynucleotide strand or second DNA strand comprises an overhang. The overhang can be any length. For example, the overhang can be at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 15, at least about 20, or at least 25 nucleotides in length. The overhang is preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. The overhang may comprise any of the nucleotide discussed above, including ribonucleotides. Preferred overhangs are discussed in more detail below.

[0099] Adaptor synthesis

[0100] The adaptors of the invention are typically synthetic or semi-synthetic. For example, DNA or RNA may be purely synthetic, synthesised by conventional DNA synthesis methods such as phosphoramidite based chemistries. Synthetic polynucleotides subunits may be joined together by known means, such as ligation or chemical linkage, to produce longer strands. Internal self-forming structures (e.g., hairpins, quadruplexes) can be designed into the substrate, e.g., by ligating appropriate sequences. Synthetic polynucleotides can be copied and scaled up for production by means known in the art, including PCR, incorporation into bacterial factories, and the like.

[0101] First embodiment

[0102] In a first embodiment of the adaptor of the invention, the overhang is preferably capable of hybridising or specifically hybridising to the 3' end of the RNA molecule. The skilled person is capable of designing an overhang capable of hybridising or specifically hybridising to the 3' end of the RNA molecule. Specific hybridisation is defined above.

[0103] The overhang may have any of the features discussed above, including the various lengths.

[0104] The overhang typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of the 3' end of the RNA molecule. The overhang preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the 3' end of the RNA molecule. The overhang most preferably comprises or consists of a sequence which is the reverse complement of the 3' end of the RNA molecule. The overhang preferably comprises or consists of a sequence 100% identical or homologous to the reverse complement of the 3' end of the RIMA molecule. Complementarity is typically determined using canonical Watson-Crick base pairing.

[0105] Eukaryotic RNA molecules typically have poly(A) tails. The overhang is preferably capable of hybridising or specifically hybridising to the poly(A) tail of the RNA molecule. The overhang preferably comprises or consists of one or more of (i) thymine-containing nucleotides, (ii) uracil-containing nucleotides and (iii) universal nucleotides. The overhand may comprise (i), (ii), (iii), (i) and (ii), (i) and (iii), (ii) and (iii) or (i), (ii) and (iii).

[0106] Overhangs comprising (i) and (ii) are capable of specifically hybridising to the poly(A) tails or polyadenylated 3' ends as described above. Nucleotides are defined above. The thymine- containing nucleotides may be TMP or dTMP. The uracil-containing nucleotides may be UMP or dUMP.

[0107] A universal nucleotide is one which will hybridise or bind to some degree to all of the nucleotides in a polynucleotide. A universal nucleotide is preferably one which will hybridise or bind to some degree to nucleotides comprising the nucleosides adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). The universal nucleotide may hybridise or bind more strongly to some nucleotides than to others. For instance, a universal nucleotide (I) comprising the nucleoside, 2'-deoxyinosine, will show a preferential order of pairing of I- C>I-A>I-G approximately=I-T.

[0108] The universal nucleotide preferably comprises one of the following nucleobases: hypoxanthine, 4-nitroindole, 5-nitroindole, 6-nitroindole, formylindole, 3-nitropyrrole, nitroimidazole, 4-nitropyrazole, 4-nitrobenzimidazole, 5-nitroindazole, 4- aminobenzimidazole or phenyl (C6-aromatic ring). The universal nucleotide more preferably comprises one of the following nucleosides: 2'-deoxyinosine, inosine, 7-deaza-2'- deoxyinosine, 7-deaza-inosine, 2-aza-deoxyinosine, 2-aza-inosine, 2-O'-methylinosine, 4- nitroindole 2'-deoxyribonucleoside, 4-nitroindole ribonucleoside, 5-nitroindole 2'- deoxyribonucleoside, 5-nitroindole ribonucleoside, 6-nitroindole 2'-deoxyribonucleoside, 6- nitroindole ribonucleoside, 3-nitropyrrole 2'-deoxyribonucleoside, 3-nitropyrrole ribonucleoside, an acyclic sugar analogue of hypoxanthine, nitroimidazole 2'- deoxyribonucleoside, nitroimidazole ribonucleoside, 4-nitropyrazole 2'-deoxyribonucleoside, 4-nitropyrazole ribonucleoside, 4-nitrobenzimidazole 2'-deoxyribonucleoside, 4- nitrobenzimidazole ribonucleoside, 5-nitroindazole 2'-deoxyribonucleoside, 5-nitroindazole ribonucleoside, 4-aminobenzimidazole 2'-deoxyribonucleoside, 4-aminobenzimidazole ribonucleoside, phenyl C-ribonucleoside, phenyl C-2'-deoxyribosyl nucleoside, 2'- deoxynebularine, 2'-deoxyisoguanosine, K-2'-deoxyribose, P-2'-deoxyribose and pyrrolidine. The universal nucleotide more preferably comprises 2'-deoxyinosine. The universal nucleotide is more preferably IMP or dIMP. The universal nucleotide is most preferably dPMP (2'-Deoxy-P-nucleoside monophosphate) or dKMP (N6-methoxy-2, 6- diaminopurine monophosphate).

[0109] The overhang preferably comprises or consists of (in the 5' to 3' direction) one or more consecutive repeating units of UT or UTT, such as 2 or more, 3 or more, 4 or more or 5 more consecutive repeating units of UT or UTT. The overhang preferably comprises or consists of five or more, such as five, consecutive repeating units of UT. In these embodiments, U is preferably dUMP and T is preferably dTMP. These overhangs minimise cross-contamination between barcodes. The U nucleotides can be digested with USER (a combination of a glycosylase and an endonuclease, NEB) after reverse transcription (RT) to minimise the chance of the wrong barcode attaching to an "unlabelled with barcode" sample.

[0110] Prokaryotic RNA molecules do not have a poly(A) tail. The Poly A Polymerase enzyme can be utilised to add a poly(dA) tail onto the 3' end of a prokaryotic RNA molecule. Any of the overhangs discussed above may then be used.

[0111] The overhang comprising one or more of (i), (ii) or (iii) may comprise any number of nucleotides, such at least about 5, at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 15, at least about 20, or at least 25 nucleotides. The overhang preferably comprises less than about 50 nucleotides, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides.

[0112] In the first embodiment, the second strand is a second DNA strand. The overhang is preferably at the 3' end of the second DNA strand.

[0113] The 5' end of the second DNA strand preferably does not hybridise to the first RNA strand. This results in the adaptors of the first embodiment being Y adaptors. Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strands form overhangs. The presence of a non-complementary region in the Y adaptors gives them their Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The hybridised stem end of the Y adaptors typically comprises the overhang that allows them to specifically hybridise to and be ligated to the RNA molecule.

[0114] The first RNA strand preferably comprises a sequence capable of hybridising to a sequencing adaptor. In the context of the invention, this sequence may be known as a first sequence to contrast it with a second, corresponding sequence in the sequencing adaptor. The first sequence may from about 5 to about 30 nucleotides, such as from about 10 to about 25 nucleotides or from about 15 to about 20 nucleotides in length. The first sequence typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first sequence preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first sequence most preferably comprises or consists of a sequence which is the reverse complement of a second sequence in the sequencing adaptor. The first sequence preferably comprises or consists of a sequence 100% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first and second sequences are typically the same length. Complementarity is typically determined using canonical Watson-Crick base pairing. The sequence capable of hybridising to a sequencing adaptor is preferably at the 3' end of the first RNA strand. The sequencing adaptor may be any of the adaptors known in the art, including those described in WO 2021 / 255476 (incorporated herein by reference in its entirety). The sequence adaptor is preferably a sequencing adaptor of the invention.

[0115] The first RNA strand may comprise or consist of the sequence shown in any one of SEQ ID NOs: 1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31, 33, 35, 37, 39, 41, 43, 45,

[0116] 47, 49, 51, 53, 55, 57, 59, 61, 63, 65, 67, 69, and 71.

[0117] The second DNA strand may comprise or consist of the sequence shown in any one of SEQ ID NOs: 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30, 32, 34, 36, 38, 40, 42, 44,

[0118] 46, 48, 50, 52, 54, 56, 58, 60, 62, 64, 66, 68, 70 and 72

[0119] The double stranded adaptor preferably comprises:

[0120] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 2;

[0121] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 3 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 4;

[0122] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 5 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 6;

[0123] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 7 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 8; • a first RIMA strand comprising or consisting of the sequence shown in SEQ ID NO: 9 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 10;

[0124] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 11 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 12;

[0125] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 13 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 14;

[0126] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 15 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 16;

[0127] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 17 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 18;

[0128] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 19 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 20;

[0129] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 21 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 22;

[0130] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 23 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 24;

[0131] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 25 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 26;

[0132] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 27 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 28;

[0133] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 29 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID • a first RIMA strand comprising or consisting of the sequence shown in SEQ ID NO: 31 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 32;

[0134] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 33 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 34;

[0135] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 35 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 36;

[0136] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 37 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 38;

[0137] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 39 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 40;

[0138] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 41 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 42;

[0139] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 43 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 44;

[0140] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 45 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 46; or

[0141] • a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 47 and a second DNA strand comprising or consisting of the sequence shown in SEQ ID NO: 48.

[0142] Second embodiment

[0143] In a second embodiment of the adaptor of the invention, the overhang is preferably capable of hybridising or specifically hybridising to the tail on a transfer RNA (tRNA) molecule. The skilled person is capable of designing an overhang capable of hybridising or specifically hybridising to the tail on tRNA molecule. Specific hybridisation is defined above. The overhang may have any of the features discussed above, including the various lengths. The overhang typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of the tail of the tRNA molecule. The overhang preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the tail of the tRNA molecule. The overhang most preferably comprises or consists of a sequence which is the reverse complement of the tail of the tRNA molecule. The overhang preferably comprises or consists of a sequence 100% identical or homologous to the reverse complement of the tail of the tRNA molecule. Complementarity is typically determined using canonical Watson-Crick base pairing.

[0144] The overhang preferably comprises 5' to 3' the sequence UGG or TGG. This sequence is capable of specifically hybridising to the CCA (5' to 3') tail of the tRNA molecule.

[0145] In the second embodiment, the second strand preferably comprises a second RNA strand and a reverse transcription (RT) primer. The second RNA strand and the reverse transcription (RT) primer may be attached together. The second RNA strand and the reverse transcription (RT) primer are preferably separate polynucleotides and are hybridised to the first RNA strand. The RT primer represents the non-RNA polynucleotide. Reverse transcriptase requires a DNA sequence as a primer site. The overhang is preferably at the 3' end of the second RNA strand. The overhang is preferably at the 3' end of the second RNA strand and comprises 5' to 3' the sequence UGG.

[0146] The 5' end of the second RNA strand preferably does not hybridise to the first RNA strand. This results in the adaptors of the second embodiment being or comprising Y adaptors. Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non-complementary parts of the strands form overhangs. The hybridised stem end of the Y adaptor comprising the second RNA strand typically comprises the overhang that allows the adaptor to specifically hybridise to and be ligated to the tRNA molecule.

[0147] The RT primer is typically preferably a single stranded polynucleotide RT primer. The RT primer may be any length. The RT primer is preferably at least about 15 nucleotides, at least about 20 nucleotides, at least about 25 nucleotides, at least about 30 nucleotides, at least about 32 nucleotides, at least about 35 nucleotides, or at least about 40 nucleotides in length. The RT primer is preferably less than about 50 nucleotides in length, such as less than about 49 nucleotides, less than about 45 nucleotides, less than about 42 nucleotides, or less than about 40 nucleotides in length. The nucleotides in the RT primer are preferably DNA nucleotides. They may be selected from dAMP, dTMP, dGMP and dCMP. The skilled person is capable of designing sequences which function as a RT primer sequence. The 5' end of the RT primer preferably does not hybridise to the first RNA strand. This embodiment results from both the 5' end of the second RNA strand and the 5' end of the RT primer not hybridising to the first RNA strand. An Example of this is shown in Figure 7.

[0148] The first RNA strand preferably comprises a sequence capable of hybridising to a sequencing adaptor. In the context of the invention, this sequence may be known as a first sequence to contrast it with a second, corresponding sequence in the sequencing adaptor. The first sequence may from about 5 to about 30 nucleotides, such as from about 10 to about 25 nucleotides or from about 15 to about 20 nucleotides in length. The first sequence typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first sequence preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first sequence most preferably comprises or consists of a sequence which is the reverse complement of a second sequence in the sequencing adaptor. The first sequence preferably comprises or consists of a sequence 100% identical or homologous to the reverse complement of a second sequence in the sequencing adaptor. The first and second sequences are typically the same length. Complementarity is typically determined using canonical Watson-Crick base pairing. The sequence capable of hybridising to a sequencing adaptor is preferably at the 3' end of the first RNA strand. The sequencing adaptor may be any of the adaptors known in the art, including those described in WO 2021 / 255476 (incorporated herein by reference in its entirety). The sequence adaptor is preferably a sequencing adaptor of the invention.

[0149] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 49 and the sequence shown in SEQ ID NO: 50. In this embodiment, SEQ ID NO: 49 is the second RNA strand and SEQ ID NO: 50 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0150] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 55 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 55 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0151] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 56 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 56 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0152] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 57 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 57 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0153] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 58 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 58 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0154] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 59 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 59 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0155] The double stranded adaptor preferably comprises a first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 49 and the sequence shown in SEQ ID NO: 60. In this embodiment, SEQ ID NO: 49 is the second RNA strand and SEQ ID NO: 60 is the RT primer. They are separate polynucleotides hybridised to the first RNA strand comprising or consisting of the sequence shown in SEQ ID NO: 1.

[0156] Population of adaptors of the invention

[0157] The invention also provides a population of double stranded polynucleotide adaptors of the invention. The population is for uniquely labelling a plurality of RNA molecules. The RNA molecules being uniquely labelled may also be known as the target RNA molecules. In the context of the invention, a plurality of RNA molecules are uniquely labelled when each RNA molecule is labelled with a different barcode.

[0158] The double stranded polynucleotide adaptors in the population may be any of the adaptors of invention defined above. The double stranded polynucleotide adaptors may be the adaptors of the first embodiment discussed above. The double stranded polynucleotide adaptors may be the adaptors of the second embodiment discussed above.

[0159] Each double stranded polynucleotide adaptor comprises a different barcode. The different barcodes are different RIMA barcodes. Such barcodes are defined above. The skilled person is capable of designing a suitable population of barcodes for use in the invention. The different barcodes may differ on the basis of their length and / or sequence. The different barcodes preferably comprise sequences selected from the sequences underlined in Table 1.

[0160] The plurality of RNA molecules being uniquely labelled may also be known as the plurality of target RNA molecules. The RNA molecules may have any of the lengths and may be any of the types of RNA molecules defined above. The RNA molecules may be the same. The RNA molecules are typically different in terms of one or more of their (a) length, (b), type, (c) sequence and (d) source. The RNA molecules are typically different in terms of (a); (b); (c); (d); (a) and (b); (a) and (c); (a) and (d); (b) and (c); (b) and (d); (c) and (d); (a), (b) and (c); (a), (b) and (d); (a), (c) and (d); (b), (c) and (d); or (a), (b), (c) and (d).

[0161] Any number of RNA molecules can be uniquely labelled. For instance, the population may be used to uniquely labelling about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 20, about 30, about 50, about 100, about 500 or about 1000 or more RNA molecules. The adaptor may be used to label all of the RNA molecules in a cell or a sample of cells.

[0162] The population may comprise any number of adaptors of the invention. The population comprise about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 20, about 30, about 50, about 100 or more adaptors of the invention.

[0163] The number of adaptors in the population are typically the same as the number of RNA molecules in the plurality.

[0164] The overhangs in the adaptors may be the same or different. If different, the population of adaptors preferably comprise overhangs comprising every possible combination of sequences based on nucleotides comprising adenosine (A), thymine (T), uracil (U), guanine (G) and cytosine (C). Such overhangs are capable of hybridising or specifically hybridising to the 3' ends of the RNA molecules. Adaptors comprising overhangs can be produced randomly to include the combinations of sequences.

[0165] All of the adaptors in the population preferably comprise the same overhang. All of the adaptors in the population may comprise the same second strand. All of the adaptors in the population may comprise the same first RNA strand apart from the different barcodes. All of the adaptors in the population may comprise the same RT primer. All of the adaptors in the population may be identical apart from the different barcodes. This allows the use of a standard set of adaptors in which only the barcodes differ between adaptors.

[0166] Methods of making the double stranded polynucleotide adaptors

[0167] The invention also provides a method of making a double stranded polynucleotide adaptor of the invention, the method comprising hybridising the first RIMA strand to the second strand. The invention also provides a method of making a population of double stranded polynucleotide adaptors of the invention, the method comprising hybridising the first RNA strands to the second strands. The double stranded polynucleotide adaptor(s) may be any of those defined above. The first RNA strand(s) and second strands may be any of those defined above.

[0168] Polynucleotide sequencing adaptor

[0169] The invention also provides a polynucleotide sequencing adaptor comprising a RNA polynucleotide upstream 5' to 3' of a loading site for a polynucleotide binding protein. Sequencing adaptors are known in the art, including in WO 2021 / 255476 (incorporated herein by reference in its entirety). The RNA polynucleotide may be any of those defined above for the first RNA strand. The sequencing adaptor of the invention provides more consistent movement of a RNA molecule and / or barcode with respect to a detector, such as a nanopore.

[0170] A sequencing adaptor typically comprises a polynucleotide strand capable of being attached to the end of a target polynucleotide. The target polynucleotide is typically intended for characterisation in accordance with methods disclosed herein, and includes the double stranded polynucleotide adaptors.

[0171] A sequencing adaptor may be added to both ends of the target polynucleotide. Alternatively, different adaptors may be added to the two ends of the target polynucleotide. An adaptor may be added to just one end of the target polynucleotide. Methods of adding adaptors to polynucleotides are known in the art. Adaptors may be attached to polynucleotides, for example, by ligation, by click chemistry, by tagmentation, by topoisomerisation or by any other suitable method.

[0172] The sequencing adaptor may be synthetic or artificial. Typically, an adaptor comprises a polynucleotide. The polynucleotide may be any of those defined herein. An adaptor may comprise a single-stranded polynucleotide strand. An adaptor may comprise a doublestranded polynucleotide. A sequencing adaptor may comprise any of the polynucleotide discussed above and includes DNA, RNA, modified DNA (such as a basic DNA), RNA, PNA, LNA, BNA and / or PEG. Usually, the adaptor comprises single stranded and / or double stranded DNA or RNA. The sequencing adaptors may be Y adaptors. Y adaptors are typically double stranded and comprise (a) at one end, a region where the two strands are hybridised together and (b), at the other end, a region where the two strands are not complementary. The non- complementary parts of the strands form overhangs. The hybridised stem of the adaptors typically attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 3' end of a second strand of a double-stranded polynucleotide; or to the 3' end of a first strand of a double-stranded polynucleotide and the 5' end of a second strand of a doublestranded polynucleotide. The presence of a non-complementary region in the Y adaptors gives them their Y shape since the two strands typically do not hybridise to each other unlike the double stranded portion. The hybridised stem end of the Y adaptors may also comprise a short overhang that allows them to specifically hybridise to and be attached to the first adaptors, RT primers, second adaptors and barcoded constructs.

[0173] The adaptor preferably comprises a double stranded region and a region where the two strands are not complementary. The adaptor preferably comprises a first strand comprising the RNA polynucleotide and a second strand. The first strand and / or the second strand may be any length. The first strand and / or the second strand may be from about 10 to about 150 nucleotides in length, such as from about 20 to about 140 nucleotides in length, from about 30 to about 120 nucleotides in length, from about 40 to about 100 nucleotides in length, or from about 50 to about 80 nucleotides in length. The first strand and / or the second strand may be at least about 10 nucleotides in length, such at least about 20, at least about 30, at least about 40, at least about 50, at least about 60, at least about 70 or at least about 80 nucleotides in length. The first strand may comprise two or more, such as 2, 3, 4, 5 or 6, polynucleotide sections separated by one or more, such as 1, 2, 3, 4 or 5, spacers. The one or more spacers may be any of those defined below. One of the two or more polynucleotide sections comprises or consists of the RNA polynucleotide.

[0174] At least a part of the first strand is hybridised to at least a part of the second strand. A skilled person is capable of hybridising two polynucleotide strands. Any part of the first strand may hybridise to at least a part of the second strand. At least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 72%, at least about 80%, at least about 90% or at least about 95% of the first strand may hybridise to at least a part of the second strand. At least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 72%, at least about 80%, at least about 90% or at least about 95% of the second strand may hybridise to at least a part of the first strand. These %s are typically calculated based on numbers of nucleotides. The same number of nucleotides in the first and second strands typically hybridise. From about 5 to about 80 nucleotides, such as from about 10 to about 60 nucleotides or from about 13 to about 40 nucleotides, in the first and second strands may hybridise. At least about 5, at least about 10, at least about 12, at least about 13 or at least about 20 nucleotides in the first and second strands may hybridise.

[0175] At least part of the first strand preferably specifically hybridises to at least part of the second strand. Specific hybridise is defined above.

[0176] In any of these embodiments, the at least part of the first strand preferably comprises the barcode. The barcode is preferably hybridised to or specifically hybridised to the second strand.

[0177] Each strand or part thereof typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of the other strand or part thereof. Each strand or part thereof preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of the other strand or part. Each strand or part thereof most preferably comprises or consists of a sequence which is the reverse complement the other strand or part. By complementary, it is meant each strand or part thereof comprises or consists of a sequence 100% identical or homologous to the other strand or part. Complementarity is typically determined using canonical Watson-Crick base pairing.

[0178] Standard methods in the art may be used to determine identity or homology as described above.

[0179] The loading site is preferably present in one of the two strands that are not complementary. The RIMA polynucleotide is preferably part of the strand comprising the loading site and forms the double stranded region.

[0180] Those skilled in the art will also appreciate that when the sequencing adaptors comprise polynucleotide strand(s), the sequences of the adaptors are typically not determinative and can be controlled or chosen according to the polynucleotide binding protein and other experimental conditions such as any polynucleotides to be characterised. Exemplary sequences are provided solely by way of illustration in the Examples. For example, the adaptors may comprise a sequence such as one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021 / 255476 (incorporated herein by reference in its entirety) or polynucleotide sequences having at least 20%, such as at least 30%, e.g., at least 40% such as at least 50%, e.g., at least 60% such as at least 70%, e.g., at least 80%, for example at least 90% e.g., at least 95% sequence similarity or identity to one or more of SEQ ID NOs: 21-26 or 28-33 in WO 2021 / 255476 (incorporated herein by reference in its entirety). The sequences of the adaptors can typically be altered without negatively affecting the efficacy of the method of the invention. The sequencing adaptor of the invention comprises a loading site for a polynucleotide binding protein. The polynucleotide binding protein may be any of those defined below. The loading site may be present on an overhang of an adaptor such as a Y adaptor. The loading site may be in the double stranded region. The loading site may form part of a singlestranded and / or a double-stranded region of the adaptor. Suitable loading sites are known in the art. The loading site preferably comprises from about 5 to about 20 consecutive thymine-containing nucleotides, such as about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19 or about 20 consecutive thymine-containing nucleotides.

[0181] WO 2015 / 110813 and WO 2020 / 234612 describe the loading of polynucleotide binding proteins onto a target polynucleotide such as an adaptor and are hereby incorporated by reference in their entireties.

[0182] The sequencing adaptor preferably further comprises a polynucleotide binding protein loaded on the loading site. The polynucleotide binding protein may be any of those defined below.

[0183] The sequencing adaptor preferably further comprises a sequence capable of hybridising to a double stranded polynucleotide adaptor of the invention. In the context of the invention, this sequence may be known as a second sequence to contrast it with a first, corresponding sequence in the double stranded polynucleotide adaptor. The second sequence may be any length. The second sequence may from about 5 to about 30 nucleotides, such as from about 10 to about 25 nucleotides or from about 13in 15 to about 20 nucleotides in length. The second sequence typically comprises or consists of a sequence at least about 80% identical or homologous to the reverse complement of a first sequence in the double stranded polynucleotide adaptor. The second sequence preferably comprises or consists of a sequence at least about 85%, at least about 90%, at least about 95%, at least about 98% or at least about 99% identical or homologous to the reverse complement of a first sequence in the double stranded polynucleotide adaptor. The second sequence most preferably comprises or consists of a sequence which is the reverse complement of a first sequence in the double stranded polynucleotide adaptor. The second sequence preferably comprises or consists of a sequence 100% identical or homologous to the reverse complement of a first sequence in the double stranded polynucleotide adaptor. The first and second sequences are typically the same length. Complementarity is typically determined using canonical Watson-Crick base pairing. The second sequence capable of hybridising to a sequencing adaptor is preferably at the 3' end of the second strand in the sequencing adaptor. Sequencing adaptors are discussed in more detail below.

[0184] The sequencing adaptor preferably comprises a membrane anchor or a pore anchor. The anchor may be attached to a polynucleotide which hybridises to, specifically hybridises to or is complementary to one of the non-complementary regions / overhangs. The non- complementary region / overhang preferably comprises the loading site for the polynucleotide binding protein.

[0185] The sequencing adaptor may comprise a leader sequence. One of the non-complementary regions / overhangs may comprise a leader sequence. Leader sequences are capable of threading into are nanopore. The leader sequence typically comprises a polymer such as a polynucleotide, for instance DNA or RIMA, a modified polynucleotide (such as abasic DNA), PNA, LNA, polyethylene glycol (PEG) or a polypeptide. The leader sequence preferably comprises a single strand of DNA, such as a poly dT section. The leader sequence can be any length, but are typically from about to about 150 nucleotides in length, such as from about 20 to about 120, from about 30 to about 100, from about 40 to about 80 or from about 50 to about 70 nucleotides in length.

[0186] The sequencing adaptor may be a hairpin loop adaptor. Hairpin loop adaptors are adaptors comprising a single polynucleotide strand, wherein the ends of the polynucleotide strand are capable of hybridising to each other, or are hybridized to each other, and wherein the middle section of the polynucleotide forms a loop. Suitable hairpin loop adaptors can be designed using methods known in the art. Typically, the 3' end of a hairpin loop adaptor attaches to the 5' end of a first strand of a double-stranded polynucleotide and the 5' end of the hairpin loop adaptor attaches to the 3' end of a second strand of a double-stranded polynucleotide; or the 5' end of a hairpin loop adaptor attaches to the 3' end of a first strand of a double-stranded polynucleotide and the 3' end of the hairpin loop adaptor attaches to the 5' end of a second strand of a double-stranded polynucleotide. As explained in more detail below, sequencing adaptors can be attached to a target polynucleotide in order to characterise the target polynucleotide.

[0187] The sequencing adaptor of the invention preferably comprises a first strand comprising or consisting of the sequence shown in SEQ ID NO: 51 and a second strand comprising or consisting of the sequence shown in SEQ ID NO: 52.

[0188] Population of sequencing adaptors

[0189] The invention also provides a population of two or more sequencing adaptors of the invention. This allows the characterisation or sequencing of a plurality of RNA molecules. The sequencing adaptors may be any of those defined above. The sequencing adaptors are typically identical. The population may comprise any number of sequencing adaptors, such as about 10, about 100, about 500, about 1000, about 5000 or about 10,000 or more sequencing adaptors.

[0190] Methods of making the sequencing adaptors The invention also provides a method of making a polynucleotide sequencing adaptor of the invention, the method comprising modifying a sequencing adaptor to comprise a RIMA polynucleotide upstream 5' to 3' of a loading site for a polynucleotide binding protein. The invention also provides a method of making a population of polynucleotide sequencing adaptor of the invention, the method comprising modifying sequencing adaptors to comprise a RNA polynucleotide upstream 5' to 3' of loading sites for a polynucleotide binding protein. The sequencing adaptors, RNA polynucleotide and polynucleotide binding sites may be any of those defined above. The method preferably further comprises loading a polynucleotide binding protein onto the sequencing adaptor(s).

[0191] General embodiments for adaptors Spacers

[0192] Any of the polynucleotides described herein, including the double stranded polynucleotide adaptors and sequencing adaptors, may comprise one or more spacers, e.g., from about one to about 10 spacers, e.g., from about 1 to about 5 spacers, e.g., about 1, 2, 3, 4 or 5 spacers. The spacer may comprise any suitable number of spacer units. A spacer typically provides an energy barrier which impedes movement of a polynucleotide binding protein. For example, a spacer may impede movement of a polynucleotide binding protein by reducing the traction of the protein, e.g., using an abasic spacer. A spacer may physically block movement of the protein, for instance by introducing a bulky chemical group to physically impede the movement of the polynucleotide binding protein.

[0193] One or more spacers are typically included in the adaptor to provide a distinctive signal when they pass through or across a nanopore. One or more spacers may be used to define or separate one or more regions of a polynucleotide, e.g., to separate an adaptor from the target polynucleotide.

[0194] A spacer may comprise a linear molecule, such as a polymer, e.g., a polypeptide or a polyethylene glycol (PEG). Typically, such a spacer has a different structure from the target polynucleotide. For instance, if the target polynucleotide is DNA, the or each spacer typically does not comprise DNA. In particular, if the target polynucleotide is deoxyribonucleic acid (DNA) or ribonucleic acid (RNA), the or each spacer preferably comprises peptide nucleic acid (PNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), locked nucleic acid (LNA) or a synthetic polymer with nucleotide side chains. A spacer may comprise one or more nitroindoles, one or more inosines, one or more acridines, one or more 2- aminopurines, one or more 2-6-diaminopurines, one or more 5-bromo-deoxyuridines, one or more inverted thymidines (inverted dTs), one or more inverted dideoxy-thymidines (ddTs), one or more dideoxy-cytidines (ddCs), one or more 5-methylcytidines, one or more 5-hydroxymethylcytidines, one or more 2'-O-Methyl RNA bases, one or more Isodeoxycytidines (Iso-dCs), one or more Iso-deoxyguanosines (Iso-dGs), one or more C3 (OC3H6OPO3) groups, one or more photo-cleavable (PC) [OC3H6-C(O)NHCH2-C6H3NO2- CH(CH3)OPO3] groups, one or more hexandiol groups, one or more spacer 9 (iSp9) [(OCH2CH2)3OPO3] groups, or one or more spacer 18 (iSplS) [(OCH2CH2)6OPO3] groups; or one or more thiol connections. A spacer may comprise any combination of these groups. Many of these groups are commercially available from IDT® (Integrated DNA Technologies®). For example, C3, iSp9 and iSpl8 spacers are all available from IDT®. A spacer may comprise any number of the above groups as spacer units.

[0195] A spacer may comprise one or more chemical groups, e.g., one or more pendant chemical groups. The one or more chemical groups may be attached to one or more nucleobases in an adaptor. The one or more chemical groups may be attached to the backbone of an adaptor. Any number of appropriate chemical groups may be present, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more. Suitable groups include, but are not limited to, fluorophores, streptavidin and / or biotin, cholesterol, methylene blue, dinitrophenols (DNPs), digoxigenin and / or anti-digoxigenin and dibenzylcyclooctyne groups.

[0196] A spacer may comprise one or more abasic nucleotides ( / .e., nucleotides lacking a nucleobase), such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more abasic nucleotides. The nucleobase can be replaced by -H (idSp) or -OH in the abasic nucleotide. Abasic spacers can be inserted into target polynucleotides by removing the nucleobases from one or more adjacent nucleotides. For instance, polynucleotides may be modified to include 3- methyladenine, 7-methylguanine, l,N6-ethenoadenine inosine or hypoxanthine and the nucleobases may be removed from these nucleotides using Human Alkyladenine DNA Glycosylase (hAAG). Alternatively, polynucleotides may be modified to include uracil and the nucleobases removed with Uracil-DNA Glycosylase (UDG). The one or more spacers preferably do not comprise any abasic nucleotides.

[0197] Suitable spacers can be designed or selected depending on the nature of the adaptor, the polynucleotide binding protein, and the conditions under which the method is to be carried out.

[0198] Tags

[0199] Any of the polynucleotides, including the double stranded polynucleotide adaptors and sequencing adaptors, used in the invention may comprise a tag or tether. For example, a polynucleotide can bind to a tag on a nanopore, e.g., via its adaptor, and release at some point, e.g., during characterization of the polynucleotide by the nanopore. A strong non- covalent bond (e.g., biotin / avidin) is still reversible and can be useful in some embodiments of the methods described herein.

[0200] The pair of pore tag and adaptor can be configured such that the binding strength or affinity of a binding site on the polynucleotide (e.g., a binding site provided by an anchor or a leader sequence of an adaptor or by a capture sequence within the duplex stem of an adaptor) to a tag on a nanopore is sufficient to maintain the coupling between the nanopore and polynucleotide until an applied force is placed on it to release the bound polynucleotide from the nanopore.

[0201] The tags or tethers are preferably uncharged. This can ensure that the tags or tethers are not drawn into the nanopore under the influence of a potential difference.

[0202] One or more molecules that attract or bind the polynucleotide or adaptor may be linked to the detector (e.g., the pore). Any molecule that hybridizes to the adaptor and / or target polynucleotide may be used. The molecule attached to the pore may be selected from a PNA tag, a PEG linker, a short oligonucleotide, a positively charged amino acid and an aptamer. Pores having such molecules linked to them are known in the art. For example, pores having short oligonucleotides attached thereto are disclosed in Howarka et al (2001) Nature Biotech. 19: 636-639 and WO 2010 / 086620, and pores comprising PEG attached within the lumen of the pore are disclosed in Howarka et al (2000) J. Am. Chem. Soc. 122(11): 2411- 2416.

[0203] A short oligonucleotide attached to the detector (e.g., a nanopore), which oligonucleotide comprises a sequence complementary to a sequence in the leader sequence or another single stranded sequence in the adaptor may be used to enhance capture of the target polynucleotide in the methods described herein.

[0204] The tag or tether may comprise or be an oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino). The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The oligonucleotide (e.g., DNA, RNA, LNA, BNA, PNA, or morpholino) for use in the tag or tether can have at least one end (e.g., 3'- or 5'-end) modified for conjugation to other modifications or to a solid substrate surface including, e.g., a bead. The end modifiers may add a reactive functional group which can be used for conjugation. Examples of functional groups that can be added include, but are not limited to amino, carboxyl, thiol, maleimide, aminooxy, and any combinations thereof. The functional groups can be combined with different length of spacers (e.g., C3, C9, C12, Spacer 9 and 18) to add physical distance of the functional group from the end of the oligonucleotide sequence.

[0205] The tag or tether may comprise or be a morpholino oligonucleotide. The morpholino oligonucleotide can have about 10-30 nucleotides in length or about 10-20 nucleotides in length. The morpholino oligonucleotides can be modified or unmodified. For example, the morpholino oligonucleotide can be modified on the 3' and / or 5' ends of the oligonucleotides. Examples of modifications on the 3' and / or 5' end of the morpholino oligonucleotides include, but are not limited to 3' affinity tag and functional groups for chemical linkage (including, e.g., 3'-biotin, 3'-primary amine, 3'-disulfide amide, 3'-pyridyl dithio, and any combinations thereof); 5' end modifications (including, e.g., 5'-primary ammine, and / or 5'- dabcyl), modifications for click chemistry (including, e.g., 3'-azide, 3'-alkyne, 5'-azide, 5'- alkyne), and any combinations thereof.

[0206] The tag or tether may further comprise a polymeric linker, e.g., to facilitate coupling to a detector e.g., a nanopore. An exemplary polymeric linker includes, but is not limited to, polyethylene glycol (PEG). The polymeric linker may have a molecular weight of about 500 Da to about 10 kDa (inclusive), or about 1 kDa to about 5 kDa (inclusive). The polymeric linker (e.g., PEG) can be functionalized with different functional groups including, e.g., but not limited to maleimide, NHS ester, dibenzocyclooctyne (DBCO), azide, biotin, amine, alkyne, aldehyde, and any combinations thereof. The tag or tether may further comprise a 1 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 2 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 3 kDa PEG with a 5'-maleimide group and a 3'-DBCO group. The tag or tether may further comprise a 5 kDa PEG with a 5'-maleimide group and a 3'-DBCO group.

[0207] The tag or tether may be a peptide, polypeptide or protein. The tag or tether may be a CsgA polypeptide. The CsgA polypeptide may be any of those described in WO 2024 / 100270 (incorporated herein in its entirety).

[0208] Other examples of a tag or tether include, but are not limited to His tags, biotin or streptavidin, antibodies that bind to analytes, aptamers that bind to analytes, analyte binding domains such as DNA binding domains (including, e.g., peptide zippers such as leucine zippers, single-stranded DNA binding proteins (SSB)), and any combinations thereof.

[0209] The tag or tether may be attached to the external surface of a nanopore, e.g., on the cis side of a membrane, using any methods known in the art. For example, one or more tags or tethers can be attached to the nanopore via one or more cysteines (cysteine linkage), one or more primary amines such as lysines, one or more non-natural amino acids, one or more histidines (His tags), one or more biotin or streptavidin, one or more antibody-based tags, one or more enzyme modification of an epitope (including, e.g., acetyl transferase), and any combinations thereof. Suitable methods for carrying out such modifications are well-known in the art. Suitable non-natural amino acids include, but are not limited to, 4-azido-L- phenylalanine (Faz) and any one of the amino acids numbered 1-71 in Figure 1 of Liu C. C. and Schultz P. G., Annu. Rev. Biochem., 2010, 79, 413-444.

[0210] Where one or more tags or tethers are attached to a nanopore via cysteine linkage(s), the one or more cysteines can be introduced to one or more monomers that form the nanopore by substitution. The nanopore may be chemically modified by attachment of (i) Maleimides including diabromomaleimides such as: 4-phenylazomaleinanil, l.N-(2- Hydroxyethyl)maleimide, N-Cyclohexylmaleimide, 1.3-Maleimidopropionic Acid, 1.1-4- Aminophenyl-lH-pyrrole,2,5,dione, l.l-4-Hydroxyphenyl-lH-pyrrole,2,5,dione, N- Ethylmaleimide, N-Methoxycarbonylmaleimide, N-tert-Butylmaleimide, N-(2- Aminoethyl)maleimide , 3-Maleimido-PROXYL , N-(4-Chlorophenyl)maleimide, l-[4- (dimethylamino)-3,5-dinitrophenyl]-lH-pyrrole-2, 5-dione, N-[4-(2- Benzimidazolyl)phenyl]maleimide, N-[4-(2-benzoxazolyl)phenyl]maleimide, N-(l-naphthyl)- maleimide, N-(2,4-xylyl)maleimide, N-(2,4-difluorophenyl)maleimide , N-(3-chloro-para- tolyl)-maleimide, l-(2-amino-ethyl)-pyrrole-2, 5-dione hydrochloride, l-cyclopentyl-3- methyl-2,5-dihydro-lH-pyrrole-2, 5-dione, l-(3-aminopropyl)-2,5-dihydro-lH-pyrrole-2,5- dione hydrochloride, 3-methyl-l-[2-oxo-2-(piperazin-l-yl)ethyl]-2,5-dihydro-lH-pyrrole- 2, 5-dione hydrochloride, l-benzyl-2,5-dihydro-lH-pyrrole-2, 5-dione, 3-methyl-l-(3,3,3- trifluropropyl)-2,5-dihydro-lH-pyrrole-2, 5-dione, l-[4-(methylamino)cyclohexyl]-2,5- dihydro-lH-pyrrole-2, 5-dione trifluroacetic acid, SMILES O=C1C=CC(=O)N1CC=2C=CN=CC2, SMILES O=C1C=CC(=O)N1CN2CCNCC2, l-benzyl-3- methyl-2,5-dihydro-lH-pyrrole-2, 5-dione, l-(2-fluorophenyl)-3-methyl-2,5-dihydro 1H- pyrrole-2, 5-dione, N-(4-phenoxyphenyl)maleimide , N-(4-nitrophenyl)maleimide (ii) lodocetamides such as :3-(2-Iodoacetamido)-proxyl, N-(cyclopropylmethyl)-2- iodoacetamide, 2-iodo-N-(2-phenylethyl)acetamide, 2-iodo-N-(2,2,2- trifluoroethyl)acetamide, N-(4-acetylphenyl)-2-iodoacetamide, N-(4- (aminosulfonyl)phenyl)-2-iodoacetamide, N-(l,3-benzothiazol-2-yl)-2-iodoacetamide, N- (2,6-diethylphenyl)-2-iodoacetamide, N-(2-benzoyl-4-chlorophenyl)-2-iodoacetamide, (iii) Bromoacetamides: such as N-(4-(acetylamino)phenyl)-2-bromoacetamide , N-(2- acetylphenyl)-2-bromoacetamide , 2-bromo-n-(2-cyanophenyl)acetamide, 2-bromo-N-(3- (trifluoromethyl)phenyl)acetamide, N-(2-benzoylphenyl)-2-bromoacetamide , 2-bromo-N- (4-fluorophenyl)-3-methylbutanamide, N-Benzyl-2-bromo-N-phenylpropionamide, N-(2- bromo-butyryl)-4-chloro-benzenesulfonamide, 2-Bromo-N-methyl-N-phenylacetamide, 2- bromo-N-phenethyl-acetamide,2-adamantan-l-yl-2-bromo-N-cyclohexyl-acetamide, 2- bromo-N-(2-methylphenyl)butanamide, Monobromoacetanilide, (iv) Disulphides such as: aldrithiol-2 , aldrithiol-4 , isopropyl disulfide, l-(Isobutyldisulfanyl)-2-methylpropane, Dibenzyl disulfide, 4-aminophenyl disulfide, 3-(2-Pyridyldithio)propionic acid, 3-(2- Pyridyldithio)propionic acid hydrazide, 3-(2-Pyridyldithio)propionic acid N-succinimidyl ester, am6amPDPl-[3CD and (v) Thiols such as: 4-Phenylthiazole-2-thiol, Purpald, 5, 6, 7, 8- tetrahydro-quinazoline-2-thiol.

[0211] The tag or tether may be attached directly to a nanopore or via one or more linkers. The tag or tether may be attached to the nanopore using the hybridization linkers described in WO 2010 / 086602 (incorporated herein by reference in its entirety). Alternatively, peptide linkers may be used. Peptide linkers are amino acid sequences. The length, flexibility and hydrophilicity of the peptide linker are typically designed such that it does not to disturb the functions of the monomer and pore. Preferred flexible peptide linkers are stretches of 2 to 20, such as 4, 6, 8, 10 or 16, serine and / or glycine amino acids. More preferred flexible linkers include (SG)i, (SG)2, (SG)3, (SG)4, (SG)5and (SG)8wherein S is serine and G is glycine. Preferred rigid linkers are stretches of 2 to 30, such as 4, 6, 8, 16 or 24, proline amino acids. More preferred rigid linkers include (P)i2wherein P is proline.

[0212] Suitable pore tags are also described in WO 2018 / 100370, which describes non-hairpin methods for characterising double-stranded polynucleotides and is herein incorporated by reference in its entirety.

[0213] Anchor

[0214] Any of the polynucleotides, including the double stranded polynucleotide adaptors and sequencing adaptors, of the invention may comprise a membrane anchor. The anchor typically assists in the characterisation of a target polynucleotide in accordance with the methods disclosed herein. For example, a membrane anchor may promote localisation of the selected polynucleotides around a nanopore.

[0215] The anchor may be a polypeptide anchor and / or a hydrophobic anchor that can be inserted into the membrane. The hydrophobic anchor is preferably a lipid, fatty acid, sterol, carbon nanotube, polypeptide, protein, or amino acid, for example cholesterol, palmitate, or tocopherol. The anchor may comprise thiol, biotin, or a surfactant.

[0216] The anchor may be biotin (for binding to streptavidin), amylose (for binding to maltose binding protein or a fusion protein), Ni-NTA (for binding to poly-histidine or poly-histidine tagged proteins) or peptides (such as an antigen).

[0217] The anchor preferably comprises a linker, or 2, 3, 4 or more linkers. Preferred linkers include, but are not limited to, polymers, such as polynucleotides, polyethylene glycols (PEGs), polysaccharides and polypeptides. These linkers may be linear, branched, or circular. For instance, the linker may be a circular polynucleotide. The adaptor may hybridise to a complementary sequence on a circular polynucleotide linker. The one or more anchors or one or more linkers may comprise a component that can be cut or broken down, such as a restriction site or a photolabile group. The linker may be functionalised with maleimide groups to attach to cysteine residues in proteins. Suitable linkers are described in WO 2010 / 086602 (incorporated herein by reference in its entirety).

[0218] The anchor is preferably cholesterol or a fatty acyl chain. For example, any fatty acyl chain having a length of from 6 to 30 carbon atom, such as hexadecanoic acid, may be used. Examples of suitable anchors and methods of attaching anchors to adaptors are disclosed in WO 2012 / 164270 and WO 2015 / 150786 (incorporated herein by reference in their entireties).

[0219] The anchor may consist or comprise a hydrophobic modification to the adaptor. The hydrophobic modification may comprise a modified phosphate group comprised within the polynucleotide or polynucleotide anchor. The hydrophobic modification may for example comprise a phosphorothioate such as a charge-neutralized alkyl-phosphorothioate (PPT) as described in Jones et al, J. Am. Chem. Soc. 2021, 143, 22, 8305, the entire contents of which are hereby incorporated by reference. Suitable alkyl groups include for example Ci- Cio alkyl groups such as C2-C6alkyl groups, e.g., methyl, ethyl, propyl, butyl, pentyl and hexyl groups. Incorporation of the charge-neutralized alkyl-phosphorothioate into a polynucleotide allows for the polynucleotide to anchor to a hydrophobic region such as a lipid bilayer.

[0220] Biotin enrichment

[0221] Any of the polynucleotides, including the double stranded polynucleotide adaptors and sequencing adaptors, preferably comprise biotin. The biotin may be used to isolate barcoded constructs. Suitable methods for biotin-based enrichment are known in the art. For instance, a surface, such as a bead, comprising avidin / or streptavidin may be used to barcoded constructs comprising biotin.

[0222] Kits

[0223] The invention also provides kits for uniquely labelling a RIMA molecule or a plurality of RNA molecules. In one embodiment, the kit comprises (a) a double stranded polynucleotide adaptor of the invention and (b) a polynucleotide sequencing adaptor of the invention. The kit may further comprise a polynucleotide binding protein loaded on the sequencing adaptor.

[0224] In another embodiment, the kit comprises (a) a population of double stranded polynucleotide adaptors of the invention and (b) a polynucleotide sequencing of the invention or a population of sequencing adaptors of the invention. The kit may further comprise a polynucleotide binding protein loaded on the sequencing adaptor(s).

[0225] The double stranded polynucleotide, the polynucleotide sequencing adaptor, the population and the polynucleotide binding protein may be any of those defined above.

[0226] The kit may additionally comprise one or more other reagents or instruments which enable any of the embodiments mentioned above or below to be carried out. Such reagents or instruments include one or more of the following: suitable buffer(s) (aqueous solutions), means to obtain a sample from a subject (such as a vessel or an instrument comprising a needle), means to amplify and / or express polynucleotides, a membrane as defined below or voltage or patch clamp apparatus. Reagents may be present in the kit in a dry state such that a fluid sample is used to resuspend the reagents. The kit may also, optionally, comprise instructions to enable the kit to be used in the methods described herein or details regarding for which organism the method may be used. The kit may comprise a magnet or an electromagnet. The kit may, optionally, comprise nucleotides.

[0227] Methods

[0228] The invention also provides a method of uniquely labelling a RNA molecule or a plurality of RNA molecules. The RNA molecule(s) may be any of those defined above. The plurality of RNA molecules may comprise any number of RNA molecules discussed above.

[0229] In one embodiment, the method comprises (a) hybridising the RNA molecule to a double stranded polynucleotide adaptor of the invention, and (b) ligating the hybridised double stranded polynucleotide adaptor to the RNA molecule. The double stranded polynucleotide adaptor may be any of those defined above. The barcode in the double stranded polynucleotide adaptor uniquely labels the RNA molecule.

[0230] In another embodiment, the method comprises (a) hybridising the RNA molecules to a population of double stranded polynucleotide adaptors of the invention, and (b) ligating the hybridised double stranded polynucleotide adaptors to the RNA molecules. The population of double stranded polynucleotide adaptors of the invention may be any of those defined above. Each RNA molecule in the plurality is typically uniquely labelled with a different double stranded polynucleotide adaptor comprising a different barcode.

[0231] Any hybridisation conditions may be used for step (a). Suitable conditions for hybridisation are known in the art and described above and in the Examples.

[0232] The invention also provides a method of uniquely labelling a plurality of RNA molecules, the method comprises ligating a population of single stranded polynucleotide adaptors to the RNA molecules, wherein each single stranded polynucleotide adaptor comprises a different barcode. The single stranded polynucleotide adaptors may be any of the first strands defined above. The barcodes may be any of those defined above. The population may have any of the characteristics defined above. Each RNA molecule in the plurality is typically uniquely labelled with a different single stranded polynucleotide adaptor comprising a different barcode.

[0233] Ligation is the joining of two polynucleotides most commonly through the action of an enzyme or by chemical means. The ends of polynucleotides are joined together by the formation of phosphodiester bonds between the 3'-hydroxyl of one polynucleotide terminus with the 5'-phosphoryl of another. A co-factor is generally involved in the reaction, and this is usually ATP or NAD+. In the context of the invention, the double stranded polynucleotide adaptor(s) and the RNA molecule(s) are held adjacent to each other by hybridisation, and this facilitates the ligation reaction.

[0234] The double stranded polynucleotide adaptor(s) may be ligated to either end of the RNA molecule(s), i.e. the 5' or the 3' end. The double stranded polynucleotide adaptor(s) may be ligated to both ends of the RNA molecule(s). Preferably, the double stranded polynucleotide adaptor(s) are ligated to the 3’ end of the RNA polynucleotide(s).

[0235] Any method of ligation may be used in accordance with the invention, including the method described in the Examples. The skilled person understands the ligases that may be used in the invention, such as T4 DNA ligase, E. coli DNA ligase, Taq DNA ligase, Tma DNA ligase, 9°N DNA ligase, T4 Polymerase I, T4 Polymerase 2, Thermostable 5’ App DNA / RNA ligase, SplintR, circ Ligase, T4 RNA ligase 1 or T4 RNA ligase 2. The double stranded polynucleotide adaptor(s) may be ligated to the RNA molecule(s) in the absence of ATP or using gamma-S- ATP (ATPyS) instead of ATP. Suitable conditions for ligation reactions are known in the art and described in the Examples.

[0236] Step (b) preferably comprises ligating the hybridised double stranded polynucleotide adaptor(s) to the RNA molecule(s) using a ligase. The method preferably comprises ligating a population of single stranded polynucleotide adaptors to the RNA molecules using a ligase. The method preferably further comprises (c) removing the ligase from the method conditions.

[0237] Preferably, the RNA molecule(s) is / are tRNA. In this embodiment, the double stranded polynucleotide adaptor(s) is / are as defined in the second embodiment above, i.e., the overhang is capable of hybridising to the tail on tRNA molecule(s). In this embodiment, the method further comprises (c) reverse transcribing part of the first RNA strand(s) and the tRNA molecule(s) to form a linear double stranded construct or linear double stranded constructs. The RT primer(s) in the second strand(s) may be used to prime the reverse transcription reaction. The linear double stranded construct(s) comprise one strand comprising the linearised tRNA ligated to the first RNA strand and another strand comprising the complementary DNA (cDNA) produced by RT. The first RNA strand comprises the Barcode. The one strand from each linear double stranded construct may be characterised or sequenced as described below.

[0238] Methods for conducting RT are well known in the art and any suitable conditions may be used. Preferred conditions are described in the Examples. RT involves the use of the enzyme reverse transcriptase to convert RNA into cDNA. Reverse transcriptases are commercially available (e.g. Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753), Superscript® II reverse transcriptase (Invitrogen) and Affinity script (Agilent)). The reverse transcriptase is preferably Maxima H minus reverse transcriptase (ThermoFisher Scientific, EP0753).

[0239] The method of the invention further preferably further comprises hybridising and ligating a sequencing adaptor or a population of sequencing adaptors to the double stranded polynucleotide adaptor(s) or linear double stranded construct(s). The sequencing adaptor or population of sequencing adaptors are preferably a sequencing adaptor or population of the invention. They may be any of the sequencing adaptors or populations defined above. Hybridisation and ligation are also described above and any of those embodiments equally apply here.

[0240] The method produces a uniquely labelled RNA molecule or a plurality of uniquely labelled RIMA molecules. The RNA molecule is uniquely labelled with a barcode. The RNA molecules are each uniquely labelled with a different barcode. In the following discussion, all of these will be collectively called "uniquely labelled RNA molecule(s)". They may also be called "barcoded RNA molecule(s)". These terms are interchangeable.

[0241] The method preferably further comprises isolating the uniquely labelled RNA molecule(s). Any method of isolation may be used including any of the ones discussed below. Any of the adaptors can comprise sequences and / or molecules which facilitate isolation.

[0242] The method preferably further comprises amplifying the uniquely labelled RNA molecule(s). Any amplification method may be used, including polymerase chain reaction (PCR), linear RNA amplification and terminal continuation RNA amplification methodology. Any of the adaptors can comprise primer sites that allow amplification. This facilitates preparation of the uniquely labelled RNA molecule(s) for amplification. Amplification methods are routine in the art.

[0243] The method preferably further comprises isolating the uniquely labelled RNA molecule(s) and amplifying the uniquely labelled RNA molecule(s).

[0244] The method preferably further comprises characterising or sequencing the uniquely labelled RNA molecule(s). Any method may be used for sequencing or characterising the uniquely labelled RNA molecule(s), including next generation sequencing (NGS). The uniquely labelled RNA molecule(s) is / are preferably characterised or sequenced using a nanopore. This is discussed in more detail below. In all these embodiment, the RNA molecule(s) and the barcode(s) are typically characterised or sequenced, for instance using a nanopore.

[0245] Characterising methods

[0246] The invention provides a method of characterising a RNA molecule or a plurality of RNA molecules, the method comprising (i) uniquely labelling the RNA molecule(s) using a method of the invention and (ii) characterising the uniquely labelled RIMA molecule(s). Step (ii) typically comprises characterising the RNA molecule(s) and the barcode(s). The characterisation is preferably sequencing of the uniquely labelled RNA molecule(s), i.e., the RNA molecule(s) and barcode(s), preferably using a nanopore. Step (ii) preferably comprises sequencing the RNA molecule(s) and the barcode(s). Step (ii) preferably comprises characterising or sequencing the RNA molecule(s) and the barcode(s) using a nanopore. Any embodiments discussed above, especially in relation to the adaptors of the invention, equally apply to these embodiments. As explained in more detail above, the labelling method of the invention preferably comprises hybridising and ligating a sequencing adaptor or a population of sequencing adaptors to the double stranded polynucleotide adaptor(s) or linear double stranded construct(s). The sequencing adaptor(s) may be used to characterise or sequence the uniquely labelled RNA molecule(s).

[0247] Any method of characterisation may be used. The method preferably uses next generation sequencing (NGS).

[0248] The uniquely labelled RNA molecule(s) are preferably moved with respect to a detector. The detector may be selected from (i) a zero-mode waveguide, (ii) a field-effect transistor, optionally a nanowire field-effect transistor; (iii) an AFM tip; (iv) a nanotube, optionally a carbon nanotube; and (v) a nanopore. Preferably, the detector is a nanopore.

[0249] The uniquely labelled RNA molecule(s) may be characterised in the method of the invention in any suitable manner. The uniquely labelled RNA molecule(s) are preferably characterised by detecting an ionic current or optical signal as they move with respect to a nanopore. This is described in more detail herein. The method is amenable to these and other methods of characterising polynucleotides.

[0250] In another non-limiting example, the uniquely labelled RNA molecule(s) are characterised by detecting the by-products of a polynucleotide-processing reaction, such as a sequencing by synthesis reaction. The method may thus involve detecting the product of the sequential addition of (poly)nucleotides by an enzyme such as a polymerase to the uniquely labelled RNA molecule(s). The product may be a change in one or more properties of the enzyme such as in the conformation of the enzyme. Such methods may thus comprise subjecting an enzyme such as polymerase or a reverse transcriptase to the uniquely labelled RNA molecule(s) as templates under conditions such that the template-dependent incorporation of nucleotide bases into a growing oligonucleotide strand causes conformational changes in the enzyme in response to sequentially encountering template nucleic acid bases and / or incorporating template-specified natural or analog bases ( / .e., an incorporation event), detecting the conformational changes in the enzyme in response to such incorporation events, and thereby detecting the sequence of the templates. In such methods, the uniquely labelled RNA molecule(s) may be moved in accordance with the method of the invention. Such methods may involve detecting and / or measuring incorporation events using methods known to those skilled in the art, such as those described in US 2017 / 0044605.

[0251] In another embodiment, by-products may be labelled so that a phosphate labelled species is released upon the addition of a nucleotide to a synthesised nucleic acid strand that is complementary to the template uniquely labelled RNA molecule(s), and the phosphate labelled species is detected e.g., using a detector as described herein. The uniquely labelled RNA molecule(s) being characterised in this way may be moved in accordance with the methods herein. Suitable labels may be optical labels that are detected using a nanopore, or a zero-mode wave guide, or by Raman spectroscopy, or other detectors. Suitable labels may be non-optical labels that are detected using a nanopore, or other detectors.

[0252] In another approach, nucleoside phosphates (nucleotides) are not labelled and upon the addition of nucleotides to synthesised nucleic acid strands that are complementary to the uniquely labelled RNA molecule(s), natural by-product species are detected. Suitable detectors may be ion-sensitive field-effect transistors, or other detectors.

[0253] These and other detection methods are suitable for use in the methods described herein. Any suitable measurements can be taken using a detector as the uniquely labelled RNA molecule(s) move with respect to the detector.

[0254] Nanopore characterisation

[0255] The uniquely labelled RNA molecule(s) are preferably characterised using a nanopore.

[0256] The method preferably comprises (i) contacting the uniquely labelled RNA molecule(s) with a nanopore such that the uniquely labelled RNA molecule(s) move with respect to the nanopore and (ii) taking one or more measurements as the uniquely labelled RNA molecule(s) move with respect to the nanopore wherein the measurements are indicative of one or more characteristics of the uniquely labelled RNA molecule(s) and thereby characterising the uniquely labelled RNA molecule(s). The one or more characteristics are preferably selected from (i) the length of the uniquely labelled RNA molecule(s), (ii) the identity of the uniquely labelled RNA molecule(s), (iii) the sequence of the uniquely labelled RNA molecule(s), (iv) the secondary structure of the uniquely labelled RNA molecule(s) and (v) whether or not the uniquely labelled RNA molecule(s) is modified. The uniquely labelled RNA molecule(s) may be modified by methylation, by oxidation, by damage, with one or more proteins or with one or more labels, tags, or spacers. The one or more characteristics of the uniquely labelled RNA molecule(s) are preferably measured by electrical measurement and / or optical measurement. The electrical measurement is preferably a current measurement, an impedance measurement, a tunnelling measurement, or a field effect transistor (FET) measurement. The method more preferably comprises (i) contacting the uniquely labelled RIMA molecule(s) with a nanopore such that the uniquely labelled RNA molecule(s) move through the nanopore and (ii) measuring the current moving through the nanopore as the uniquely labelled RNA molecule(s) move through the nanopore wherein the current is indicative of one or more characteristics of the uniquely labelled RNA molecule(s) and thereby characterising the uniquely labelled RNA molecule(s). The one or more characteristics may be any of those defined above.

[0257] The movement of the uniquely labelled RNA molecule(s) with respect to the nanopore or through the nanopore is preferably controlled using a polynucleotide binding protein. The use of such proteins in nanopore sequencing is known. Examples of suitable proteins are discussed in more detail below.

[0258] Any suitable nanopore can be used. The nanopore is preferably a transmembrane pore. A transmembrane pore is a structure that crosses the membrane to some degree. It permits hydrated ions driven by an applied potential to flow across or within the membrane. The transmembrane pore typically crosses the entire membrane so that hydrated ions may flow from one side of the membrane to the other side of the membrane. However, the transmembrane pore does not have to cross the membrane. It may be closed at one end. For instance, the pore may be a well, gap, channel, trench or slit in the membrane along which or into which hydrated ions may flow.

[0259] The nanopore typically has a first opening and a second opening. The first opening is typically the cis opening and the second opening is typically the trans opening. However, the first opening may be the trans opening and the second opening may be the cis opening. Any polynucleotide binding protein used in the method of the invention is typically provided at the first opening of the nanopore and thus controls the movement of the target polynucleotide in the direction from the second opening of the nanopore towards the first opening of the nanopore.

[0260] Any transmembrane pore may be used in the method of the invention. The pore may be biological or artificial. Suitable pores include, but are not limited to, protein pores, polynucleotide pores and solid-state pores. The pore may be a DNA origami pore (Langecker et al., Science, 2012; 338: 932-936). Suitable DNA origami pores are disclosed in WO2013 / 083983.

[0261] The nanopore is preferably a transmembrane protein pore. A transmembrane protein pore is a polypeptide or a collection of polypeptides that permits hydrated ions, such as polynucleotide, to flow from one side of a membrane to the other side of the membrane. In the method of the invention, the transmembrane protein pore is capable of forming a pore that permits hydrated ions driven by an applied potential to flow from one side of the membrane to the other. The transmembrane protein pore preferably permits polynucleotides to flow from one side of the membrane, such as a triblock copolymer membrane, to the other. The transmembrane protein pore allows a polynucleotide to be moved through the pore.

[0262] The nanopore may be a transmembrane protein pore which is a monomer or an oligomer. The pore is preferably made up of several repeating subunits, such as at least about 6, at least about 7, at least about 8, at least about 9, at least about 10, at least about 11, at least about 12, at least about 13, at least about 14, at least about 15, or at least about 16 subunits. The pore is preferably a hexameric, heptameric, octameric or nonameric pore. The pore may be a homo-oligomer or a hetero-oligomer.

[0263] The transmembrane protein pore may comprise a barrel or channel through which the ions may flow. The subunits of the pore typically surround a central axis and contribute strands to a transmembrane [3-barrel or channel or a transmembrane a-helix bundle or channel.

[0264] Typically, the barrel or channel of the transmembrane protein pore comprises amino acids that facilitate interaction with an analyte, such as a target polynucleotide (as described herein). These amino acids are preferably located near a constriction of the barrel or channel. The transmembrane protein pore typically comprises one or more positively charged amino acids, such as arginine, lysine or histidine, or aromatic amino acids, such as tyrosine or tryptophan. These amino acids typically facilitate the interaction between the pore and nucleotides, polynucleotides, or nucleic acids.

[0265] The transmembrane protein pore may be from or derived from Wza, Iota toxin, Anthrax protective antigen, Vibrio cholerae cytolysin, Cytotoxin K (CytK), CELIII, CsgG, CsgF, CsgG- CsgF, Aerolysin, alpha hemolysin, MspA, MspB, MspC, PorARr, PorBRr, PorARc, PilQ, necrotic enteritis B-like toxin (NetB), FraC, portal proteins including G20c, P23_45, T4, SPP1, P22 and Phi29, gamma hemolysin, Monalysin, Lysenin, ClyA, an actinoporin, Clostridium perfringens beta toxin, parasporin-2, epsilon toxin, lectin from the parasitic mushroom Laetiporus sulphureus (LSL), volvatoxin, Cry toxins, CytlAa, Cyt2Aa, Complement component 9 (C9), Perfringolysin O, Pleurotolysin, Listeriolysin, Perforin-2, Gasdermin-A3, L-, P- and M-ring protein, Type II secretion system protein D, GspD, InvG, VirB7, SpoIIIAG, Cag8, Cag3, Cag or other proteins in the Type IV secretion system apparatus protein CagY, WzzB, Pentraxin, Afp2, Major vault protein, Thioredoxin-dependent peroxidase reductase, Arf-GAP, Respiratory syncytial virus ribonucleoprotein, Chikungunya virus nonstructural protein 1, PRC, YaxA, XaxA, HfaB, NfpAB, leukocidin and PrgH.

[0266] Suitable transmembrane protein pores for use in the invention include those described in WO 2016 / 034591, WO 2017 / 149316, WO 2017 / 149317, WO 2017 / 149318, WO 2018 / 211241, WO 2019 / 002893, WO 2023 / 118404, WO 2023 / 198911, WO 2024 / 033421, WO 2024 / 033422, WO 2024 / 033443, and WO 2024 / 089270 (all incorporated by reference herein in their entirety).

[0267] The transmembrane protein pore may also be any of the CsgG pores described in WO 2023 / 060420, WO 2023 / 60418, WO 2023 / 60422, WO 2023 / 060421, WO 2023 / 019470, CN114957412, WO 2023 / 019471, W02023 / 060419 and WO 2023 / 050031 (all incorporated herein by reference in their entireties) or a variant thereof.

[0268] The transmembrane protein pore may be any of the pores described in WO 2023 / 123370, WO 2024 / 138470, WO 2024 / 138472, WO 2024 / 138424, WO 2024 / 138425, WO 2024 / 138512 and WO 2024 / 138565 (all incorporated herein by reference in their entireties) or a variant thereof.

[0269] The transmembrane pore may be formed from a chimeric pore monomer comprising two or more regions, wherein at least two of the two or more regions are from at least two different pores. The chimeric pore monomer may comprise any number of regions, such as three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more or ten or more regions, from different pores. The chimeric pore monomer may comprise two or three regions. The regions are preferably selected from a cap region, a constriction region, and a transmembrane region. The regions may be a cap region and a constriction region. The regions may be a cap region, a constriction region, and a transmembrane region. The at least two different pores are typically at least two different pores that appear in nature. The at least two different pores are typically at least two different wild-type or naturally occurring pores. The at least two different pores are preferably different before any artificial or synthetic modifications, such as additions, deletions and / or substitutions, are made to them. The at least two different pores are preferably homologues, for example structural homologues. A structural homologue refers to a protein or molecule that shares a similar three-dimensional structure with another protein or molecule. This can be determined using standard methods in the art (e.g., AlphaFold or PSIPRED). Structural homologues typically have similar sequences. Structural homologues are normally identified in similar species. The at least two different pores may be selected from any of the pores listed above. The at least two different pores may be two different PorARc pores or three different PorARc pores. The at least two different pores may be two different CsgG pores or three different CsgG pores. The chimeric pore monomer may be any of those described in PCT / EP2023 / 080135 (incorporated by reference herein in its entirety).

[0270] The transmembrane pore may be formed from a pore monomer comprising (a) a CsgG monomer and (b) a fusion polypeptide comprising a first portion comprising a CsgF peptide and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the pore monomer. The pore monomer may be derived from a protein transmembrane pore complex comprising (a) a CsgG transmembrane pore comprising a lumen and (b) a fusion polypeptide comprising a first portion comprising a CsgF protein and a second portion comprising a helix-forming auxiliary protein, wherein the fusion protein is attached to the transmembrane pore. The auxiliary protein can be designed de novo using computer-based structural analysis tools to confer certain desirable features to the CsgG monomer (e.g., modulation of pore width, lengthening of pore lumen, formation of one or more additional constrictions, etc.). The de novo designed auxiliary protein may form one or more additional constrictions in the lumen of a CsgG pore formed from the monomer, and improve discrimination of polymer units as an analyte moves through the pore. The pore monomer may be any of the pore monomers described in WO 2024 / 033447 (incorporated by reference herein in its entirety).

[0271] The transmembrane protein pore may be any of those described in UK Application No. 2413506.3 being co-filed with this application.

[0272] Membrane

[0273] The detector or nanopore is typically present in a membrane. Any suitable membrane may be used.

[0274] The membrane is preferably an amphiphilic layer. An amphiphilic layer is a layer formed from amphiphilic molecules, such as phospholipids, which have both hydrophilic and lipophilic properties. The amphiphilic molecules may be synthetic or naturally occurring. Non-naturally occurring amphiphiles and amphiphiles which form a monolayer are known in the art and include, for example, block copolymers (Gonzalez-Perez et al., Langmuir, 2009, 25, 10447-10450). Block copolymers are polymeric materials in which two or more monomer sub-units that are polymerized together to create a single polymer chain. Block copolymers typically have properties that are contributed by each monomer sub-unit. However, a block copolymer may have unique properties that polymers formed from the individual sub-units do not possess. Block copolymers can be engineered such that one of the monomer sub-units is hydrophobic ( / .e., lipophilic), whilst the other sub-unit(s) are hydrophilic whilst in aqueous media. In this case, the block copolymer may possess amphiphilic properties and may form a structure that mimics a biological membrane. The block copolymer may be a diblock (consisting of two monomer sub-units) but may also be constructed from more than two monomer sub-units to form more complex arrangements that behave as amphipiles. The copolymer may be a triblock, tetrablock or pentablock copolymer. The membrane may be a triblock copolymer membrane.

[0275] Archaebacterial bipolar tetraether lipids are naturally occurring lipids that are constructed such that the lipid forms a monolayer membrane. These lipids are generally found in extremophiles that survive in harsh biological environments, thermophiles, halophiles and acidophiles. Their stability is believed to derive from the fused nature of the final bilayer. It is straightforward to construct block copolymer materials that mimic these biological entities by creating a triblock polymer that has the general motif hydrophilic-hydrophobic- hydrophilic. This material may form monomeric membranes that behave similarly to lipid bilayers and encompass a range of phase behaviours from vesicles through to laminar membranes. Membranes formed from these triblock copolymers hold several advantages over biological lipid membranes. Because the triblock copolymer is synthesised, the exact construction can be carefully controlled to provide the correct chain lengths and properties required to form membranes and to interact with pores and other proteins.

[0276] Block copolymers may also be constructed from sub-units that are not classed as lipid submaterials; for example, a hydrophobic polymer may be made from siloxane or other non- hydrocarbon-based monomers. The hydrophilic sub-section of block copolymer can also possess low protein binding properties, which allows the creation of a membrane that is highly resistant when exposed to raw biological samples. This head group unit may also be derived from non-classical lipid head-groups.

[0277] Triblock copolymer membranes also have increased mechanical and environmental stability compared with biological lipid membranes, for example a much higher operational temperature or pH range. The synthetic nature of the block copolymers provides a platform to customise polymer-based membranes for a wide range of applications.

[0278] The membrane may be one of the membranes disclosed in International Application No. WO2014 / 064443 or WO2014 / 064444 (both of which are incorporated herein by reference in their entireties).

[0279] The amphiphilic molecules may be chemically modified or functionalised to facilitate coupling of the polynucleotide. The amphiphilic layer may be a monolayer or a bilayer. The amphiphilic layer is typically planar. The amphiphilic layer may be curved. The amphiphilic layer may be supported.

[0280] Amphiphilic membranes are typically naturally mobile, essentially acting as two-dimensional fluids with lipid diffusion rates of approximately IO-8cm s4. This means that the pore and coupled polynucleotide can typically move within an amphiphilic membrane.

[0281] The membrane may be a lipid bilayer. Lipid bilayers are models of cell membranes and serve as excellent platforms for a range of experimental studies. For example, lipid bilayers can be used for in vitro investigation of membrane proteins by single-channel recording. Alternatively, lipid bilayers can be used as biosensors to detect the presence of a range of substances. The lipid bilayer may be any lipid bilayer. Suitable lipid bilayers include, but are not limited to, a planar lipid bilayer, a supported bilayer, or a liposome. The lipid bilayer is preferably a planar lipid bilayer. Suitable lipid bilayers are disclosed in WO 2008 / 102121, \N0 2009 / 077734, and WO 2006 / 100484 (incorporated herein by reference in their entireties).

[0282] Methods for forming lipid bilayers are known in the art. Lipid bilayers are commonly formed by the method of Montal and Mueller (Proc. Natl. Acad. Sci. USA., 1972; 69: 3561-3566).

[0283] A lipid bilayer may be formed as described in WO 2009 / 077734 (incorporated herein by reference in its entirety). In this method, the lipid bilayer is formed from dried lipids. A lipid bilayer may be formed across an opening as described in W02009 / 077734.

[0284] The membrane may comprise a solid-state layer. Solid state layers can be formed from both organic and inorganic materials including, but not limited to, microelectronic materials, insulating materials such as Si3N4, A12O3, and SiO, organic and inorganic polymers such as polyamide, plastics such as Teflon® or elastomers such as two-component addition-cure silicone rubber, and glasses. The solid-state layer may be formed from graphene. Suitable graphene layers are disclosed in WO 2009 / 035647 (incorporated herein by reference in its entirety). If the membrane comprises a solid-state layer, the pore is typically present in an amphiphilic membrane or layer contained within the solid-state layer, for instance within a hole, well, gap, channel, trench or slit within the solid-state layer. The skilled person can prepare suitable solid state / amphiphilic hybrid systems. Suitable systems are disclosed in WO 2009 / 020682 and WO 2012 / 005857 (incorporated herein by reference in their entireties). Any of the amphiphilic membranes or layers discussed above may be used.

[0285] The methods disclosed herein are typically carried out using (i) an artificial amphiphilic layer comprising a pore, (ii) an isolated, naturally occurring lipid bilayer comprising a pore, or (iii) a cell having a pore inserted therein. The methods are typically carried out using an artificial amphiphilic layer, such as an artificial triblock copolymer layer. The layer may comprise other transmembrane and / or intramembrane proteins as well as other molecules in addition to the pore. Suitable apparatus and conditions are discussed below. The method of the invention is typically carried out in vitro.

[0286] Polynucleotide binding protein

[0287] As those skilled in the art will appreciate, any suitable polynucleotide binding protein can be used in the adaptors, kits and methods of invention. The polynucleotide binding protein may be any protein that is capable of binding to a polynucleotide and controlling its movement with respect to a detector, e.g., a nanopore. The polynucleotide binding protein is preferably any protein that is capable of binding to RNA and controlling its movement with respect to a detector, e.g., a nanopore.

[0288] In more detail, polynucleotide binding proteins such as helicases can typically control the movement of polynucleotides in at least two active modes of operation (when is provided with all the necessary components to facilitate movement e.g., ATP and Mg2+) and one inactive mode of operation (when not provided with the necessary components to facilitate movement; or when the polynucleotide binding protein is modified in order to prevent the active mode).

[0289] When provided with all the necessary components to facilitate movement, a polynucleotide binding protein may move along a polynucleotide in either a 5'-3' direction or a 3'-5' direction. Many polynucleotide binding proteins process polynucleotides in a 5'-3' direction. Polynucleotide binding proteins which control the movement of polynucleotides in this manner are typically suitable for use in the method of the invention.

[0290] However, when a polynucleotide binding protein is not provided with the necessary components to facilitate movement or is modified in order to prevent it from actively controlling the movement of the polynucleotide with respect to the nanopore, it can still passively control the movement of the polynucleotide with respect to the nanopore. For example, the polynucleotide binding protein can bind to the polynucleotide and act as a brake slowing the movement of the polynucleotide when it is pulled into the pore by an applied field (e.g., by the first force in the method of the invention). In the "inactive" mode it typically does not matter whether the polynucleotide is captured either 3' or 5' down ( / .e., moves through the nanopore in a 5'-3' direction or in a 3'-5' direction), as the applied force provides the impetus to move the polynucleotide through the nanopore. However, in such embodiments, the polynucleotide binding protein may still control the movement of the polynucleotide with respect to the nanopore e.g., by acting as a brake. When in the inactive mode the movement control of a polynucleotide by a polynucleotide binding protein can be described in a number of ways including ratcheting, sliding, and braking. Typically the method of the invention do not comprise the use of a polynucleotide binding protein operating in the passive mode. However, when a polynucleotide binding protein the polynucleotide binding protein is used, it may be a polynucleotide binding protein operating in the passive mode.

[0291] Some methods of the invention may comprise use of a polynucleotide binding protein as a pausing moiety to impede the movement of the polynucleotide strand through the nanopore. The polynucleotide binding protein may be a protein which binds to polynucleotides but which does not have polynucleotide processing capacity, i.e., it is not a polynucleotide binding protein.

[0292] A polynucleotide-handling enzyme is a polypeptide that is capable of interacting with a polynucleotide. The enzyme may modify the polynucleotide by cleaving it to form individual nucleotides or shorter chains of nucleotides, such as di- or trinucleotides. The enzyme may modify the polynucleotide by orienting it or moving it to a specific position. A polynucleotide binding protein as used herein may be, or may be derived from a polynucleotide handling enzyme. A polynucleotide binding protein may be, or may be derived from a polynucleotide- handling enzyme.

[0293] The polynucleotide binding protein may be derived from a member of any of the Enzyme Classification (EC) groups 3.1.11, 3.1.13, 3.1.14, 3.1.15, 3.1.16, 3.1.21, 3.1.22, 3.1.25, 3.1.26, 3.1.27, 3.1.30 and 3.1.31.

[0294] Typically, the polynucleotide binding protein is a helicase, a polymerase, an exonuclease, a topoisomerase, or a variant thereof.

[0295] The polynucleotide binding protein may be modified to prevent the polynucleotide binding protein disengaging from the polynucleotide. Thus, the target polynucleotide preferably does not disengage from the polynucleotide binding protein.

[0296] As used herein, the term "disengaging" refers to the dissociation of the polynucleotide binding protein from the target polynucleotide. Thus, a polynucleotide binding protein may be modified to prevent it from dissociating from the target polynucleotide, e.g., into the reaction medium. It is important to distinguish potential "disengagement" of a polynucleotide binding protein from "unbinding" of a polynucleotide binding protein from a target polynucleotide. As used herein, "unbinding" refers to the transient release of the target polynucleotide the active site of the polynucleotide binding protein (described in more detail herein) but does not imply disengagement. Thus, for example, a polynucleotide binding protein may be modified to prevent the polynucleotide binding protein from disengaging from a polynucleotide, but without preventing the polynucleotide binding protein from unbinding from the polynucleotide. When unbound, the polynucleotide binding protein remains engaged with the target polynucleotide. For example, the polynucleotide binding protein may remain engaged with the target polynucleotide ( / .e., it may be prevented from disengaging from the target polynucleotide) because it is topologically closed around the target polynucleotide. The polynucleotide binding site may remain free to bind or unbind the target polynucleotide such that the polynucleotide binding protein may bind or unbind to the target polynucleotide, whilst the polynucleotide binding protein remains engaged with the target polynucleotide. When the polynucleotide binding protein is unbound from the target polynucleotide it may be able to move on (e.g., along) the target polynucleotide under an applied force and may be capable of re-binding to the target polynucleotide. When engaged on the target polynucleotide but unbound from the target polynucleotide, the polynucleotide binding protein is not capable of dissociating from the target polynucleotide.

[0297] The polynucleotide binding protein can be adapted to prevent disengagement in any suitable way. For example, the polynucleotide binding protein can be loaded on the polynucleotide and then modified in order to prevent it from disengaging from the polynucleotide. Alternatively, the polynucleotide binding protein can be modified to prevent it from disengaging from the polynucleotide before it is loaded onto the polynucleotide. Modification of a polynucleotide binding protein and / or a polynucleotide binding protein in order to prevent it from disengaging from a polynucleotide can be achieved using methods known in the art, such as those discussed in WO 2014 / 013260, which is hereby incorporated by reference in its entirety, and with particular reference to passages describing the modification of polynucleotide binding proteins such as helicases in order to prevent them from disengaging with polynucleotide strands. For example, a polynucleotide binding protein can be modified by treating with tetramethylazodicarboxamide (TMAD). Various other closing moieties are described in WO 2021 / 255476 (incorporated herein by reference in its entirety).

[0298] For example, a polynucleotide binding protein and / or a polynucleotide binding protein may have a polynucleotide-unbinding opening, e.g., a cavity, cleft or void through which a polynucleotide strand may pass when the polynucleotide binding protein disengages from the strand. The polynucleotide-unbinding opening may be the opening through which a polynucleotide may pass when the polynucleotide binding protein disengages from the polynucleotide. The polynucleotide-unbinding opening for a given polynucleotide binding protein can be determined by reference to its structure, e.g., by reference to its X-ray crystal structure. The X-ray crystal structure may be obtained in the presence and / or the absence of a polynucleotide substrate. The location of a polynucleotide-unbinding opening in a given polynucleotide binding protein may be deduced or confirmed by molecular modelling using standard packages known in the art. The polynucleotide-unbinding opening may be transiently produced by movement of one or more parts e.g., one or more domains of the polynucleotide binding protein.

[0299] The polynucleotide binding protein may be modified by closing the polynucleotide-unbinding opening. The polynucleotide-unbinding opening may be closed with a closing moiety.

[0300] Closing the polynucleotide-unbinding opening may therefore prevent the polynucleotide binding protein from disengaging from the polynucleotide. For example, the polynucleotide binding protein may be modified by covalently closing the polynucleotide-unbinding opening. However, as explained above closing the polynucleotide-unbinding opening does not necessarily prevent the target polynucleotide from unbinding from the polynucleotide binding site of the polynucleotide binding protein. A preferred protein for addressing in this way is a helicase.

[0301] The polynucleotide binding protein may be modified with a closing moiety for (i) topologically closing the polynucleotide binding site of the polynucleotide binding protein around the target polynucleotide and (ii) promoting unbinding of the target polynucleotide from the polynucleotide binding site of the polynucleotide binding protein and / or retarding re-binding of the target polynucleotide to the polynucleotide binding site of the polynucleotide binding protein. The polynucleotide binding protein may be modified in any suitable manner to facilitate attachment of such a closing moiety.

[0302] A closing moiety may comprise a bifunctional cross-linking moiety. The closing moiety may comprise a bifunctional cross-linker. The bifunctional crosslinker may attach at two points on the polynucleotide binding protein and close the polynucleotide-unbinding opening of the polynucleotide binding protein thereby preventing disengagement of the polynucleotide from the polynucleotide binding protein whilst allowing unbinding of the polynucleotide from the polynucleotide-binding site of the polynucleotide binding protein.

[0303] The closing moiety may attach at any suitable positions on the polynucleotide binding protein. For example, the closing moiety may crosslink two amino acid residues of the polynucleotide binding protein. Typically, at least one amino acid crosslinked by the closing moiety is a cysteine or a non-natural amino acid. The cysteine or non-natural amino acid may be introduced into the polynucleotide binding protein by substitution or modification of a naturally occurring amino acid residue of the polynucleotide binding protein. Methods for introducing non-natural amino acids are well known in the art and include for example native chemical ligation with synthetic polypeptide strands comprising such non-natural amino acids. Methods for introducing cysteines into a polynucleotide binding protein are likewise within the capability of one of skill in the art, for example using techniques disclosed in references such as Sambrook et al., Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Press, Plainsview, New York (2012); and Ausubel et al., Current Protocols in Molecular Biology (Supplement 114), John Wiley & Sons, New York (2016).

[0304] The closing moiety may have a length of from about 1 A to about 100 A. The length of the closing moiety may be calculated according to static bond lengths or more preferably using molecular dynamics simulations. The length may for example be from about 2 A to about 80 A, such as from about 5 A to about 50 A, e.g., from about 8 to about 30 A such as from about 10 to about 25 A or about 20 A, e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or 19 A.

[0305] Polynucleotide binding proteins suitable for being closed using a closing moiety as described above are discussed in more detail herein. The polynucleotide binding protein is preferably a helicase.

[0306] The polynucleotide binding protein may be or may be derived from an exonuclease. Suitable enzymes include, but are not limited to, exonuclease I from E. coli, exonuclease III enzyme from E. coli, RecJ from T. thermophilus and bacteriophage lambda exonuclease, TatD exonuclease and variants thereof.

[0307] The polynucleotide binding protein may be a polymerase. The polymerase may be PyroPhage® 3173 DNA Polymerase (which is commercially available from Lucigen® Corporation), SD Polymerase (commercially available from Bioron®), Klenow from NEB or variants thereof. In one embodiment, the enzyme is Phi29 DNA polymerase or a variant thereof. Modified versions of Phi29 polymerase that may be used in the invention are disclosed in US Patent No. 5,576,204.

[0308] The polynucleotide binding protein may be a topoisomerase. In one embodiment, the topoisomerase is a member of any of the Moiety Classification (EC) groups 5.99.1.2 and 5.99.1.3. The topoisomerase may be a reverse transcriptase, which are enzymes capable of catalysing the formation of cDNA from a RNA template. They are commercially available from, for instance, New England Biolabs® and Invitrogen®.

[0309] The polynucleotide binding protein is preferably a helicase. Any suitable helicase can be used in accordance with the method of the invention. The helicase is preferably a member of superfamily 1 or superfamily 2. The helicase is more preferably a member of one of the following families: Pifl-like, Upfl-like, UvrD / Rep, Ski-like, Rad3 / XPD, NS3 / NPH-II, DEAD, DEAH / RHA, RecG-like, REcQ-like, TIR-like, Swi / Snf-like and Rig-I-like. The first three of those families are in superfamily 1 and the second ten families are in superfamily 2. The helicase is more preferably a member of one of the following subfamilies: RecD, Upfl , PcrA, Rep, UvrD, Hel308, Mtr4, XPD, NS3, Mssll6, Prp43, RecG, RecQ, T1R, RapA and Hef. The first five of those subfamilies are in superfamily 1 and the second eleven subfamilies are in superfamily 2. Members of the Upfl, Mtr4, NS3, Mssll6, Prp43 and Hef subfamilies are RNA helicases. The polynucleotide is preferably a member of one of the Upfl, Mtr4, NS3, Mssll6, Prp43 and Hef subfamilies. Members of the remaining subfamilies are DNA helicases.

[0310] The polynucleotide binding protein is preferably a NS3 helicase or a modified NS3 helicase. The NS3 helicase or modified NS3 helicase may be any of those described in PCT / EP2024 / 063118 (incorporated herein by reference in its entirety).

[0311] For example, the or each enzyme used in accordance with the present disclosure may be independently selected from a Hel308 helicase, a RecD helicase, a Tral helicase, a TrwC helicase, an XPD helicase, and a Dda helicase, or a variant thereof. Monomeric helicases may comprise several domains attached together. For instance, Tral helicases and Tral subgroup helicases may contain two RecD helicase domains, a relaxase domain and a C- terminal domain. The domains typically form a monomeric helicase that is capable of functioning without forming oligomers. Particular examples of suitable helicases include Hel308, NS3, Dda, UvrD, Rep, PcrA, Pifl and Tral. These helicases typically work on single stranded DNA. Examples of helicases that can move along both strands of a double stranded DNA include FtfK and hexameric enzyme complexes, or multisubunit complexes such as RecBCD. The polynucleotide binding protein may be a Dda (DNA-dependent ATPase) helicase. Hel308 helicases are described in publications such as WO 2013 / 057495, the entire contents of which are incorporated by reference. RecD helicases are described in publications such as WO 2013 / 098562, the entire contents of which are incorporated by reference. XPD helicases are described in publications such as WO 2013 / 098561, the entire contents of which are incorporated by reference. Dda helicases are described in publications such as WO 2015 / 055981 and WO 2016 / 055777, the entire contents of each of which are incorporated by reference.

[0312] The helicase may be Trwc Cba or a variant thereof, Hel308 Mbu or a variant thereof or Dda or a variant thereof. Variants may differ from the native sequences in any of the ways discussed herein.

[0313] Diagnosing or prognosinq diseases or conditions

[0314] RIMA molecule(s) characterised or sequenced in accordance with the invention may be used to diagnose or prognose a disease or condition. Some diseases or conditions are associated with an altered amount (or level) of RNA molecule(s), such as mRNA. The mRNA may be normal or wild-type mRNA, i.e. not alternately spliced. The amount (or level) of the RNA molecule(s), such as mRNA, may be increased or decreased in the disease or condition compared with the amount (or level) in a patient without the disease or condition. Such diseases or conditions may be diagnosed or prognosed by determining the amount of the RNA molecule(s), such as mRNA, in a sample from the patient using a method of the invention.

[0315] Many genetic diseases or conditions are caused by mutations that cause alternate mRNA splicing, such as mRNA splicing defects. A number of diseases or conditions are associated with alternate mRNA splicing which are not attributed to overt mutations. The presence or absence of alternate splicing can be identified by determining the presence or absence of an alternately spliced mRNA in a sample from the patient using the method of the invention. In some instances, alternate mRNA splicing may be the normal function of a cell. In such instances, an increased or decreased amount (or level) of the alternately spliced mRNA compared with the normal amount i.e. the amount in a patient without the disease or condition) may be used to diagnose or prognose the disease or condition.

[0316] The invention provides a method of diagnosing or prognosing a disease or condition associated with an altered amount and / or alternate splicing of messenger RNA (mRNA) in a patient. The invention provides a method of determining whether or not a patient has or is at risk of developing a disease or condition associated with an altered amount and / or alternate splicing of messenger RNA (mRNA). In each instance, the method comprises determining the amount and / or identity of the mRNA in a sample from the patient using a method of characterising a RNA molecule or a plurality of RNA molecules of the invention. The RNA molecule is mRNA or the plurality of RNA molecules are a plurality of mRNAs. The mRNA(s) are uniquely labelled in accordance with the invention and then characterised in accordance with the invention.

[0317] The disease or condition may be any of those discussed below. The disease or condition is preferably cystic fibrosis, familial dysautonomia, frontotemporal lobar dementia, amyotrophic lateral sclerosis, Hutchinson-Gilford progeria syndrome, medium-chain acyl- CoA dehydrogenase (MCAD) deficiency, myotonic dystrophy, Prader-Willi syndrome, spinal muscular atrophy, tauopathy, hypercholesterolemia or cancer. These diseases, their causes and possible treatments are discussed in Tazi et al. (Biochimica et Biophysica Acta (BBA) - Molecular Basis of Disease, Volume 1792, Issue 1, January 2009, Pages 14-26).

[0318] The presence of an altered ( / .e. increased or decreased) amount (or level) of the mRNA in the sample from the patient typically diagnoses or prognoses the disease or condition, i.e. indicates that the patient has or is at risk of developing the disease or condition. The absence of an altered i.e. increased or decreased) amount (or level) of the mRNA in the sample from the patient typically indicates that the patient does not have or is not at risk of developing the disease or condition.

[0319] The presence of the alternately spliced mRNA in the sample from the patient typically diagnoses or prognoses the disease or condition, i.e. indicates that the patient has or is at risk of developing the disease or condition. The absence of the alternately spliced mRNA in the sample from the patient typically indicates that the patient does not have or is not at risk of developing the disease or condition. The presence or absence of the alternately spliced mRNA can be determined by identifying RNA in the sample as discussed above.

[0320] An increased or decreased amount (or level) of the alternately spliced mRNA in the sample from the patient typically diagnoses or prognoses the disease or condition, i.e. indicates that the patient has or is at risk of developing the disease or condition. No change in the amount of the alternately spliced mRNA in the sample from the patient (compared with the amount or level in a patient without the disease or condition) typically indicates that the patient does not have or is not at risk of developing the disease or condition. The amount of the alternately spliced mRNA can be determined as discussed above. miRNA is preferably used in the invention to diagnose or prognose a disease or condition. The invention provides a method of diagnosing or prognosing a disease or condition associated with a miRNA. The invention provides a method of determining whether or not a patient has or is at risk of developing a disease or condition associated with a miRNA. The method comprises determining the presence or absence of the miRNA in a sample from the patient using a method of the invention. The disease or condition may be any of those discussed below. The presence of the miRNA in the sample from the patient typically indicates that the patient has or is at risk of developing the disease or condition. The absence of the miRNA in the sample from the patient typically indicates that the patient does not have or is not at risk of developing the disease or condition. The presence or absence of the miRNA can be determined by identifying any miRNAs in the sample as discussed above.

[0321] The disease or condition is preferably cancer, coronary heart disease, cardiovascular disease or sepsis. The disease or condition is more preferably abdominal aortic aneurysm, acute lymphoblastic leukemia (ALL), acute myeloid leukemia (AML), acute myocardial infarction, acute promyelocytic leukemia (APL), adenoma, adrenocortical carcinoma, alcoholic liver disease, Alzheimer's disease, anaplastic thyroid carcinoma (ATC), anxiety disorder, asthma, astrocytoma, atopic dermatitis, autism spectrum disorder (ASD), B-cell chronic lymphocytic leukemia, B-cell lymphoma, Becker muscular dystrophy (BMD), bladder cancer, brain neoplasm, breast cancer, Burkitt lymphoma, cardiac hypertrophy, cardiomyopathy, cardiovascular disease, cerebellar neurodegeneration, cervical cancer, cholangiocarcinoma, cholesteatoma, choriocarcinoma, chronic lymphocytic leukemia, chronic myeloid leukemia, chronic pancreatitis, colon carcinoma, colorectal cancer, congenital heart disease, coronary artery disease, cowden syndrome, dermatomyositis (DM), diabetic nephropathy, diarrhea predominant irritable bowel syndrome, diffuse large B-cell lymphoma, dilated cardiomyopathy, down syndrome (DS), duchenne muscular dystrophy (DMD), endometrial cancer, endometrial endometrioid adenocarcinoma, endometriosis, epithelial ovarian cancer, esophageal cancer, esophagus squamous cell carcinoma, essential thrombocythemia (ET), facioscapulohumeral muscular dystrophy (FSHD), follicular lymphoma (FL), follicular thyroid carcinoma (FTC), frontotemporal dementia, gastric cancer (stomach cancer), glioblastoma, glioblastoma multiforme (GBM), glioma, glomerular disease, glomerulosclerosis, hamartoma, HBV-related cirrhosis, HCV infection, head and neck cancer, head and neck squamous cell carcinoma (HNSCC), hearing loss, heart disease, heart failure, hepatitis B, hepatitis C, hepatocellular carcinoma (HCC), hilar cholangiocarcinoma, Hodgkin's lymphoma, homozygous sickle cell disease (HbSS), Huntington's disease (HD), hypertension, hypopharyngeal cancer, inclusion body myositis (IBM), insulinoma, intrahepatic cholangiocarcinoma (ICC), kidney cancer, kidney disease, laryngeal carcinoma, late insomnia (sleep disease), leiomyoma of lung, leukemia, limb-girdle muscular dystrophies types 2A (LGMD2A), lipoma, lung adenocarcinoma, lung cancer, lymphoproliferative disease, malignant lymphoma, malignant melanoma, malignant mesothelioma (MM), mantle cell lymphoma (MCL), medulloblastoma, melanoma, meningioma, metabolic disease, miyoshi myopathy (MM), multiple myeloma (MM), multiple sclerosis, MYC-rearranged lymphoma, myelodysplastic syndrome, myeloproliferative disorder, myocardial infarction, myocardial injury, myoma, nasopharyngeal carcinoma (NPC), nemaline myopathy (NM), nephritis, neuroblastoma (NB), neutrophilia, Niemann-Pick type C (NPC) disease, non-alcoholic fatty liver disease (NAFLD), non-small cell lung cancer (NSCLC), obesity, oral carcinomaosteosarcoma ovarian cancer (OC), pancreatic cancer, pancreatic ductal adenocarcinoma (PDAC), pancreatic neoplasia, panic disease, papillary thyroid carcinoma (PTC), Parkinson's disease, PFV-1 infection, pharyngeal disease, pituitary adenoma, polycystic kidney disease, polycystic liver disease, polycythemia vera (PV), polymyositis (PM), primary biliary cirrhosis (PBC), primary myelofibrosis, prion disease, prostate cancer, psoriasic arthritis, psoriasis, pulmonary hypertension, recurrent ovarian cancer, renal cell carcinoma, renal clear cell carcinoma, retinitis pigmentosa (RP), retinoblastoma, rhabdomyosarcoma, rheumatic heart disease and atrial fibrillation, rheumatoid arthritis, sarcoma, schizophrenia, sepsis, serous ovarian cancer, Sezary syndrome, skin disease, small cell lung cancer, spinocerebellar ataxia, squamous carcinoma, T-cell leukemia, teratocarcinoma, testicular germ cell tumor, thalassemia, thyroid cancer, tongue squamous cell carcinoma, tourette's syndrome, type 2 diabetes, ulcerative colitis (UC), uterine leiomyoma (ULM), uveal melanoma, vascular disease, vesicular stomatitis or Waldenstrom macroglobulinemia (WM).

[0322] The patient may be any of the mammals discussed above. The patient is preferably human. The patient is an individual.

[0323] The sample may be any of those discussed above. The sample is typically from any tissue or bodily fluid. The sample typically comprises a body fluid and / or cells of the patient and may, for example, be obtained using a swab, such as a mouth swab. The sample may be, or be derived from, blood, urine, saliva, skin, cheek cell or hair root samples. The RNA molecule(s) is / are typically extracted from the sample before it is used in the method of the invention.

[0324] The method may concern diagnosis of the disease or condition in the patient, i.e. determining whether or not the patient has the disease or condition. The patient may be symptomatic.

[0325] The method may concern prognosing the disease or condition in the patient, i.e. determining whether or not the patient is likely to develop the disease or condition. The patient can be asymptomatic. The patient can have a genetic predisposition to the disease or condition. The patient may have one or more family member(s) with the disease or condition.

[0326] General methods

[0327] As mentioned above, the method of the invention may be operated using any suitable detector, and as such any suitable apparatus for detecting polynucleotides can be used.

[0328] The method of the invention may be carried out using any apparatus that is suitable for nanopore sensing. For example, the apparatus may comprise a chamber comprising an aqueous solution and a barrier that separates the chamber into two sections. The barrier may have an aperture in which a membrane containing a transmembrane pore is formed. Transmembrane pores are described herein.

[0329] The methods may be carried out using the apparatus described in WO 2008 / 102120, WO 2010 / 122293, or WO 00 / 28312 (incorporated herein by reference in their entireties). In brief, the binding of the uniquely labelled RIMA molecule(s) in the channel of a pore will have an effect on the open-channel ion flow through the pore, which is the essence of "molecular sensing" of pore channels. Variation in the open-channel ion flow can be measured using suitable measurement techniques by the change in electrical current. The degree of reduction in ion flow, as measured by the reduction in electrical current, is related to the size of the obstruction within, or in the vicinity of, the pore. Binding of the uniquely labelled RNA molecule(s) in or near the pore therefore provides a detectable and measurable event, thereby forming the basis of a "biological sensor".

[0330] When used to characterize the uniquely labelled RNA molecule(s), the presence, absence or one or more characteristics of the the uniquely labelled RNA molecule(s) are determined. The methods may be for determining the presence, absence or one or more characteristics of the uniquely labelled RNA molecule(s). The methods may concern determining the presence, absence or one or more characteristics of two or more the uniquely labelled RNA molecules. The methods may comprise determining the presence, absence or one or more characteristics of any number of uniquely labelled RNA molecules, such as 2, 5, 10, 15, 20, 30, 40, 50, 100 or more the uniquely labelled RNA molecules. Any number of characteristics of the one or more target polynucleotides may be determined, such as 1, 2, 3, 4, 5, 10 or more characteristics. Characteristics amenable to being detected in the methods provide herein include the identity or sequence of the uniquely labelled RNA molecule(s), the length of the uniquely labelled RNA molecule(s), whether or not the uniquely labelled RNA molecule(s) is / are modified, etc. In some embodiments, the method of the invention is a method of sequencing the uniquely labelled RNA molecule(s). In some embodiments the sequences of the uniquely labelled RNA molecule(s) may be determined in real-time by aligning real-time signal or basecalling to known references. Exemplary methods of determining a polynucleotide sequence are described in WO 2016 / 059427 (incorporated herein by reference in its entirety).

[0331] The method may involve measuring the ion current flow through the pore, typically by measurement of a current. Alternatively, the ion flow through the pore may be measured optically, such as disclosed by Heron et al: J. Am. Chem. Soc. 9 Vol. 131, No. 5, 2009. Therefore, the apparatus may also comprise an electrical circuit capable of applying a potential and measuring an electrical signal across the membrane and pore. The characterisation methods may be carried out using a patch clamp or a voltage clamp. The characterisation methods preferably involve the use of a voltage clamp. The method may involve measuring an optical signal as described in Chen et al, Nature Communications (2018)9: 1733, the entire contents of which are hereby incorporated by reference. For example, a nanopore such as an optically engineered nanopore structure (e.g., a plasmonic nanoslit) may be used to locally enable single-molecule surface enhanced Raman spectroscopy (SERS) to allow the characterisation of the polynucleotide through direct Raman spectroscopic detection.

[0332] The method may be carried out on a silicon-based array of wells where each array comprises 128, 256, 512, 1024, 2000, 3000, 4000, 6000, 10000, 12000, 15000 or more wells.

[0333] The method may involve the measuring of a current flowing through the pore. The method is typically carried out with a voltage applied across the membrane and pore. The voltage used is typically from +2 V to -2 V, typically -400 mV to +400mV. The voltage used is preferably in a range having a lower limit selected from -400 mV, -300 mV, -200 mV, -150 mV, -100 mV, -50 mV, -20mV and 0 mV and an upper limit independently selected from + 10 mV, + 20 mV, +50 mV, +100 mV, +150 mV, +200 mV, +300 mV and +400 mV. The voltage used is more preferably in the range 100 mV to 240mV and most preferably in the range of 120 mV to 220 mV. It is possible to increase discrimination between different nucleotides by a pore by using an increased applied potential.

[0334] The methods are typically carried out in the presence of any charge carriers, such as metal salts, for example alkali metal salts, halide salts, for example chloride salts, such as alkali metal chloride salt. Charge carriers may include ionic liquids or organic salts, for example tetramethyl ammonium chloride, trimethylphenyl ammonium chloride, phenyltrimethyl ammonium chloride, or l-ethyl-3-methyl imidazolium chloride. In the exemplary apparatus discussed above, the salt is present in the aqueous solution in the chamber. Potassium chloride (KCI), sodium chloride (NaCI) or caesium chloride (CsCI) is typically used. KCI is preferred. The salt may be an alkaline earth metal salt such as calcium chloride (CaCI2).

[0335] The salt concentration may be at saturation. The salt concentration may be 3M or lower and is typically from 0.1 to 2.5 M, from 0.3 to 1.9 M, from 0.5 to 1.8 M, from 0.7 to 1.7 M, from 0.9 to 1.6 M or from 1 M to 1.4 M. The salt concentration is preferably from 150 mM to 1 M. The method is preferably carried out using a salt concentration of at least 0.3 M, such as at least 0.4 M, at least 0.5 M, at least 0.6 M, at least 0.8 M, at least 1.0 M, at least 1.5 M, at least 2.0 M, at least 2.5 M or at least 3.0 M. High salt concentrations provide a high signal to noise ratio and allow for currents indicative of binding / no binding to be identified against the background of normal current fluctuations.

[0336] The methods are typically carried out in the presence of a buffer. In the exemplary apparatus discussed above, the buffer is present in the aqueous solution in the chamber. Any suitable buffer may be used. Typically, the buffer is HEPES. Another suitable buffer is Tris-HCI buffer. The methods are typically carried out at a pH of from 4.0 to 12.0, from 4.5 to 10.0, from 5.0 to 9.0, from 5.5 to 8.8, from 6.0 to 8.7 or from 7.0 to 8.8 or 7.5 to 8.5. The pH used is preferably about 7.5.

[0337] The methods may be carried out at from 0 °C to 100 °C, from 15 °C to 95 °C, from 16 °C to 90 °C, from 17 °C to 85 °C, from 18 °C to 80 °C, 19 °C to 70 °C, or from 20 °C to 60 °C. The methods are typically carried out at room temperature. The methods are optionally carried out at a temperature that supports enzyme function, such as about 37 °C.

[0338] Any of the proteins described herein, such as the protein pores, may be made synthetically or by recombinant means. For example, the pore may be synthesised by in vitro translation and transcription (IVTT). The amino acid sequence of the pore may be modified to include non-naturally occurring amino acids or to increase the stability of the protein. When a protein is produced by synthetic means, such amino acids may be introduced during production. The pore may also be altered following either synthetic or recombinant production.

[0339] Any of the proteins described herein, such as the protein pores, can be produced using standard methods known in the art. Polynucleotide sequences encoding a pore or construct may be derived and replicated using standard methods in the art. Polynucleotide sequences encoding a pore or construct may be expressed in a bacterial host cell using standard techniques in the art. The pore may be produced in a cell by in situ expression of the polypeptide from a recombinant expression vector. The expression vector optionally carries an inducible promoter to control the expression of the polypeptide. These methods are described in Sambrook, J. and Russell, D. (2001). Molecular Cloning: A Laboratory Manual, 3rd Edition. Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY.

[0340] The pore may be produced in large scale following purification by any protein liquid chromatography system from protein producing organisms or after recombinant expression. Typical protein liquid chromatography systems include FPLC, AKTA systems, the Bio-Cad system, the Bio-Rad BioLogic system, and the Gilson HPLC system.

[0341] Systems

[0342] The invention also provides a system for conducting the method of the invention. The system is for uniquely labelling and characterising a RNA molecule or a plurality of RNA molecules.

[0343] In one embodiment, the system comprises (a) a double stranded polynucleotide adaptor of the invention and (b) a nanopore. The system preferably further comprises a polynucleotide sequencing adaptor of the invention. In another embodiment, the system comprises (a) a population of double stranded polynucleotide adaptors of the invention and (b) a nanopore. The system preferably further comprises a polynucleotide sequencing adaptor or a population of sequencing adaptors of the invention.

[0344] Any of the embodiments discussed above with reference to the adaptors of the invention, kits of the invention and method of the invention equally apply to the system of the invention.

[0345] The nanopore is preferably present in a membrane. Suitable membranes are discussed above. The system may comprise any of the membranes disclosed above, such as an amphiphilic layer, a triblock copolymer membrane or a solid-state layer. The membrane is typically part of an array of membranes, wherein each membrane preferably comprises a nanopore. The array may be any of those described in WO 2018 / 060740 (incorporated herein by reference in its entirety).

[0346] The system is preferably adapted to apply a voltage across the membrane and to take one or more electrical measurements. Suitable adaptations are discussed in WO 2018 / 060740 (incorporated herein by reference in its entirety).

[0347] SEQUENCE LISTING

[0348] Table 1 - Preferred double stranded polynucleotide adaptors of the invention. The two sequences in each row hybridise to form an adaptor. The barcodes are underlined. / 5Phos / = 5' Phosphate. dU = deoxyuridine.

[0349] SEQ ID NO: 49 (RTB01_BOT_tRNA; all RNA)

[0350] UACAUUCGAAGUGUCCGAAGAAGCCUGGN

[0351] SEQ ID NO: 50 (RTB01_BOT_PRIMER; all DNA) CGATGCTGGTTACAAGAT SEQ ID NO: 51 (RNA Ligation Sequencing Adaptor Top Strand) where / 5Phos / is 5' Phosphate, RNA is underlines, X is a spacer, Y is a polynucleotide binding protein binding site and Z is a leader. The two sequences in SEQ ID NO: 51 are shown as SEQ ID NOs: 61 and 62 in the ST.26 sequence listing.

[0352] / 5Phos / GGUUGUUUCUGUUGGUGCUGAUAUUGC / X / Y / ATGATGCAAGATACGCAC / Z /

[0353] SEQ ID NO: 52 (RNA Ligation Sequencing Adaptor Bottom Strand)

[0354] CAATTTGCAATATCAGCACCAACAGAAACAACCACGAACCCACAAATTGG

[0355] SEQ ID NO: 53 (Blocker Strand)

[0356] TTGTGCGTATCTTGCATCATA

[0357] SEQ ID NO: 54 (100mer_RNA)

[0358] AUACUCGACAUAGAUAGGACUCUUUAGCUAGUGAACCCUAGCCUCCGGAGACAGGUCGCGACCU G UG U AG AU G AG AG AAC UG AG U GC AC AAAAAAAAAAA

[0359] SEQ ID NOs: 55-60: tRNA adapters

[0360] SEQ ID NOs: 61 and 62 form part of SEQ ID NO: 51 as explained above.

[0361] The following Examples illustrate the invention. It is to be understood that although particular embodiments, specific configurations as well as materials and / or molecules, have been discussed herein for methods according to the invention, various changes or modifications in form and detail may be made without departing from the scope and spirit of this invention. The following examples are provided to better illustrate particular embodiments, and they should not be considered limiting the application. The application is limited only by the claims.

[0362] EXAMPLES

[0363] Example 1 This Example describes the preparation of RIMA Barcoding adaptor, ligation of the RNA barcoding adaptor to RNA molecules and ligation of the RNA barcoding adaptor to sequencing adaptor.

[0364] Oligonucleotides were reconstituted from lyophilised form in TE buffer (10 mM Tris-HCI pH 7.5, 0.1 mM EDTA). The reconstituted oligonucleotides in Table 1 were annealed in a pairwise fashion such that the RNA strand and its complementary DNA strand were mixed together (e.g. SEQ ID NO 1 & 2). The individual mixtures were heated and slow cooled to anneal the strands. These were the RNA Barcoding Adaptors, shown schematically in Figure 2. To assess annealing efficiency, the RNA Barcoding Adaptors were analysed by HPLC. OpenLAB CDS ChemStation Edition was used for automated data analysis and peak integration. Figure 3 presents chromatographic traces that show the typical profile of an RNA barcoding adaptor and its constituent oligonucleotides [1 and 2]. The shift in retention time for the annealed RNA Barcoding adaptor [3] demonstrates that the constituent oligonucleotides are annealed.

[0365] A ligation reaction was set up using the RNA Barcoding adaptor, synthetic RNA strand (SEQ ID NO 54) and a ligase in a reaction buffer. The ligation reaction was analysed by denaturing polyacrylamide gel electrophoresis (PAGE), results are shown in Figure 4. Figure 4 presents a denaturing PAGE gel image, showing successful annealing of the DNA constituent strand of the RNA Barcoding adaptor [1] and ligation [higher molecular weight species 4] of the RNA constituent strand of the RNA barcoding adaptor [2] to an RNA strand [3].

[0366] A sequencing adaptor is prepared as described in Example 3 and shown schematically in Figure 1. A ligation reaction was set up using the RNA Barcoding adaptor, sequencing adaptor and a ligase in a reaction buffer. The ligation reaction was analysed on a denaturing PAGE gel, results are shown in Figure 5. Figure 5 presents a denaturing PAGE gel image, showing successful annealing of the DNA constituent strand of the RNA Barcoding adaptor [3] and ligation [higher molecular weight species 4] of the RNA constituent strand of the RNA barcoding adaptor [3] to the sequencing adaptor [1].

[0367] Example 2

[0368] This Example describes the preparation of tRNA Barcoding adaptor for uniquely labelling tRNA molecules, ligation of the tRNA RNA barcoding adaptor to tRNA molecules and ligation of the tRNA barcoding adaptor to sequencing adaptor.

[0369] Oligonucleotides are reconstituted from lyophilised form in TE buffer (10 mM Tris-HCI pH 7.5, 0.1 mM EDTA). The reconstituted oligonucleotides (SEQ ID NO 1, SEQ ID NO 49 and SEQ ID NO 50) are annealed in a pairwise fashion such that the RNA strand and its complementary DNA strand are mixed together. The individual mixtures are heated and slow cooled to anneal the strands. These are the tRNA Barcoding Adaptors. To assess annealing efficiency, the tRNA Barcoding Adaptors are analysed by HPLC. OpenLAB CDS ChemStation Edition is used for automated data analysis and peak integration.

[0370] A ligation reaction is prepared using the tRNA Barcoding adaptor, tRNA molecules and a ligase in a reaction buffer. The ligation reaction is analysed by denaturing PAGE.

[0371] A ligation reaction is prepared using the tRNA Barcoding adaptor, sequencing adaptor and a ligase in a reaction buffer. The ligation reaction is analysed on a denaturing PAGE gel.

[0372] Example 3

[0373] The sequencing adaptor assembly method is described in PCT / EP2024 / 063118 (incorporated herein by reference in its entirety) where the top and bottom strands were SEQ ID 51 and 52.

[0374] Example 4

[0375] This Example describes a method for labelling RNA or tRNA molecules with RNA or tRNA Barcoding adaptors. tRNA or RNA Barcoding adaptors were prepared and ligated to tRNA molecules or RNA molecules as described in examples 1 and 2. The ligation reaction was quenched using EDTA. Ligation reactions were pooled and a SPRI clean-up was performed using RNACIean AMPure XP beads, and the ligated sample was eluted in nuclease free water. A reverse transcription reaction was performed on the ligated sample as described in SQK-RNA004 Oxford Nanopore Technologies protocol. A SPRI clean-up was performed using RNACIean AMPure XP beads and the reverse transcribed sample was eluted in nuclease free water. USER (#M5508 New England Biolabs) was added to the reverse transcribed sample and incubated following manufacturer's instructions. A SPRI clean-up was performed using RNACIean AMPure XP beads. The USER treated sample was ligated to the sequencing adaptor as described in Example 1 and 2. A SPRI clean-up was performed, and the sample was eluted, prepared for sequencing and sequenced following the SQK-RNA004 Oxford Nanopore Technologies protocol.

[0376] Example 5

[0377] This Example describes a method for labelling tRNA molecules with tRNA Barcoding adaptors with the option to use an RT step or not. tRNA Barcoding adaptors were prepared and ligated to synthetic tRNA molecules as described in Example 4. The method steps of this Example are shown in Figure 7.

[0378] Each tRNA adaptor was formed using SEQ ID NO: 1 (RTB01_TOP) as the first RNA strand, one of sequences SEQ ID NOs: 55-59 and 49 as the second RNA strand and SEQ ID NO: 60 as the RT primer. The ligation reaction was quenched using EDTA. Ligation reactions were pooled and a SPRI clean-up was performed using RNACIean AMPure XP beads, and the ligated sample was eluted in nuclease free water. A reverse transcription reaction was performed on the ligated sample as described in SQK-RNA004 Oxford Nanopore Technologies protocol. Several alternative reverse transcriptase enzymes were also evaluated. A SPRI clean-up was performed using RNACIean AMPure XP beads and the reverse transcribed sample was eluted in nuclease free water. The reverse transcribed sample was ligated to the sequencing adaptor as described in Example 1 and 2. A SPRI clean-up was performed, and the sample was eluted, prepared for sequencing and sequenced following the SQK-RNA004 Oxford Nanopore Technologies protocol.

[0379] Samples were analysed with a custom pipeline consisting of the following steps, 1) basecalling of the sequencing data (POD5) with the latest RNA model (sup@v5.1.0), 2) Optional detection of 5' and 3' adapters, 3) Optional subtractive mappings to non-tRNA references such as mRNAs and rRNAs, 4) Mapping to tRNA references with a custom local alignment algorithm and 5) identification of reads spanning the tRNA reference and at least a part of the sequencing adapters in both 5' and 3'.

[0380] The results are shown in Figures 8 and 9.

Claims

1. CLAIMS1. A double stranded polynucleotide adaptor for uniquely labelling a RNA molecule, the double stranded adaptor comprising an overhang capable of hybridising to the RNA molecule and a barcode, wherein the adaptor comprises a first RNA strand comprising the barcode and a second strand comprising the overhang and a non-RNA polynucleotide.

2. A double stranded polynucleotide adaptor according to claim 1, wherein the second strand is a second DNA strand.

3. A double stranded polynucleotide adaptor according to claim 2, wherein the overhang is capable of hybridising to the 3' end of the RNA molecule.

4. A double stranded polynucleotide adaptor according to claim 2 or 3, wherein the overhang is at the 3' end of the second DNA strand.

5. A double stranded polynucleotide adaptor according to claim 4, wherein the 5' end of the second DNA strand does not hybridise to the first RNA strand.

6. A double stranded polynucleotide adaptor according to any one of claims 2-5, wherein the first RNA strand comprises a sequence capable of hybridising to a sequencing adaptor.

7. A double stranded polynucleotide adaptor according to claim 6, wherein the sequence is at the 3' end of the first RNA strand.

8. A double stranded polynucleotide adaptor according to claim 1, wherein the overhang is capable of hybridising to the tail on a transfer RNA (tRNA) molecule.

9. A double stranded polynucleotide adaptor according to claim 8, wherein the second strand comprises a second RNA strand and a reverse transcription (RT) primer.

10. A double stranded polynucleotide adaptor according to claim 9, wherein the overhang is at the 3' end of the second RNA strand.

11. A double stranded polynucleotide adaptor according to any one of claims 8-10, wherein the first RNA strand comprises a sequence capable of hybridising to a sequencing adaptor.

12. A double stranded polynucleotide adaptor according to claim 11, wherein the sequence is at the 3' end of the first RNA strand.

13. A double stranded polynucleotide adaptor according to any one of the preceding claims, wherein the barcode comprises one or more of (a) up to about 40 nucleotides in length, (b) no homopolymers longer than about 3 nucleotides, and (c) no secondary structure.

14. A population of double stranded polynucleotide adaptors for uniquely labelling a plurality of RIMA molecules, the population comprising two or more double stranded polynucleotide adaptors according to any one of the preceding claims, wherein each double stranded polynucleotide adaptor comprises a different barcode.

15. A polynucleotide sequencing adaptor comprising a RNA polynucleotide upstream 5' to 3' of a loading site for a polynucleotide binding protein.

16. A polynucleotide sequencing adaptor according to claim 15, wherein the adaptor comprises a double stranded region and a region where the two strands are not complementary.

17. A polynucleotide sequencing adaptor according to claim 16, wherein the loading site is present in one of the two strands that are not complementary.

18. A polynucleotide sequencing adaptor according to claim 17, wherein the RNA polynucleotide is part of the strand comprising the loading site and forms the double stranded region.

19. A polynucleotide sequencing adaptor according to any one of claims 15-18, wherein the adaptor further comprises a sequence capable of hybridising to a double stranded polynucleotide adaptor according to any one of claims 1-13.

20. A kit for uniquely labelling a RNA molecule, the kit comprising a double stranded polynucleotide adaptor according to any one of claims 1-13 and a polynucleotide sequencing adaptor according to any one of claims 15-19.

21. A kit for uniquely labelling a plurality of RNA molecules, the kit comprising a population of double stranded polynucleotide adaptors according to claim 14 and a polynucleotide sequencing adaptor according to any one of claims 15-19.

22. A method of uniquely labelling a RNA molecule, the method comprising (a) hybridising the RNA molecule to a double stranded polynucleotide adaptor according to any one of claims 1-13, and (b) ligating the hybridised double stranded polynucleotide adaptor to the RNA molecule.

23. A method of uniquely labelling a plurality of RNA molecules, the method comprising (a) hybridising the RNA molecules to a population of double stranded polynucleotide adaptors according to claim 14, and (b) ligating the hybridised double stranded polynucleotide adaptors to the RNA molecules.

24. A method according to claim 22 or 23, wherein the RNA molecule(s) is / are tRNA, wherein the double stranded polynucleotide adaptor(s) are as defined in any one of claims8-13 and wherein the method further comprises (c) reverse transcribing part of the first RIMA strand(s) and the tRNA molecule(s) to form a linear double stranded construct or linear double stranded constructs.

25. A method of characterising a RNA molecule or a plurality of RNA molecules, the method comprising (i) uniquely labelling the RNA molecule(s) using a method according to any one of the claims 22-24 and (ii) characterising the RNA molecule(s) and the barcode(s).

Citation Information

Patent Citations

  • Novel pore protein monomer and application thereof

    CN114957412A

  • Biomolecular sensors and methods

    US20170044605A1

  • phi 29 DNA polymerase

    US5576204A

  • A miniature support for thin films containing single channels or nanopores and methods for using same

    WO2000028312A1

  • Deliver of molecules to a li id bila

    WO2006100484A2