Identification of polyadenylated and non -polyadenylated mRNA mixtures through long-read direct RNA-sequencing

Cytidine tailing and ligation with a custom adapter enable the efficient sequencing and analysis of both polyadenylated and non-polyadenylated RNA molecules, addressing the limitations of existing methods and improving RNA characterization.

WO2026055078A1PCT designated stage Publication Date: 2026-03-12BIORELIANCE CORP
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing methods for direct RNA sequencing struggle to efficiently identify and differentiate between polyadenylated and non-polyadenylated RNA molecules, leading to incomplete characterization and potential contamination from truncated or antisense-stranded products.

Method used

A method involving cytidine tailing of RNA molecules, followed by ligation with a custom-designed adapter containing guanine and/or inosine bases, allows for the sequencing of both polyadenylated and non-polyadenylated RNAs using a double-stranded reverse transcription adapter.

Benefits of technology

Enables the simultaneous sequencing and analysis of both poly(A) and non-poly(A) RNA molecules, providing accurate characterization and quantification of poly(A) tail lengths, full-length RNA estimation, and detection of double-strand RNA contamination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025044245_12032026_PF_FP_ABST
    Figure US2025044245_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention provides materials and methods for RNA sequencing and analysis of both polyadenylated and non-polyadenylated RNA molecules.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: P24-157-SEC-WO01IDENTIFICATION OF POLYADENYLATED AND NON -POLYADENYLATED mRNA MIXTURES THROUGH LONG-READ DIRECT RNA-SEQUENCINGRelated Applications

[0001] The present application claims the benefit of priority of U.S. Provisional Patent Application No.: 63 / 832,616, filed June 30, 2025, and U.S. Provisional Patent Application No.: 63 / 691 , 123, filed September s, 2024, the entire content of each of which is incorporated herein by reference.Sequence Listing

[0002] This application contains a Sequence Listing that has been submitted in XML format via Patent Center and is hereby incorporated by reference in its entirety. The XML copy, created on August 13, 2025, is named P24-157-SEC- WO01_SL.xml, and is 3.72 kilobytes in size.Background

[0003] In a standard Oxford Nanopore Technologies Direct RNA Sequencing experiment, the double-stranded reverse transcription adapter (RTA) is designed to ligate to the 3’-end of the RNAs containing a Poly(A) tail of at least 10 nucleotides (nt) in length. A “typical” mRNA has a Poly(A) tail closer to 100 nt long. This is accomplished through base-pairing of the terminal region of the target RNA Poly(A) tail to at least a 10 nt Poly(T) overhang on the first paired strand of the RTA, followed by ligation of the second pared strand of the RTA to the 3’-end of the target mRNA. The RTA may not ligate to, and will therefore not include for characterization, RNAs lacking a Poly(A) tail, including contaminating truncated sense-stranded products or antisense- stranded products (dsRNA).

[0004] Other technologies have attempted to solve the problem of identification and analysis of non-Polyadenylated RNA with limited success.

[0005] United States Patent Application Publication No.: US 2022 / 0090056 describes the use of a Poly(A) tailing approach using immobilized Poly(A)- Polymerase and immobilized ligase to prepare libraries for Oxford Nanopore Direct RNA sequencing. The core of the invention involves the improvementAttorney Docket No.: P24-157-SEC-WO01 of sequencing efficiency by using solid phase-coupled enzymes. However, it does not identify non-Poly(A)-tailed RNA.

[0006] United States Patent Application Publication No.: US 2008 / 0300142 describes the tailing of an RNA target followed by ligating an adapter utilizing a splinted ligation approach. Split ligation is generally thought of as when specific oligonucleotides are ligated together using a ligase, e.g., T4 DNA ligase, and a bridging oligonucleotide complementary to the target oligonucleotides. However, the procedure is designed for microarray analysis and not direct RNA sequencing (“direct RNA-seq”).

[0007] United States Patent Application Publication No.: US 2021 / 0198730 describes the addition of molecular barcodes onto RNA targets using a Poly(C) or Poly(G / l) tailing reaction, followed by long read sequencing of PCR-amplified DNA. However, this approach does not use a splinted ligation step prior to reverse transcription, and the sequenced molecule is the PCR amplified DNA, not the RNA.

[0008] United States Patent Application Publication No.: US 2022 / 0002797 describes the utilization of Poly(G / l)-tai ling to install a Poly(G / l) tail onto polyadenylated mRNA, followed by reverse transcription, PCR amplification, and long read sequencing of PCR-amplified DNA. Similar to the above patent document, this approach does not use a splinted ligation step prior to reverse transcription, and the sequenced molecule is the PCR amplified DNA, not the RNA.

[0009] Begik, et al. attempted to develop a non-tailing approach to detect both Poly(A) and non-Poly(A) RNA with a template-switching-based sequencing protocol. The protocol, Nano3P-seq, can sequence a RNA molecule from its 3' end, regardless of its polyadenylation status, and without the need for PCR amplification or ligation of RNA adapters. (See, Begik, et al. Nano3P-seq: transcriptome-wide analysis of gene expression and tail dynamics using endcapture nanopore cDNA sequencing. Nat Methods 20, 75-85 (2023). doi.org / 10.1038 / s41592-022-01714-w). The technique fails to differentiate between RNA with and without a Poly(A) tail.

[0010] Drexler, et al. developed a protocol for nanopore analysis of Co- transcriptional Processing (nano-COP), for monitoring early RNA processing by directly sequencing long nascent RNAs. (See, Drexler et al. RevealingAttorney Docket No.: P24-157-SEC-WO01 nascent RNA processing dynamics with nano-COP. Nat Protoc. 2021 Mar; 16(3): 1343-1375. doi: 10.1038 / s41596-020-00469-y. Epub 2021 Jan 29. PMID: 33514943; PMCID: PMC8713461 ). nano-COP combines stringent purification of nascent RNA, long-read sequencing without reverse transcription or PCR amplification, and rigorous computational analyses. However, nano-COP utilizes an l-tailing procedure which does not sufficiently discriminate l-tails from A-tails.

[0011] Vo, et al. developed an l-tailing procedure with the goal of allowing for broad capture of more diverse populations of RNA without loss of native 3' end information, but the technique is not sufficiently developed to discriminate l-tails from A-tails and deconvolute read sample origins. (See, Vo et al., RNA (2021 ) 27:1497-1511 )

[0012] Yuan, et al. developed a method of 3’ -tailing of RNA with modified adenosines including, 2’-O-methyladenosine (mA). (See, Yuan, et al., Efficient 3'-end tailing of RNA with modified adenosine for nanopore direct total RNA sequencing, pre-print published February 25, 2024 by bioRxiv, doi. org / 10.1101 / 2024.02.24.581884) Yuan provides for RNA tailing with several modified adenosines but does not provide teachings or suggestions of RNA tailing without use of modified adenosines. Additionally, Yuan shows their several various modified adenosines generate variable results at least with regard to electric signal detected during sequencing indicating a significant level of unpredictability in the field for different tailing moieties. Further, Yuan reports a difference in the efficiency of addition of (mA) to Poly(A)-tailed RNA and non-Poly(A)-tailed RNA which may lead to imprecise or erroneous results.

[0013] Ibrahim, et al. developed a non-tailing approach to identify mixtures of Poly(A) and non-Poly(A) RNA using direct RNA-sequencing, termed True End-to-End RNA Sequencing (TERA-Seq). (See, Fadia et al., TERA-Seq: true end-to-end sequencing of native RNA molecules for transcriptome characterization, Nucleic Acids Research, Volume 49, Issue 20, 18 November 2021 , Page e115, doi.org / 10.1093 / nar / gkab713). However, the technique requires large samples, which may be limiting in certain settings such as with clinical samples.

[0014] Therefore, compositions and methods for the improved detection and analysis of both Poly(A) and non-Poly(A) RNA, including differentiationAttorney Docket No.: P24-157-SEC-WO01 between RNA with and without a poly(A) tail, would be an improvement in the art.Summary of the Invention

[0015] Sequencing and analyzing RNA, both Poly(A) and non-Poly(A) RNA is important in genomic analysis and other biological research areas and is instrumental in accelerating biological and medical advancement.

[0016] Some embodiments described here may be used to adapt mRNA molecules in, for example, synthetically produced mRNA samples for producing direct RNA-sequencing libraries, and these libraries can be subjected to direct RNA-sequencing. Analysis of produced direct RNA- sequencing data can then be used to assess the Poly(A) tail lengths and fraction of RNA molecules having a Poly(A) tail in synthetically produced mRNA samples. Poly(A) length affects mRNA fate and translation efficiency - knowing Poly(A) length will, for example, provide a better understanding of the physiological role of tail heterogeneity, and recognition of a correlation between tail length and RNA translatability. (See, for example, Jalkanen et al. Determinants and implications of mRNA poly(A) tail size-does this protein make my tail look big? Semin Cell Dev Biol. 2014 Oct;34:24-32).

[0017] In some embodiments, analysis of produced direct RNA-sequencing data can also then be used to assess in the fraction of Poly(A) containing mRNA molecules in synthetically produced mRNA samples and / or the full length of mRNA molecules in synthetically produced mRNA samples and / or the proportion of double-stranded contaminating molecules in synthetically produced mRNA samples. Analysis of double-stranded RNA contamination is an important measure of fidelity of synthetically produced mRNA.

[0018] In some embodiments, compositions and methods for preparing a nucleotide library for sequencing and analysis are described herein. More specifically, some embodiments provide for more efficient nucleotide sequencing by permitting the sequencing of both Poly(A) tagged ( / .e., Poly(A) tailing) and non-Poly(A) tagged nucleotide sequences.

[0019] The current state of the art provides for a double-stranded reverse transcription adapter (RTA) designed to ligate to the 3’-end of RNAs containing a Poly(A) tail. The Poly(A) tail is typically about 100 nucleotides (nt)Attorney Docket No.: P24-157-SEC-WO01 but may be shorter or longer. The ligation is accomplished through baseparing of the overhang on the bottom strand of the RTA followed by ligation of the top strand of the RTA to the 3’-end of the target mRNA. The RTA will not ligate, and therefore, will not include for characterization, the RNA molecules lacking a Poly(A) tail, including contaminating truncated sense-stranded products of antisense-stranded products (dsRNA).

[0020] Some embodiments described herein solve this problem in the prior art by providing compositions and methods to enable the sequencing of RNA that have or do not have a Poly(A) tail.

[0021] In some embodiments, a cytidine polymer is enzymatically added at the 3’-end of the RNA molecule, referred to herein as “C-taili ng .” In some embodiments, the process of C-tailing provides a consistent 3’-end sequence installed onto any RNA regardless of the existing 3’-end sequence. In some embodiments, Poly(A) containing or non-Poly(A) containing sequences would all have a consistent 3’ -end sequence. In some embodiments, a custom designed RTA containing complementary sequence overhang, e.g., a mixture of guanine (G) and / or inosine (I) bases ranging from 3 - 9 nt in length, in the bottom (first) paired strand of the adapter is used. The terms “top strand” and “bottom strand,” as used herein, do not refer to relative positions of the strands to each other but, rather, indicate a first paired and a second paired strand. The bottom strand of the adapter will allow for the ligation of the top strand of the RTA to any RNA that has undergone C-tailing thereby allowing for the capture of both polyadenylated (Poly(A)) and non-polyadenylated RNAs in a single process run.

[0022] In some embodiments, cytidine triphosphate (CTP) is fused to C-tail RNAs by using poly(A) polymerase. The poly(A) polymerase may be of any origin. In some embodiments, yeast or E. coli derived poly(A) polymerase is used due to availability and low cost. In some embodiments, sequence lengths of 1 - 200 cytidine(s) nt(s) may be added to the 3’-end of RNA molecules. In some embodiments, this process takes from about 1 to about 10 minutes. In some embodiments, a double-stranded cyti dine- RTA (C- RTA) containing a DNA overhang on the bottom strand of the C-RTA (comprised of G and / or I bases and ranging from 3 - 9 nt in length but may be longer or shorter) is used to anneal to poly(C)-containing RNAs, followed by ligation ofAttorney Docket No.: P24-157-SEC-WO01 the top strand of the C-RTA to the 3’-end of the RNA using T4 DNA ligase, or equivalent.

[0023] In some embodiments, reverse transcription of the ligated RNA provides stability to the RNA (or, optionally, no transcription prior to the direct RNA-sequencing), followed by ligation of RMX / R LA adaptor (for example, Oxford Nanopore Technologies (ONT), Lexington, MA and Oxford, UK, direct RNA-sequencing kit SQK-RNA002 or SQK-RNA-004 or latter versions or equivalent) followed by direct RNA sequencing using, for example, Nanopore flowcell (FLO-MIN 160D / FLO-PR0002 of FLO-MIN004RA / FLO-PR0004RA or later versions or equivalent). Analysis of direct RNA-sequencing data for RNA poly(A) tail length, fraction of poly(A)-containing RNAs, full-length RNA length estimation and contamination antisense RNA (dsRNA) can be ascertained with the compositions and methods described herein.

[0024] In some embodiments, the enzymes are coupled to a solid support. Examples of suitable solid supports include beads, tubes, multi-well plates, membranes, woven fibers, non-woven fibers, etc. In another aspect, the enzymes are not coupled to a solid support. In some embodiments, the use of solid supports for binding enzymes is specifically excluded in the present invention.

[0025] In some embodiments, a method of adapting a mixed population of polyadenylated and non-polyadenylated RNA molecules for analysis comprises: contacting the mixed population of RNA molecules with a cytidine tailing enzyme to add one or more cytidine residues at the 3’ -end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter- tagged cytidine-tailed RNA molecules, the sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and isolating the adapter-tagged cytidine-tailed RNA molecules.

[0026] In some embodiments, the method further comprises subjecting adapter-tagged cytidine-tailed RNA molecules to direct sequencing or indirect sequencing.

[0027] In some embodiments, the method further contemplates that the number of cytidine residues added to the RNA molecule is from 1 to about 200, 2 to about 100, 3 to about 50, 4 to about 20 or 5 to about 12.Attorney Docket No.: P24-157-SEC-WO01

[0028] In some embodiments, the method further contemplates that the length of the tail sequence comprising guanine and / or inosine bases on the RNA sequencing adapter is from 3 nt to 15 nt or 3 nt to 9 nt in length.

[0029] In some embodiments, the ligase is selected from the group consisting of T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, E. coli DNA ligase and T4 RNA Ligase 2. In some embodiments, the ligase is T4 DNA ligase.

[0030] In some embodiments, the cytidine tailing enzyme is a Poly(A) Polymerase or a Poly(U) Polymerase. In some embodiments, the cytidine tailing enzyme is selected from the group consisting of E. coli Poly(A) Polymerase I, Poly(A) Polymerase and S. pombe Poly(ll) Polymerase In some embodiments, the cytidine tailing enzyme is E. coli Poly(A) Polymerase I.

[0031] In some embodiments, the isolated adapter-tagged cytidine-tailed RNA molecules are sequenced and analyzed to determine the length of Poly(A) tails on the polyadenylated RNA molecules in the mixed population of RNA molecules.

[0032] In some embodiments, the isolated adapter-tagged cytidine-tailed RNA molecules are sequenced and analyzed to identify the Poly(A) and nonPolyp) tailed RNA molecules in a sample and quantify the proportion of Poly(A) to non-Poly(A) in the mixed population of RNA molecules.

[0033] In some embodiments, the isolated adapter-tagged cytidine-tailed RNA molecules are sequenced and analyzed to estimate the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules.

[0034] In some embodiments, the isolated adapter-tagged cytidine-tailed RNA molecules are sequenced and analyzed to identify contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNA in the mixed population of RNA molecules.

[0035] In some embodiments, the analysis involves one or more steps of applying an algorithm to identify the presence or absence of a poly(A) tail for each read in a generated dataset; and / or applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset; and / or applying an algorithm to calculate the poly(A) tail length of a poly(A)Attorney Docket No.: P24-157-SEC-WO01 region given the starting and ending locations of the poly(A) tail region for each read identified as containing a poly(A) tail in a generated dataset; and / or applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence to estimate the full length of the mRNA for each read in a generated dataset; and / or applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence in both the forward and reverse direction to identify antisense mRNAs.

[0036] In some embodiments, a method of determining the length of polyadenylated (poly(A)) tails on RNA molecules in a sample comprises: contacting a population of RNA molecules with poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’ -end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter- tagged cytidine-tailed RNA molecules, the sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and isolating the adapter-tagged cytidine-tailed RNA molecules and then direct or indirect sequencing the adapter-tagged cytidine-tailed RNA molecules; and determining one or both of: i) poly(A) tail lengths of the individual RNA molecules, and / or ii) the distribution of Poly(A) tail lengths of RNA molecules in the sample by applying an algorithm to identify the presence or absence of a poly(A) tail for each read in a generated dataset and / or applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset.

[0037] In some embodiments, a method of identify both poly(A) and non- poly(A) tailed RNA molecules in a sample comprises: contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’ -end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter- tagged RNA molecules, the sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and isolating the adapter-taggedAttorney Docket No.: P24-157-SEC-WO01RNA molecules and then direct or indirect sequencing the adapter-tagged molecules; and determining the one or both of: i) the proportion of poly(A) containing RNA molecules to non-poly(A) containing molecules, and ii) identifying the non-poly(A) containing RNA molecule in the sample by applying an algorithm to identify the presence or absence of a poly(A) tail for each read in a generated dataset.

[0038] In some embodiments, a method to estimate the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules comprises: contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter- tagged RNA molecules, the sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and isolating the adapter-tagged RNA molecules and then direct or indirect sequencing the adapter-tagged molecules; and determining the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules by applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence to estimate the full length of the mRNA for each read in a generated dataset.

[0039] In some embodiments, a method to identify contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNA in the mixed population of RNA molecules comprises: contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter-tagged RNA molecules, the sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and isolating the adapter-tagged RNA molecules and then direct or indirect sequencing the adapter-tagged molecules; and identifying contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNAAttorney Docket No.: P24-157-SEC-WO01 in the mixed population of RNA molecules by applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence in both the forward and reverse direction to identify antisense mRNAs which is indicative of contamination of dsRNA.Description of the Figures

[0040] Fig. 1 shows an example of a distribution (histogram) of Poly(A) tail lengths for sequenced mRNA sample and statistics table for Poly(A) tail length.

[0041] Fig. 2 shows an example of a table displaying the read counts of Poly(A)-plus and Poly(A)-minus mRNAs for sequenced mRNA sample and percentage of Poly(A)-containing molecules.

[0042] Fig. 3 shows an example of a distribution (histogram) of full mRNA lengths for sequenced mRNA sample.

[0043] Fig. 4 shows are read counts and coverage maps of sense and antisense reads for sequenced mRNA sample. Antisense reads indicate potential for dsRNA presence.Definitions

[0044] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0045] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton, et al., Dictionary of Microbiology and Molecular Biology (2nd ed. 1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger, et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991 ). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.Attorney Docket No.: P24-157-SEC-WO01

[0046] When introducing elements of the present disclosure or the preferred embodiments(s) thereof, the articles "a," "an," "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising," "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.

[0047] The transitional phrases “comprising,” “consisting essentially of” and “consisting of” have the meanings as given in MPEP 2111.03 (Manual of Patent Examining Procedure; United States Patent and Trademark Office, 9thEd., Revision Feb 2023 [R-07.2022]). Any claims using the transitional phrase “consisting essentially of” will be understood as reciting only essential elements ( / .e., the basic and novel characteristics) of the invention and any other elements recited in dependent claims depending from the claim using the transitional phrase “consisting essentially of” are understood to be non- essential to the invention recited in the claim from which they depend.Likewise, any additional elements over those claimed that are described in a prior art reference(s) are excluded from the claim by use of the transitional phrase “consisting essentially of.”

[0048] Any numerical ranges provided herein include all numbers, whole, fractional and decimal, within the range as though they had been specifically recited. Thus, for example, a range of 10 - 100 also includes each and every number between 10 and 100 including without exception every whole number and every fractional or decimal number between the recited numbers. Further, phrases such as “at least” and “no more than” are understood to recite ranges starting with or ending with the recited number, respectively, and include all numbers within that range as though they had been specifically recited.

[0049] As used herein, the term "poly(A)" or "poly-A" or similar expressions are used interchangeably and refer to polyadenosine. Polyadenosine is a long chain molecule containing more than one or dozens to hundreds of adenosine residues formed by multiple adenosine molecules linked together by phosphodiester bonds. This polymer is characterized by being composed of adenosine residues, wherein the adenosine residues contain an adenine base and a ribose, and part of the phosphate group after the formation of the phosphodiester bond. One of the most common applications of "poly(A)” refers to the polyadenylic acid tail on ribonucleic acid (RNA), which is usuallyAttorney Docket No.: P24-157-SEC-WO01 called, for example, a “poly(A) tail”, “poly(A) structure" or "polyadenylation tail." In the RNA of eukaryotic organisms, the 3' end of the mRNA (messenger RNA) molecule usually has a poly(A) tail. This poly(A) tail is added to the 3' end of the mRNA molecule by a specific enzyme during the post-transcription process of RNA. The poly(A) tail plays an important role in many biological processes. In in vitro synthesized mRNA technology, a poly(A) tail consisting of a small number (generally 5 to 8) of non-A bases is also widely used.

[0050] The term “synthetic RNA” refers to RNA that is “machine made,” i.e., made outside of a cell by oligonucleotide synthesis devices and techniques.

[0051] As used herein, the term “tailing enzyme” refers to templateindependent enzymes (e.g., polymerases, transferases) that add one or more nucleotides or ribonucleotides to the 3' end of a polynucleotide. Tailing enzymes may add one or more adenines (A), one or more guanines (G), one or more thymine (T), one or more cytosines (C), or one or more uracils (II). Tailing enzymes may be selected for specific applications based on their preference for adding a particular nucleotide or ribonucleotide, for example, to compliment the end of an adapter to which the tailed polynucleotide. Nonlimiting examples of tailing enzymes include poly(A) polymerases, poly(G) polymerases, poly(ll) polymerases, and terminal deoxynucleotidyl transferase (TdT).

[0052] As used herein, the term “ligase” refers to enzymes that join polynucleotide ends together. Ligases include ATP-dependent double-strand polynucleotide ligases, NAD+-dependent double-strand DNA or RNA ligases and single-strand polynucleotide ligases. Ligases may include any of the ligases described in EC 6.5.1.1 (ATP-dependent ligases), EC 6.5.1.2 (NAD+- dependent ligases), EC 6.5.1 .3 (RNA ligases) (see ExPASy Bioinformatics Resource Portal having a URL of enzyme.expasy.org which is a repository of information concerning nomenclature of enzymes based on the recommendations of the Nomenclature Committee of the International Union of Biochemistry and Molecular Biology (IUBMB) describing each type of characterized enzyme for which an EC (Enzyme Commission) number has been provided). Specific examples of ligases include bacterial ligases such as E. coli DNA ligase and Taq DNA ligase, Ampligase® thermostable DNA ligase (Epicentre® Technologies Corp., part of Illumina®, Madison, Wis.) and phageAttorney Docket No.: P24-157-SEC-WO01 ligases such as T3 DNA ligase, T4 DNA ligase, T7 DNA ligase, 9° N DNA ligase, and mutants thereof.

[0053] As used herein, the term, “adapter” refers to a sequence that is joined to or can be joined to another molecule (e.g., ligated or copied onto via primer extension). An adapter can be DNA or RNA, or a mixture of the two. An adapter may be 15 to 100 bases, e.g., 50 to 70 bases, although adapters outside of this range are contemplated. Adapters may be designed to serve a specific purpose. For example, adapters may be designed for use in sequencing applications. Sequencing adapters may comprise, for example, an oligo-(dT) overhang, a barcode sequence, an overhang (other than oligo- (dT)) to anneal to another adapter, a site for anchoring a motor protein, and a sequence to bind to tethering oligos with affinity to polymer membrane for guiding a DNA or RNA fragment (on which it resides) to the vicinity of a nanopore, and combinations thereof.

[0054] As used herein, the term "endogenous sequence" refers to a chromosomal sequence that is native to the cell.

[0055] The term "exogenous" or “exogenous sequence,” as used herein, refers to a sequence that is not native to the cell or a chromosomal sequence whose native location in the genome of the cell is in a different chromosomal location.

[0056] A "gene," as used herein, refers to a DNA region (including exons and introns) encoding a gene product, as well as all DNA regions which regulate the production of the gene product, whether or not such regulatory sequences are adjacent to coding and / or transcribed sequences. Accordingly, a gene includes, but is not necessarily limited to, promoter sequences, terminators, translational regulatory sequences such as ribosome binding sites and internal ribosome entry sites, enhancers, silencers, insulators, boundary elements, replication origins, matrix attachment sites, and locus control regions.

[0057] The terms “complementary” or “complementarity” refer to the association of double-stranded nucleic acids by base pairing through specific hydrogen bonds. The base paring may be standard Watson-Crick base pairing (e.g., 5 -A G T C-3' pairs with the complementary sequence 3'-T GA GS'). The base pairing also may be Hoogsteen or reversed HoogsteenAttorney Docket No.: P24-157-SEC-WO01 hydrogen bonding. Complementarity is typically measured with respect to a duplex region and thus, excludes overhangs, for example. Complementarity between two strands of the duplex region may be partial and expressed as a percentage (e.g., 70%), if only some of the base pairs are complementary. The bases that are not complementary are “mismatched.” Complementarity may also be complete ( / .e., 100%), if all the base pairs of the duplex region are complementary.

[0058] The terms "nucleic acid" and "polynucleotide" refer to a deoxyribonucleotide or ribonucleotide polymer, in linear or circular conformation, and in either single- or double-stranded form. For the purposes of the present disclosure, these terms are not to be construed as limiting with respect to the length of a polymer. The terms can encompass known analogs of natural nucleotides, as well as nucleotides that are modified in the base, sugar and / or phosphate moieties (e.g., phosphorothioate backbones). In general, an analog of a particular nucleotide has the same base-pairing specificity; / .e., an analog of A will base-pair with T.

[0059] The term "nucleotide" refers to deoxyribonucleotides or ribonucleotides. The nucleotides may be standard nucleotides ( / .e., adenosine, guanosine, cytidine, thymidine, and uridine) or nucleotide analogs. A nucleotide analog refers to a nucleotide having a modified purine or pyrimidine base or a modified ribose moiety. A nucleotide analog may be a naturally occurring nucleotide (e.g., inosine) or a non-naturally occurring nucleotide. Non-limiting examples of modifications on the sugar or base moieties of a nucleotide include the addition (or removal) of acetyl groups, amino groups, carboxyl groups, carboxymethyl groups, hydroxyl groups, methyl groups, phosphoryl groups, and thiol groups, as well as the substitution of the carbon and nitrogen atoms of the bases with other atoms (e.g., 7-deaza purines). Nucleotide analogs also include dideoxy nucleotides, 2'-O-methyl nucleotides, locked nucleic acids (LNA), peptide nucleic acids (PNA), and morpholines.

[0060] The terms "polypeptide" and "protein" are used interchangeably to refer to a polymer of amino acid residues.Attorney Docket No.: P24-157-SEC-WO01Detailed Description of the InventionA. Long-read Direct RNA Sequencing

[0061] As explained by Stark, et el., (Stark, R., Grzelak, M. & Hadfield, J., RNA sequencing: the teenage years. Nat Rev Genet 20, 631-656 (2019). doi.org / 10.1038 / s41576-019-0150-2), long-read methods, like the baseline short-read platform, have relied on converting mRNA to cDNA before sequencing. Nanopore sequencing technology (e.g., Oxford Nanopore (Lexington, MA and Oxford, UK)) can be used to sequence RNA directly — that is, without modification, cDNA synthesis and / or PCR amplification during library preparation. This approach, termed direct RNA sequencing (dRNA- seq), removes the biases generated by indirect processes and enables epigenetic information to be retained.

[0062] As explained further by Stark, et al., library preparation from RNA involves sequential ligation of two adaptors. First, a duplex adaptor bearing an oligo(dT) overhang is annealed and ligated to the RNA polyadenylation (poly(A)) tail of the target RNA, which is followed by an optional (but recommended) reverse-transcription step that improves the sequencing throughput. The second ligation step attaches the sequencing adaptors, which are pre-loaded with the motor protein that drives sequencing. A motor protein drives the nucleic acid molecule through the nanopore in a step-wise manner for the sequencing step. Motor proteins are known to those of skill in the art. The library is then ready for MinlON sequencing, in which RNA is sequenced directly from the 3' poly(A) tail to the 5' cap. dRNA-seq generates read lengths of around 1 ,000 bp, with maximum lengths exceeding 10 kb. These long reads have several advantages over short reads including improved isoform detection compared with short reads. However, the advantages of long-read RNA sequencing technologies is reliant having high-quality RNA libraries. Improved RNA-library preparation, including the detection of nonpolyadenylated RNA, is an aspect of the present invention.

[0063] Some embodiments described herein contemplate that direct RNA- sequencing library preparation is performed using components of Oxford Nanopore Technologies Direct RNA-sequencing kit (or equivalent) (preferably SQK-RNA004 but also SQK-RNA002). The library can then be loaded into anAttorney Docket No.: P24-157-SEC-WO01Oxford Nanopore Technologies flow cell (preferably FLQ-MIN004RA or FLO- PRO004RA, but also FLO-MIN106D or FLO-PRO002, or equivalent) and sequenced using an Oxford Nanopore Technologies sequencing device (preferably GridlON, but PromethlON or MinlON are also suitable, or equivalent). For the present invention, modifications will be made relating to the use of poly(C) tails, as detailed below in the Exemplification section.

[0064] It is further contemplated by the present invention that sequencing may be indirect sequencing. Indirect sequencing methods, for example, use a reverse transcriptase enzyme to transcribe RNA molecules to complementary DNA (cDNA), followed by sequencing of the cDNA. During analysis of the cDNA sequence, the template RNA base sequence (adenine [A], cytosine [C], guanine [G], uracil [U]), and sometimes, additional information about modified nucleotides, is derived. (See, for example: National Academies of Sciences, Engineering, and Medicine; Health and Medicine Division; Division on Earth and Life Studies; Board on Health Sciences Policy; Board on Life Sciences; Toward Sequencing and Mapping of RNA Modifications Committee. Charting a Future for Sequencing RNA and Its Modifications: A New Era for Biology and Medicine. Washington (DC): National Academies Press (US); 2024 Jul 22. Chapter 3, Current and Emerging Tools and Technologies for Studying RNA Modifications. Available from: www.ncbi.nlm.nih.gov / books / NBK606045 / ).A. Data Analysis

[0065] Data analysis programs used with the present invention are known open source software programs and are integral with the Oxford Nanopore Technologies sequencing devices. Their uses and functions are known to those of skill in the art as is evidenced by the references provided below.

[0066] Data is analyzed by one or more of:1) applying an algorithm (e.g., Nanopolish (nanopolish. readthedocs.io / en / latest / # Nanopolish) is an open source software package for signal-level analysis of Oxford Nanopore sequencing data. Its use is known by those of skill in the art. See, e.g., Loman, Nicholas J., Joshua Quick, and Jared T. Simpson. “A complete bacterial genome assembled de novo using only nanopore sequencing data.” Nature methodsAttorney Docket No.: P24-157-SEC-WO0112.8 (2015): 733-735; Quick, Joshua, et al. “Real-time, portable genome sequencing for Ebola surveillance.” Nature 530.7589 (2016): 228-232; Simpson, Jared T, et al. “Detecting DNA cytosine methylation using nanopore sequencing.” nature methods 14.4 (2017): 407-410) to identify the presence or absence of a poly(A) tail for each read in a generated dataset;2) and / or applying an algorithm (e.g., Nanopolish, Dorado (github.com / nanoporetech / dorado) is an opensource base caller software program for open source basecaller for Oxford Nanopore reads. Its use is known by those of skill in the art; TailFindR (github.com / adnaniazi / tailfindr) is an opensource package for estimating poly(A)-tail lengths in Oxford Nanopore reads. Its use is known by those of skill in the art) to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset;3) and / or applying an algorithm (e.g., Nanopolish, Dorado, TailFindR) to calculate the poly(A) tail length of a poly(A) region given the starting and ending locations of the poly(A) tail region for each read identified as containing a poly(A) tail in a generated dataset;4) and / or applying a series of algorithms (e.g. Dorado and Minimap2 (github.com / lh3 / minimap2) used to analyze and align nucleotide sequences) to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence to estimate the full length of the mRNA for each read in a generated dataset;5) and / or applying a series of algorithms (e.g., Dorado and Minimap2) to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence in both the forward and reverse direction to identify antisense mRNAs.

[0067] As various changes could be made in the above-described processes and kits without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting senseAttorney Docket No.: P24-157-SEC-WO01ExamplesExample 1 : Preparation of U-Reverse Transcriptase Adapter (RTA)

[0068] 10X annealing buffer (100 mM Tris-HCL, pH 7.5 1 M, NaCI 5 M, diethyl pyrocarbonate (DEPC)-treated water) was sterilized. 100 pl of U-RTA mixture was made by mixing 10.0 pl of the 10X annealing buffer, 1 .4 pl Oligo-A (100 pM), 1.4 pl Oligo-B (100 pM) and 87.2 pl DEPC-treated water and well mixed.Oligo-A / 5Phos / GGCTTCTTCTTGCTCTTAGGTAGTAGGTTC [SEQ ID NO: 1]Oligo-BGAGGCGAGCGGTCAATTTTCCTAAGAGCAAGAAGAAGCCGGG / ide oxyl / GG [SEQ ID NO: 2]

[0069] A thermal cycler was used to anneal the U-RTA under the conditions in Table 1 .Table 1 : Thermal cycling conditions for U-RTA adapter annealing

[0070] After cycling, the mixture was mixed briefly and kept on ice or a cold block.

[0071] An alternative exemplification of U-Reverse Transcriptase Adapter (RTA) is performed as follows:

[0072] RxnReady™ Primer Pool (Integrated DNA Technologies) comprised of mixture of Oligo-A and Oligo-B (both primers at a final concentration in the pool of 1.4 pM) in I DTE buffer (1x TE buffer, pH 8.0). The Primer Pool is then subjected to annealing according to Table 1.Example 2: Polynucleotide kinase (PNK) treatment of mRNA

[0073] A 118ng / pl dilution (50 pl) of sample mRNA was prepared according to the formula:Attorney Docket No.: P24-157-SEC-WO01

[0074] C1 * V1 = C2 * V2

[0075] Where C1 is concentration of stock extracted mRNA (in ng / pl), V1 is volume of stock mRNA to input into dilution (in pl), C2 is final 118 ng / pl dilution concentration, and V2 is total volume of dilution (in pl).

[0076] Sample mRNA (118ng / pl) was placed on a heat block and incubated at 70°C for 5 minutes. Immediately after incubation, sample mRNA tubes were placed on ice or a cold block.

[0077] The following components were combined in a 1 ,5mL lobind tube according to Table 2:Table 2

[0078] Using a 200ul pipette, the sample prepared in Table 2 were mixed by pipetting up and down gently approximately 8 times

[0079] Sample tube was placed on a heat block and incubate at 37°C for 30 minutes.

[0080] 72 pl of RNACIean beads (Beckman Coulter, Brea, CA) was added to sample tube and mixed thoroughly by pipetting up and down approximately 8 times.

[0081] The sample and bead mixture was incubated at room temperature on a sample rotator (e.g., Hula mixer, Thermo Fisher) for 5 minutes. After incubation, sample tube was briefly spun down on mini / microcentrifuge.Sample tube was placed on a magnetic particle concentrator (MFC) for beads to migrated to the magnet and until the solution was completely clear (approximately 1-2 minutes).

[0082] Using a pipette, the clear supernatant was removed and discarded without disturbing the bead pellet.

[0083] 200pl of 80% EtOH wash was carefully pipetted into sample tube.Attorney Docket No.: P24-157-SEC-WQ01

[0084] Wash was allowed to incubate at room temperature for approximately 30 seconds, then was removed and discarded without disturbing the bead pellet.

[0085] A second wash was performed.

[0086] After removing the wash solution, sample tube was removed from the MFC and capped. Sample tube was spun down on a mini / microcentnfuge to collect any remaining wash solution at bottom of tube.

[0087] Sample tube was placed onto MPC, allowing beads to migrate to the

[0088] magnet. Using a 20pl micro pipette, for example, any residual wash solution was carefully removed.

[0089] The tubes were removed from the MPC. The bead pellet was resuspended in 20 pl nuclease-free water by carefully pipetting up and down approximately 8 times.

[0090] The resuspended beads were incubated for 5 minutes at room temperature.

[0091] After incubation, sample tube was briefly spun down on mini / microcentrifuge and placed again on the MPC for the beads to migrate to magnet and until solution was completely clear (approximately 1-2 minutes).

[0092] Using a 20pl micropipette, the clear supernatant (containing sample mRNA) was carefully transferred to a new 1 ,5mL Lobind tube and placed the tube on ice or a cold block.

[0093] Concentration of 1 pl of sample mRNA was assayed on a Qubit Fluorometer using the Qubit BR RNAAssay kit (Thermo Fisher, Waltham, MA).Example 3: C-Tailing of mRNA

[0094] C-tailing reaction mixture was made according to Table 3.Table 3: C-tailing reaction mixtureAttorney Docket No.: P24-157-SEC-WO01

[0095] A 50ng / pl dilution (50 pl) of PNK-treated sample mRNA was prepared according to the formula: i. Ci * Vi = C2* V2ii. Where Ci is concentration of stock extracted mRNA (in ng / pl), Vi is volume of stock mRNA to input into dilution (in pl), C2is final 50 ng / pl dilution concentration, and V2is total volume of dilution (in pl).

[0096] Sample mRNA (50ng / pl) was placed on a heat block and incubated at 70 °C for 5 minutes. Immediately after incubation, the sample mRNA tube was placed on ice or a cold block.

[0097] Move tube (e.g., LoBind (Eppendorf, Wesseling-Berzdorf, Germany) or equivalent) containing mRNA C-tailing mixture (from Table 3) from ice / cold block to room temperature. Sample mRNA was combined with previously prepared C-tailing mixture following instructions in Table 4:Table 4: mRNA C-tailing mixture

[0098] Using a 200ul pipette, the samples prepared in Table 4 were mixed by pipetting up and down gently approximately 4 times.

[0099] To the mRNA C-tailing mixture, 2pl E-PAP enzyme was added. Using a 200ul pipette, sample was mixed by pipetting up and down approximately 8 times.

[0100] The sample tubes were placed on a heat block and incubate at 37 °C for 6 minutes.Attorney Docket No.: P24-157-SEC-WO01

[0101] Immediately after the 6 minute incubation, sample tubes were removed from heat block. 1 ul of 500mM EDTA was added to samples. Samples were mix by pipetting up and down approximately 8 times.

[0102] 90pl of RNACIean beads (Beckman Coulter, Brea, CA) were added to sample and mixed thoroughly by pipetting up and down approximately 8 times. Sample and bead mixture were incubated at room temperature for 5 minutes.

[0103] After incubation, sample tubes were briefly spun down on mini / microcentrifuge. Sample tubes were placed on a magnetic particle concentrator (MPC) for beads to migrated to the magnet and until the solution was completely clear (approximately 1-2 minutes).

[0104] Using a pipette, the clear supernatant was removed and discarded without disturbing the bead pellet.

[0105] 200pl of 80% EtOH wash was carefully pipetted into sample tube.

[0106] Wash was allowed to incubate at room temperature for approximately 30 seconds, then was removed and discarded without disturbing the bead pellet. A second wash was performed.

[0107] After removing the wash solution, sample tubes were removed from the MPC and capped. Sample tubes were spun down on a mini / microcentrifuge to collect any remaining wash solution at bottom of tube.

[0108] Sample tubes were placed onto MPC, allowing beads to migrate to the magnet. Using a 20pl micro pipette, for example, any residual wash solution was carefully removed.

[0109] The tubes were removed from the MPC. The bead pellet was resuspended in 15pl nuclease-free water by carefully pipetting up and down approximately 8 times.

[0110] The resuspended beads were incubated for 5 minutes at room temperature.Attorney Docket No.: P24-157-SEC-WQ01

[0111] After incubation, sample tubes were briefly spun down on mini / microcentrifuge and placed again on the MFC for the beads to migrate to magnet and until solution was completely clear (approximately 1-2 minutes).

[0112] Using a 20pl micropipette, the clear supernatant (containing sample mRNA) was carefully transferred to a new 1.5mL Lobind tube and placed the tube on ice or a cold block.

[0113] Concentration of 1 pl of sample was assayed mRNA on a Qubit Fluorometer using the Qubit BR RNA Assay kit (Thermo Fisher, Waltham, MA).Example 4: Direct RNA-Sequencing

[0114] If mRNA sample preparation was performed on a prior day, retrieve samples from storage (-80 °C) and thaw on ice / cold block.

[0115] Calculate the volume of mRNA sample to input according to the following equation:C-tailed mRNA input in pl (A) = 250 ng I mRNA concentration (ng / pl)

[0116] Calculate the volume of DEPC-treated water input according to the following equation:DEPC-treated water input in pl (B) = 8.5ul - mRNA input in pl (A)

[0117] The following components were mixed in a thin-walled PCR tube according to Table 5:Table 5: U-RTA and mRNA sample mixture

[0118] The mRNA and adapter mixture was mixed and then spun down using a microcentrifuge to collect the sample mixture at the bottom of the tube.

[0119] The mRNA sample / adapter mix was incubated in a thermal cycler using the program shown in Table 6:Attorney Docket No.: P24-157-SEC-WO01Table 6: Thermal cycling conditions for mRNA sample / adapter mix

[0120] Immediately after reaching 4 °C, the samples were placed on ice or cold block.

[0121] The first ligation mixture was prepared at room temperature in a 0.2 ml PCR tube following the instructions in Table 7.Table 7: First mRNA ligation step mixture instructions

[0122] The mRNA ligation mixture was mixed by pipetting up and down approximately 8 times and then spun down using a microcentrifuge to collect the sample mixture at the bottom of the tube.

[0123] The ligation mixture was incubated for 10 minutes at room temperature.

[0124] After the 10 minute ligation incubation, the sample ligation mixture was combined with the reverse transcription reagents in a 0.2ml PCR tube following the instructions in Table 8.Table 8: Reverse transcription reaction mixtureAttorney Docket No.: P24-157-SEC-WO01

[0125] The reverse transcription reaction mixture was mixed and spun down the sample tube using a microcentrifuge to collect the sample mixture at the bottom of the tube.

[0126] The reverse transcription of sample mRNA was performed using a thermal cycler (e.g., C1000, Thermo Fisher) using the cycling conditions outlined in Table 9.Table 9: Reverse transcription thermal cycling conditions

[0127] Following reverse transcription, the sample mixture is transferred to a new 1 ,5mL LoBind tube or equivalent.

[0128] 72 pl RNACIean beads were added to reverse-transcribed mRNA and mixed thoroughly.

[0129] The sample and bead mixture was incubated at room temperature on a sample rotator (e.g., Hula mixer, Thermo Fisher) for 5 minutes.

[0130] After incubation, the tube was spun down on a mini / microcentrifuge to collect the sample at the bottom of the tube.

[0131] The sample tubes were placed on a MFC until the beads to migrate to the magnet solution was completely clear (approximately 1-2 minutes).

[0132] The clear supernatant was removed and discarded taking care not to disturb the bead pellet.

[0133] 150pl of 70% EtOH wash was carefully pipetted into each sample tube.

[0134] The sample tubes were closed and while keeping the sample tube in the MPC on the benchtop, the tubes were rotated by 180°. The beads to migrated to the magnet and reformed a pellet (approximately 30 seconds).

[0135] Once beads had re-pelleted on the magnet and while keeping the sample tubes in the MPC on the benchtop, the tube was rotated again by 180° (returning tube to original position) allowing the beads to migrate to the magnet and reform a pellet (approximately 30 seconds).Attorney Docket No.: P24-157-SEC-WO01

[0136] The wash solution was removed and discarded, taking care not to disturb the bead pellet.

[0137] After removing the wash solution, the samples tubes were capped and spun to collect any remaining wash solution at bottom of tube.

[0138] The sample tubes were again placed onto MPC and beads allowed to migrate to the magnet followed by removal of the wash solution.

[0139] The tubes were removed from the MPC and the bead pellet resuspended in 20pl DEPC-treated water by carefully pipetting up and down approximately 8 times.

[0140] The resuspended beads were incubated for about 5 minutes at room temperature.

[0141] After incubation, the sample tubes were briefly spun down on a microcentrifuge. The sample tubes were the placed on a MPC until the beads migrated to the magnet and until solution was completely clear (approximately 1-2 minutes).

[0142] The supernatant (containing sample mRNA) was carefully transferred to a new 1 ,5m L Lobind tube or equivalent.

[0143] To the 1.5mL tube containing purified reverse-transcribed sample, reagent were added for the second ligation step, following the instructions in Table 10.Table 10: Second ligation reaction mixture

[0144] The second ligation reaction mixture was and spun down the sample tube using a microcentrifuge to collect the sample mixture at the bottom of the tube followed by incubation for 10 minutes at room temperature.

[0145] 16pl RNACIean beads were added to each sample and mixed thoroughly.Attorney Docket No.: P24-157-SEC-WQ01

[0146] Sample and bead mixtures were incubated at room temperature on a sample rotator (e.g., Hula mixer) for 5 minutes.

[0147] After incubation, the tubes were spun down on a mini / microcentrifuge to collect the sample at the bottom of the tubes.

[0148] The sample tubes were placed on MPC for the beads to migrate to magnet and until the solution was completely clear (approximately 1-2 minutes).

[0149] The clear supernatant was removed and discarded without disturbing the bead pellet.

[0150] The sample tubes were removed from the MPC and 150pl of Wash Buffer (WSB)added to the sample tubes, beads completely dislodged and resuspended and then spun down to collect the sample at the bottom of the tube.

[0151] The sample tubes were placed on a MPC for beads to migrate to magnet and until solution is completely clear. The clear supernatant was then removed taking care not to disturb the bead pellet.

[0152] A second wash was performed.

[0153] After removing the second wash the tubes were spun down on a mini / microcentrifuge to collect residual wash buffer at the bottom of the tube.

[0154] The sample tubes were placed on a MPC for the beads to migrate to magnet and residual wash removed.

[0155] The sample tubes were removed from the MPC and 13pl Elution Buffer (REB) added to the sample pellet. The bead pellet was completely dislodged from the tube, resuspended and spun down on a microcentrifuge to collect the sample at the bottom of the tube.

[0156] The samples were incubated for 10 minutes at room temperature.

[0157] After incubation, sample tubs were briefly spun down on a mini / microcentrifuge, placed on a MPC for beads to migrate to magnet and until solution is completely clear (approximately 1-2 minutes).

[0158] The ~13pl of clear supernatant (containing prepared mRNA library)was transferred to a new 1 ,5m L Lobind tube.

[0159] The concentration of 1 pl of sample library was assayed on a Qubit Fluorometer using the Qubit HS DNA Assay kit (Thermo Fisher). TheAttorney Docket No.: P24-157-SEC-WO01 prepared Direct RNA-sequencing Library was then ready for loading onto a FLO-MIN004RA (RNA) flow cell for sequencing.

[0160] After loading flow cell, a 24 hour sequencing run was performed on a GridlON instrument (Oxford Nanopore technologies, Lexington, MA) using default parameters for FLO-MIN004RA.Example 5: Poly(A)-tail Length Estimation

[0161] Poly(C)-tailing is performed on sample(s) of Poly(A)-containing mRNA, followed by library preparation (ONT direct RNA-sequencing kit) using a custom designed reverse transcription adapter (C-RTA) followed by direct RNA-sequencing, as described above. This example will quantify the Poly(A) tail lengths of individual mRNA molecules and the distribution of Poly(A) tail lengths of mRNAs in the sample.

[0162] Libraries were prepared as described, above, and direct sequenced. Data was analyzed to measure the length of Poly(A) tail for each molecule identified as containing a Poly(A) tail. Shown in Fig. 1 is a distribution (histogram) of Poly(A) tail lengths for sequenced mRNA sample and statistics table for Poly(A) tail length.Example 6: Fraction of mRNA Molecules Containing Poly(A) Tails

[0163] Poly(C)-tailing is performed on a mixture of Poly(A)-containing and non-Poly(A)-containing mRNAs, followed by library preparation (ONT Direct RNA-sequencing kit) using a custom designed reverse transcription adapter (C-RTA), then Direct RNA-sequencing, as described above. This example will identify non-Poly(A)-containing individual mRNA molecules as well as Polypcontaining individual mRNA molecules, and will quantify the proportion of contaminating non-Poly(A)-containing mRNA vs. Poly(A)-containing mRNAs in a sample.

[0164] Libraries were prepared as described, above, and direct sequenced. Data was analyzed to identify reads containing and not containing Poly(A) tails. Fig. 2 shows a table displaying the expected percentage of Poly(A)- containing mRNAs for two analyzed samples, the estimated percentage of Poly(A)-containing mRNAs for these two samples, and the lower and upperAttorney Docket No.: P24-157-SEC-WO01 estimate bounds (95% confidence interval) for percentage of Polypcontaining mRNAs for these two samples.Example 7: Full Length mRNA Length Estimation

[0165] Poly(C)-tailing is performed on sample of mRNA, followed by library preparation (ONT Direct RNA-sequencing kit) using a custom designed reverse transcription adapter (C-RTA), then Direct RNA-sequencing, as described above. This example will quantify the full length of individual mRNA molecules and the distribution of the full lengths of mRNAs in a sample.

[0166] Libraries were prepared as described, above, and direct sequenced. Data was analyzed to calculate full length of sequencing reads. Fig. 3 shows a distribution (histogram) of full mRNA lengths for sequenced mRNA sample.Example 8: Contaminating dsRNA Detection

[0167] Poly(C)-tailing is performed on sample of mRNA containing dsRNA contamination, followed by library preparation (ONT Direct RNA-sequencing kit) using a custom designed reverse transcription adapter (C-RTA), then Direct RNA-sequencing, as described above. This example will identify contaminating dsRNA molecules and quantify the proportion of contaminating dsRNA in a sample.

[0168] Libraries were prepared as described, above, and direct sequenced. Data was analyzed to identify mRNA molecules with sequences aligning to both the forward direction and the reverse complement (antisense, dsRNA) of the reference sequence. Shown in Fig. 4 are read coverage maps for the reference mRNA sequence of sense and antisense reads for sequenced mRNA sample. Antisense reads indicate potential for dsRNA presence.

Claims

Attorney Docket No.: P24-157-SEC-WO01What is Claimed is:1 . A method of adapting a mixed population of polyadenylated and nonpolyadenylated RNA molecules for analysis, said method comprising: a) contacting the mixed population of RNA molecules with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter-tagged cytidine-tailed RNA molecules, said sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and b) isolating the adapter-tagged cytidine-tailed RNA molecules.

2. The method of Claim 1 , wherein said isolated adapter-tagged cytidine- tailed RNA molecules are subject to direct sequencing or indirect sequencing.

3. The method of Claim 1 , wherein the number of cytidine residues added to the RNA molecule is from 1 to about 200, 2 to about 100, 3 to about 50, 4 to about 20 or 5 to about 12.

4. The method of Claim 1 , wherein the length of the tail sequence comprising guanine and / or inosine bases on the RNA sequencing adapter is from 3 nt to 15 nt or 3 nt to 9 nt in length.

5. The method of Claim 1 , wherein said ligase is selected from the group consisting of T4 DNA ligase, T3 DNA ligase, T7 DNA ligase, E. coli DNA ligase and T4 RNA Ligase 2.

6. The method of Claim 5, wherein said ligase is T4 DNA ligase.

7. The method of Claim 1 , wherein said cytidine tailing enzyme is a Poly(A) Polymerase or a Poly(U) Polymerase.Attorney Docket No.: P24-157-SEC-WO018. The method of Claim 7, wherein said cytidine tailing enzyme is selected from the group consisting of E. coli Poly(A) Polymerase I, Poly(A) Polymerase and S. pombe Poly(U) Polymerase.

9. The method of Claim 7, wherein said cytidine tailing enzyme is E. coli Poly(A) Polymerase I.

10. The method of Claim 1 , wherein said isolated adapter-tagged cytidine- tailed RNA molecules are sequenced and analyzed to determine the length of Poly(A) tails on the polyadenylated RNA molecules in the mixed population of RNA molecules.11 .The method of Claim 1 , wherein said isolated adapter-tagged cytidine- tailed RNA molecules are sequenced and analyzed to identify the Poly(A) and non-Poly(A) tailed RNA molecules in a sample and quantify the proportion of Poly(A) to non-Poly(A) in the mixed population of RNA molecules.

12. The method of Claim 1 , wherein said isolated adapter-tagged cytidine- tailed RNA molecules are sequenced and analyzed to estimate the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules.

13. The method of Claim 1 , wherein said isolated adapter-tagged cytidine- tailed RNA molecules are sequenced and analyzed to identify contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNA in the mixed population of RNA molecules.

14. The method of anyone of Claims 10 - 13, wherein said analysis involves one or more of applying an algorithm to identify the presence or absence of a poly(A) tail for each read in a generated dataset; and / or applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset; and / orAttorney Docket No.: P24-157-SEC-WO01 applying an algorithm to calculate the poly(A) tail length of a poly(A) region given the starting and ending locations of the poly(A) tail region for each read identified as containing a poly(A) tail in a generated dataset; and / or applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence to estimate the full length of the mRNA for each read in a generated dataset; and / or applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence in both the forward and reverse direction to identify antisense mRNAs.

15. A method of determining the length of polyadenylated (poly(A)) tails on RNA molecules in a sample, the method comprising: a) contacting a population of RNA molecules with poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’ -end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adapter-tagged cytidine-tailed RNA molecules, said sequencing adapter having a tail sequence comprising guanine and / or inosine bases; and b) isolating the adapter-tagged cytidine-tailed RNA molecules and then direct or indirect sequencing said adapter-tagger cytidine-tailed RNA molecules; and c) determining one or both of: i) poly(A) tail lengths of the individual RNA molecules, and / or ii) the distribution of Poly(A) tail lengths of RNA molecules in the sample by applying an algorithm to identify the presence or absence of a poly(A) tail for each read in a generated dataset and / or applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset and / or applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset.Attorney Docket No.: P24-157-SEC-WO0116. A method of identify both poly(A) and non-poly(A) tailed RNA molecules in a sample, said method comprising: a) contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adaptor-tagged RNA molecules, said sequencing adaptor having a tail sequence comprising guanine and / or inosine bases; and b) isolating the adapter-tagger RNA molecules and then direct or indirect sequencing said adapter-tagger molecules; and c) determining the one or both of: i) the proportion of poly(A) containing RNA molecules to non-poly(A) containing molecules, and ii) identifying the non-poly(A) containing RNA molecule in the sample by applying an algorithm to estimate the starting and ending locations of the poly(A) tail region in the raw signal of for each read identified as containing a poly(A) tail in a generated dataset.

17. A method to estimate the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules, said method comprising: a) contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adaptor-tagged RNA molecules, said sequencing adaptor having a tail sequence comprising guanine and / or inosine bases; and b) isolating the adapter-tagger RNA molecules and then direct or indirect sequencing said adapter-tagger molecules; and c) determining the full length of individual molecules and the distribution of lengths of the RNA molecules in the mixed population of RNA molecules by applying a series of algorithms to convert theAttorney Docket No.: P24-157-SEC-WO01 raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence to estimate the full length of the mRNA for each read in a generated dataset.

18. A method to identify contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNA in the mixed population of RNA molecules, said method comprising: a) contacting a mixed population of RNA molecules both with and without poly(A) tails with a cytidine tailing enzyme to add one or more cytidine residues at the 3’-end of the RNA molecules to produce cytidine-tailed RNA molecules; and, ligating the tailed RNA molecules to a sequencing adapter with a ligase to produce adaptor-tagged RNA molecules, said sequencing adaptor having a tail sequence comprising guanine and / or inosine bases; and b) isolating the adapter-tagger RNA molecules and then direct or indirect sequencing said adapter-tagger molecules; and c) identifying contamination of double strand RNA (dsRNA) molecules and quantify the proportion of contaminating dsRNA in the mixed population of RNA molecules by applying a series of algorithms to convert the raw signal into the sequence of the mRNA, followed by mapping the sequence of the read to a reference sequence in both the forward and reverse direction to identify antisense mRNAs which is indicative of contamination of dsRNA.

Citation Information

Patent Citations

  • Methods, Reagents and Kits for Detection of Nucleic Acid Molecules

    US20080300142A1

  • Polynucleotide barcodes for long read sequencing

    US20210198730A1

  • Full-length RNA sequencing

    US20220002797A1

  • Application of Immobilized Enzymes for Nanopore Library Construction

    US20220090056A1

  • Poly(A) Tail Length Measurement by PCR

    US20110059453A1