Method for preparing RNA samples for sequencing and kit therefor

CircAID-p-seq addresses the preservation of 3'-P/cP groups in RNA sequencing by phosphorylating and circularizing RNA molecules, providing accurate and efficient sequencing without PCR, suitable for nanopore technology and biomarker detection.

JP7733642B2Active Publication Date: 2025-09-03IMMAGINA BIOTECH SRL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022515626
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-09-09
Filing Date
2020-09-04
Publication Date
2025-09-03
Estimated Expiration
2040-09-04

AI Technical Summary

Technical Problem

Current RNA sequencing methods fail to preserve the 3'-terminal phosphate or 2',3'-cyclic phosphate groups, leading to inaccurate results, high background noise, and PCR amplification biases, which are crucial for understanding RNA-protein interactions and biomarker detection.

Method used

A method called CircAID-p-seq that phosphorylates both ends of RNA molecules, ligates them to random RNA linkers, forms circular molecules, and undergoes reverse transcription rolling-circular amplification to generate cDNA suitable for nanopore sequencing, eliminating the need for PCR.

Benefits of technology

This method preserves the 3'-P/cP signature, reduces bias, and enables efficient sequencing of biologically relevant RNA species, particularly suitable for low-input samples and resource-limited settings, enhancing the detection of ribosome footprints and biomarkers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007733642000013
    Figure 0007733642000013
  • Figure 0007733642000014
    Figure 0007733642000014
  • Figure 0007733642000015
    Figure 0007733642000015
Patent Text Reader

Abstract

1. A method for preparing at least one RNA molecule contained within a biological sample for sequencing, comprising the steps of: (i) obtaining a biological sample containing at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end; (ii) phosphorylating the 5' end of at least one RNA molecule, thereby introducing a phosphate group into the 5' end of the at least one RNA molecule to obtain at least one RNA molecule phosphorylated at both ends; (iii) ligating the 3′ end of at least one phosphorylated RNA molecule to the 5′ end of a random RNA linker having —OH groups at both ends to obtain at least one first ligation product; (iv) self-ligating at least one first ligation product to form at least one circular RNA molecule that is mixed with the linear RNA molecule; (v) digesting the linear RNA molecules; (vi) subjecting the at least one circular RNA molecule to reverse transcription rolling-circular amplification to obtain at least one single-stranded cDNA molecule having at least one copy, preferably 2 to 500 copies, of the at least one RNA molecule; Including, wherein said at least one single-stranded cDNA molecule is suitable for sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present specification relates to novel methods for preparing RNA samples for sequencing and kits for carrying out such methods. [Background technology]

[0002] RNA-protein interactions play a fundamental role in regulating crucial aspects of cell biology, from mRNA transcription, pre-mRNA splicing and RNA signaling functions to translation and protein localization. 1 Given the importance of understanding such biological processes, chemical labeling of RNA and proteins 2 Whole-genome high-throughput sequencing of RNA footprints from 3-6 Various efforts have been made to develop methods to study and characterize these interactions. However, sequencing methods generally suffer from various limitations during sample preparation, such as extensive manipulation steps, PCR amplification, and the inability to selectively capture RNA sequences with 3'-terminal phosphate or 2',3'-cyclic phosphate (3'-P / cP) groups, thus reducing the accuracy of the results. 7 This results in library cross-reactivity with undesired RNA targets, high background noise, and poor library quality, preventing important biological information from 3'-P / cP-terminated RNA products. 3'-P / cP is produced by enzymatic cleavage, and 3'-P / cP RNA is a key component in many pathologies (e.g., cancer and amyotrophic lateral sclerosis). 8,9 ), biological processes (e.g., the protein unfolding response 10 , stress granule production 8 , RNA metabolism 11 , rRNA and tRNA biosynthesis 12 and mRNA splicing 13 ), and biological functions (e.g., neuronal survival). 14 and inflammatory responses 15Although the 3'-terminal phosphate signature is an important functional marker, most sequencing pipelines do not preserve this chemical signature during library preparation. Only a few methods are available for the detection of 3'-P or 3'-cP, but they only allow the indirect detection of 3'-P. 16 , or exclusively selective for 3'-cP 16,17 Furthermore, these protocols are labor- and time-consuming and involve a PCR amplification step, which may therefore result in uneven sequence coverage or sequencing errors (e.g., in internal repeat regions).

[0003] From a technical point of view, many RNA footprinting techniques are aimed at detecting RNA-protein interactions (large RNA-protein complexes). 18 or RNA small molecule interactions 19 An experimental setting highly affected by the lack of available library preparation protocols capable of selectively capturing 3'-P ends is ribosome profiling (Ribo-seq), an RNA footprinting method based on deep sequencing of 25-35 nt long ribosome-protected fragments (RPFs), i.e., mRNA fragments produced after nuclease digestion of unprotected single-stranded RNA. As it provides information on the position of ribosomes along the transcript captured at a specific moment, this technique represents a powerful method for studying the biology of protein synthesis. 20 Current protocols for ribosome profiling involve many sequential steps and are based on the Illumina sequencing platform. In particular, after isolation of RPFs, two alternative library preparation workflows are available: (i) a workflow based on sequential steps of adapter ligation at the 3' end of the RPFs, cDNA synthesis, circularization, and PCR amplification, with a total of four gel extraction steps; 21 (ii) Commercially available products include sequential steps of RNA 3' polyadenylation, template switching, and cDNA synthesis by PCR amplification. 22Ligation-independent workflows for low-input material are currently available, which involve PCR amplification and require a gel extraction step. The main drawbacks of the available protocols are manifested by (i) PCR amplification bias and (ii) the lack of preservation of 3'-P / cP ends (which provide an effective digestion signature) and the subsequent under-representation of RNA species with 3'-P / cP (representing actual RPFs) in sequencing datasets. In fact, both workflows require adapter ligation. 3 Alternatively, a dephosphorylation step is required prior to polyadenylation, which reduces the level of specificity in the ligation reaction and results in the capture of any short RNA molecule with an -OH group at the 3' end.

[0004] Furthermore, recent studies have revealed significant differences in the relative abundance and type of small RNA populations, e.g., tRNA-derived RNAs, piwi-interacting RNAs, and Y RNAs, in biological fluids (umbilical cord plasma, bronchoalveolar lavage fluid, adult plasma, parotid saliva, follicular fluid, serum, amniotic fluid, seminal plasma, urine, bile, submandibular / sublingual saliva, and cerebrospinal fluid). Importantly, some of them are known to have 3'P or 2',3'-cP residues and are associated with cancer, neurological disorders, and immunological disorders. 23 In this clinical scenario, these RNA species may have a potential role as biomarkers and predictive and / or prognostic significance in patient stratification. 24,25 . Summary of the Invention [Problem to be solved by the invention]

[0005] Therefore, there is a need for new methods of preparing RNA samples for sequencing that do not have the drawbacks of known methods. [Means for solving the problem]

[0006] It is an object of the present disclosure to provide novel methods for preparing RNA samples for sequencing and kits for carrying out such methods.

[0007] According to the present invention, the above objects are achieved by the subject matter expressly recalled in the following claims, which are to be understood as forming an integral part of this disclosure.

[0008] The present invention relates to a method for preparing at least one RNA molecule contained in a biological sample for sequencing, comprising the steps of: (i) obtaining a biological sample containing at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end; (ii) phosphorylating the 5' end of at least one RNA molecule, thereby introducing a phosphate group into the 5' end of the at least one RNA molecule to obtain at least one RNA molecule phosphorylated at both ends; (iii) ligating the 3′ end of at least one phosphorylated RNA molecule to the 5′ end of a random RNA linker having —OH groups at both ends to obtain at least one first ligation product; (iv) self-ligating at least one first ligation product to form at least one circular RNA molecule that is mixed with the linear RNA molecule; (v) digesting the linear RNA molecule; and (vi) subjecting the at least one circular RNA molecule to reverse transcription rolling-circular amplification to obtain at least one single-stranded cDNA molecule having at least one copy, preferably 2 to 500 copies, of the at least one RNA molecule; Including, wherein said at least one single-stranded cDNA molecule is suitable for sequencing, preferably single-molecule sequencing.

[0009] This method does not require PCR and can be applied to any 3'-P / cP-terminated RNA footprint. The goal of this method (called CircAID-p-seq) is a rapid cDNA sequencing protocol for low-input biological samples optimized for nanopore sequencing. This method overcomes some of the limitations that have traditionally plagued RNA footprinting, such as time-consuming protocols and PCR bias, and provides a powerful pipeline for deep sequencing of 3'-P / cP-terminated RNA fragments using the Oxford Nanopore platform, thus enabling real-time single-molecule detection of biologically relevant RNA species.

[0010] In a further embodiment, the present invention relates to a kit for carrying out a method (disclosed herein) of preparing at least one RNA molecule contained in a biological sample for sequencing, wherein the kit comprises random RNA linkers, and a first ligase enzyme, an exoribonuclease, and optionally a second ligase enzyme, wherein: (i) the random RNA linker has –OH groups at both ends; (ii) the ligase enzyme is suitable for ligating the 3' end of an RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end and a phosphate group at its 5' end to the 5' end of a random RNA linker having a hydroxyl group at its 3' end; (iii) the exoribonuclease is suitable for enzymatically digesting linear RNA molecules; and (iv) A second ligase enzyme is suitable for circularizing the ligation product obtained by ligating the random RNA linker to the RNA molecule.

[0011] The invention will now be described in detail, purely by way of illustration and non-limiting example, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0012] [Figure 1] A) TBE-urea PAGE analysis of fragments generated by RNAse I digestion with and without polyadenylation. B) Schematic of the CircAID-p-seq workflow. C) TBE-urea PAGE analysis of all CircAID-p-seq steps. [Figure 2] Direct cDNA sequencing of the circGFP-linkerR library. A) Length distribution of sequencing reads; B) Representative consensus sequence. Single base-called reads were split into their individual repeats, which were then aligned against each other to generate a consensus sequence. [Figure 3] A) Representative photograph of GFP-transfected HEK293T cells. B) Length distribution of GFP fragments detected by BLASTn. C) BLASTn alignment of sequencing reads against the reference GFP sequence. [Figure 4] Representative consensus sequences produced by 2-repeat (top) and 3-repeat (bottom) sequences obtained from two different reads. [Figure 5] Nucleotide sequence. DETAILED DESCRIPTION OF THE INVENTION

[0013] In the following description, numerous specific details are provided to provide a thorough understanding of the embodiments. The embodiments may be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the embodiments.

[0014] References throughout this specification to "one embodiment" or "one embodiment" mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrase "in one embodiment" or "in one embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0015] The headings provided herein are for convenience only and do not interpret the scope or meaning of the embodiments.

[0016] The present invention relates to a novel method for preparing at least one RNA molecule contained within a biological sample for sequencing, comprising the following steps: (i) obtaining a biological sample containing at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end; (ii) phosphorylating the 5' end of at least one RNA molecule, thereby introducing a phosphate group into the 5' end of the at least one RNA molecule to obtain at least one RNA molecule phosphorylated at both ends; (iii) ligating the 3′ end of at least one phosphorylated RNA molecule to the 5′ end of a random RNA linker having —OH groups at both ends to obtain at least one first ligation product; (iv) self-ligating at least one first ligation product to form at least one circular RNA molecule that is mixed with the linear RNA molecule; (v) digesting the linear RNA molecule; and (vi) subjecting the at least one circular RNA molecule to reverse transcription rolling-circular amplification to obtain at least one single-stranded cDNA molecule having at least one copy, preferably 2 to 500 copies, of the at least one RNA molecule; Including, wherein at least one single-stranded cDNA molecule is suitable for sequencing, preferably single-molecule sequencing. More preferably, sequencing is performed by the Oxford Nanopore Sequencing platform (Nanopore Sequencing).

[0017] In one embodiment, the biological sample can be selected from eukaryotes (single-celled organisms such as plants, animals, fungi, and protists), viral or prokaryotic cell lysates, tissues (including blood and cells from biopsies, in vitro and ex vivo), biological fluids (umbilical cord plasma, bronchoalveolar lavage fluid, adult plasma, parotid saliva, follicular fluid, serum, amniotic fluid, seminal plasma, urine, bile, submandibular / sublingual saliva, cerebrospinal fluid), 3D cell cultures.

[0018] In one embodiment, at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end is produced by treating a biological sample with an endoribonuclease, exoribonuclease, ribozyme, or toxin capable of cleaving mRNA, tRNA, snRNA, snoRNA, Y RNA, lncRNA, piRNA, siRNA, viral RNA (from positive-sense RNA viruses, negative-sense RNA viruses, reverse-transcribed viruses, and other RNA species produced by viruses), or rRNA.

[0019] In one embodiment, at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end is physiologically or pathologically present in a biological sample as a result of the action of an endoribonuclease, exoribonuclease, ribozyme, or toxin capable of cleaving mRNA, tRNA, snRNA, snoRNA, Y RNA, lncRNA, piRNA, siRNA, viral RNA (from positive-sense RNA viruses, negative-sense RNA viruses, reverse-transcribed viruses, and other RNA species produced by viruses) or rRNA present in the biological sample.

[0020] In one embodiment, the endoribonuclease is preferably RNase A; RNase T1; RNase T2; RNase I; S7 micrococcal nuclease; staphylococcal nuclease; RNAse L; angiogenin; colicin E5; tRNA-splicing endonuclease (SE2, SEN34); ferredoxin-like Cas6 and ferredoxin-like CasE; IRE1; poly(U)-specific endoribonuclease (PP11); Las1; RtcA; type IB topoisomerase; Cue2 endonuclease. 26 and Cas proteins.

[0021] In one embodiment, the exoribonuclease is preferably represented by USB1.

[0022] In one embodiment, the ribozyme is preferably selected from a hammerhead shark ribozyme, a hairpin ribozyme, a hepatitis delta ribozyme, and a Varkud satellite (VS) ribozyme.

[0023] In one embodiment, the toxin is preferably selected from colicin D and colicin E5, alpha-sarcin, zymosin, PaT, MazF, ChpBK, prrC.

[0024] In one embodiment, at least one RNA molecule to be sequenced is single-stranded.

[0025] In one embodiment, the at least one RNA molecule is contained in the biological sample within a concentration comprised between 10 pM and 100 μM, preferably between 1 nM and 10 μM.

[0026] In one embodiment, the method comprises the further step (vi) of generating a complementary cDNA strand of at least one single-stranded cDNA molecule to obtain at least one double-stranded cDNA molecule.

[0027] In one embodiment, the phosphorylation step (ii) is carried out using a phosphorylating enzyme selected from T4 PNK 3'minus, T4 PNK and a recombinant version of T4 PNK (eg Optikinase™).

[0028] In one embodiment, the ligation step (iii) is carried out using a first ligase enzyme selected from RtcB, Archease, Arabidopsis Thaliana tRNA ligase, and a eukaryotic tRNA ligase.

[0029] In one embodiment, the self-ligation step (iv) is carried out using a second ligase enzyme selected from T4 Rnl1, T4 Rnl2, T4 Rnl2tr, T4 Rnl2 K227Q, Mth Rnl, and an ATP-independent ligase that catalyzes intramolecular ligation (e.g., circligase™, circligase II™).

[0030] In one embodiment, the digestion step (v) is carried out using a 5'-3' exoribonuclease or a 3'-5' exoribonuclease, preferably using RNAseR.

[0031] In one embodiment, the reverse transcription rolling circular amplification step (vi) is carried out using modified M MLV-RT (Moloney Murine Leukemia Virus Reverse Transcriptase) and AMV-RT (Avian Myeoloblastosis Virus Reverse Transcriptase) (preferably selected from Maxima H minus™ Superscript™ I-II-III-IV, Sunscript™).

[0032] In one embodiment, the step (vi) of generating the complementary cDNA strand is carried out using a DNA polymerase enzyme selected from Taq polymerase with 5'-3' exonuclease activity and the Gubler-Hoffman method (e.g., Platinum II Taq Hot Start DNA Polymerase™, AB Taq™ PrimeScript™, NEBNext® Ultra™ II Non-Directional RNA Second Strand Synthesis).

[0033] In a further embodiment, the present invention relates to a kit for carrying out a method (disclosed herein) of preparing at least one RNA molecule contained in a biological sample for sequencing, the kit comprising random RNA linkers, and a first ligase enzyme, an exoribonuclease, and optionally a second ligase enzyme; (i) the random RNA linker has –OH groups at both ends; (ii) the ligase enzyme is suitable for ligating the 3′ end of an RNA molecule having a phosphate group or a 2′,3′-cyclic phosphate group at its 3′ end and a phosphate group at its 5′ end to the 5′ end of the random RNA linker; (iii) the exoribonuclease is suitable for enzymatically digesting linear RNA molecules; and iv) The second ligase enzyme is suitable for circularizing the ligation products obtained by ligating the random RNA linkers to the RNA molecules (i.e., ligating the 5' end of the ligation product having a phosphate at the 5' end to the 3' end of the ligation product having an -OH group at the 3' end).

[0034] In one embodiment, the first ligase enzyme is selected from RtcB, Archease, Arabidopsis Thaliana tRNA ligase, and a eukaryotic tRNA ligase.

[0035] In one embodiment, the second ligase enzyme is selected from T4 Rnl1 T4 Rnl1, T4 Rnl2, T4 Rnl2tr, T4 Rnl2 K227Q, Mth Rnl, and an ATP-independent ligase that catalyzes intramolecular ligation (e.g., circligase™, circligase II™).

[0036] In one embodiment, the exoribonuclease is RNaseR.

[0037] In one embodiment, the kit further comprises (a) a kinase, and / or (b) an endoribonuclease, ribozyme, or toxin capable of cleaving mRNA, tRNA, snRNA, snoRNA, Y RNA, lncRNA, piRNA, siRNA, viral RNA (from positive-sense RNA viruses, negative-sense RNA viruses, reverse-transcribed viruses, and other RNA species produced by viruses), or rRNA. Preferably, the kit further comprises an endoribonuclease, wherein the endoribonuclease is RNAse I.

[0038] In one or more embodiments, the random RNA linker has a length comprised between 50 and 500 nucleotides.

[0039] In one or more embodiments, the random RNA linker has a minimum free energy between -3 and -150 kcal / mol. Preferably, each random RNA linker has a minimum free energy between -6 kcal / mol and -24 kcal / mol and is designed to have no significant secondary structure. Some secondary structure is tolerated in the interior of the sequence, but not at the 5' / 3' ends. The minimum free energy can be calculated using software available to those skilled in the art.

[0040] In one or more embodiments, the random RNA linker has the nucleotide sequence set forth in SEQ ID NO:3.

[0041] Random RNA linkers can be chemically synthesized or transcribed and purified in vitro according to the common general knowledge of a specialist in the field.

[0042] The 5'-OH group of the random linker can be produced chemically or enzymatically. If produced enzymatically, the 5'-OH can be obtained by (i) a ribozyme acting in cis (encoded by an in vitro transcribed sequence) or trans (acting on an in vitro transcribed sequence), ii) an enzyme that releases a 5'-OH group, such as calf intestinal phosphatase, or (iii) the catalytic activity of a toxin selected from colicin D and colicin E5, α-sarcin, zymosin, Pichia acaciae killer toxin (PaT), MazF, ChpBK, and prrC.

[0043] The random RNA linker can comprise at least one, preferably 1 to 109, nucleotides modified with at least one of the following modifications: LNA, PNA, 2-aminopurine, 2,6-diaminopurine (2-amino-dA), 6mA, 5-bromodU, inverted dT, 5-methyldC, 8-aza-7-deazaguanosine, 5-hydroxybutynyl-2'-deoxyuridine, 5-nitroindole, 2'-O-methyl A, 2'-O-methyl G, 2'-O-methyl C, 2'-O-methyl U, 2'-fluorine A, 2'-fluorine C, 2'-fluorine G, 2'-fluorine U, 2-methoxyethoxy A, 2- Methoxyethoxy MeC, 2-methoxyethoxy G, 2-methoxyethoxy T, 5-bromo dU, 2-aminopurine, inverted dT, 2,6-diaminopurine, deoxyuridine, inverted dideoxy-T, 5-methyl dC, dideoxy-C, deoxyinosine, universal bases including: 5-nitroindole, morpholino, 2'-0-methyl RA base, iso-dC, iso-dG, ribonucleotides, threose nucleotide analogs, protein nucleotide analogs, glycoic nucleotide analogs, locked nucleotide analogs, chain termination The random RNA linker may contain 1 to 25 modified nucleotides within the first 25 bases from the 5' end and the last 25 bases from the 3' end.

[0044] We developed a library preparation method for nanopore sequencing of short RNA molecules with 3'-P signatures and validated it in the setting of ribosome profiling (Ribo-seq). In particular, CircAID-p-seq is a highly sensitive RT-RCA-based method that enables the detection of low-abundance short RNA molecules and is therefore potentially applicable to single-cell techniques.

[0045] This method is based on standard protocols 3 This method significantly shortens the time required for Ribo-seq library preparation, which currently takes longer than the one week required for conventional methods, and significantly reduces the number of technical steps (dephosphorylation, gel extraction, purification), thus lowering the possibility of introducing bias. This method enables any RNA footprinting study that uses enzymatic cleavage to release 3'-P / cP ends. Furthermore, studies on cancer and neurodegeneration, autoimmune and infectious disorders, and multiple cellular functions reported to involve 3'-P / cP-terminated RNA molecules are particularly benefited by this method. In particular, this method is suitable for characterizing the endonucleolytic activity of specific enzymes, ribozymes, or toxins. 27 , RNA editing CRISPR-Cas 28 Finally, the methods disclosed herein enable rapid sequencing pipelines without the need for expensive laboratory equipment, even in resource-limited settings. [Example]

[0046] result Cellular RNA can have hydroxyl (-OH), phosphate (-P), or 2',3'-cyclic phosphate (-cP) groups at its termini. RNA cleavage by many endoribonucleases often releases 3'-P or 3'-cP ends, which are not compatible substrates for ATP-dependent ligases (e.g., T4 RNA ligase). A methodologically relevant situation involving the use of endoribonucleases to cleave RNA strands and subsequent ligation events is represented by Ribo-seq for RNA footprinting, which is based on the following steps: (i) cell lysis, (ii) endonuclease (e.g., RNase I) digestion of ssRNA, (iii) recovery of 25-35 nt-long fragments (bona fide RPFs), (iv) library preparation, (v) deep sequencing, and (vi) final alignment to a reference protein-coding transcriptome.

[0047] To clarify the fraction of actual RPFs from the total population of fragments generated by RNase I cleavage, we utilized 3' polyadenylation. The results show that approximately 50% of the 25-35 nt fragments obtained from cultured cells (MCF7) reacted in the polyadenylation reaction (Figure 1A). This indicates that the size-selected RPFs are contaminated with RNA species with 3'-OH ends, which can be captured by standard ligation processes and result in higher background noise.

[0048] To overcome the limitations of currently available Ribo-seq library preparation strategies, we sought to develop a method that would (i) preserve the 3'-P / cP signature and (ii) be independent of a PCR amplification step. To achieve these goals, we developed an enzyme (RtcB ligase) capable of (i) ligating 5'-OH to 3'-P / cP ends. 29,30and (ii) a linker suitable for PCR-free nanopore sequencing. Specifically, we designed a method focused on direct cDNA nanopore sequencing. To provide proof-of-concept for the feasibility of this method, we first used a 30-nt synthetic RNA fragment (5'P-GFP-3'P) bearing -P groups at both the 5' and 3' ends as a surrogate for cell-derived and 5'-phosphorylated RPF. The GFP fragment has the nucleotide sequence shown in SEQ ID NO: 1.

[0049] In the cirAID-p-Seq approach (Figure 1B), we first ligated the 5'-OH end of a random RNA linker to the 3'-P end of the 5'P-GFP-3'P fragment using RtcB ligase. The ligation product was then separated by TBE-urea PAGE, size-selected, and gel-purified for subsequent reactions (Figure 1C). Next, we circularized the 5'P-GFP-linkerR-3'OH product using T4 RNA ligase I. To confirm the presence of a circular RNA structure, we treated the circularized product with RNaseR, an exoribonuclease that digests all linear RNA instead of preserving circular RNA. After TBE-urea PAGE separation of the RNAseR reaction (Figure 1C), the circularized product (circGFP-linkerR) was detected at the expected molecular weight, thus confirming the stability and correct circularization of the construct. The circGFP-linkerR product was then gel purified and subjected to reverse transcription rolling-circular amplification (RT-RCA) 31,32 This resulted in the generation of long, multimeric single-stranded cDNA molecules (140–15,000 nt) with multiple copies of the inserted fragment. As a final checkpoint before sequencing, the RT-RCA products were separated by TBE-urea PAGE to confirm the presence of multimeric cDNA products within the expected size range (Figure 1C).

[0050] This method allows for the enrichment of 3'-P / cP-endowed RNA fragments, as this signature is essential for the efficiency of the overall protocol. The approach is compatible with direct cDNA nanopore sequencing without downstream PCR and, when combined with barcoded linkers, allows for multiplexed assays.

[0051] The library preparation methodology of the present disclosure aims to represent an initial PCR-free protocol for the selective incorporation of 3'-P / cP-tailed RNA fragments, suitable for nanopore sequencing.

[0052] Nanopore sequencing of short 3'-P-terminated RNA fragments Both of the above GFP-based libraries were sequenced on an Oxford Nanopore Technologies (ONT) MinION by using an R9.4 flow cell and 1D chemistry for direct cDNA sequencing.

[0053] We adapted the ONT protocol for cDNA sequencing by performing second (complementary) cDNA strand synthesis, followed by end repair and dA tailing. From sequencing a 60 ng library input, we obtained approximately 1.5 million "passed" reads (MinKNOWN base calls) in 2 hours, with an average base call quality score of 10.5 and less than 5% failed reads.

[0054] The read length distribution showed a major peak at approximately 300 nt, with approximately 10% of the reads spanning a broader range, from 1 KB to over 50 KB (Figure 2A). This indicates that the RT-RCA reaction is maximally efficient after two to three rounds of reverse transcription, but can produce up to 500 copies of the original template. After BLASTn alignment of "passed" reads against a reference GFP sequence, 100% of the reads appeared to contain at least one 30-mer GFP fragment. A representative alignment of 17 repeated GFP fragments within a sampled 2.5 KB read highlights the importance of repeats for computationally generating consensus sequences, whose accuracy (% identity to the original sequence) is proportional to the number of repeats (Figure 2B), which in turn depends on the efficiency of the RT-RCA reaction.

[0055] Overall, these results provide evidence that the library preparation method (i) can incorporate short synthetic RNA molecules that resemble endogenous cleaved RNAs with a 3'-P signature and (ii) can be effectively applied to the ONT sequencing platform.

[0056] circAID-p-seq allows detection of ribosome footprints after RNase I digestion Since circAID-p-Seq provides high sequencing depth (necessary for ribosome profiling experiments) and the applicability of this method to the ONT platform with a synthetic 30-mer GFP fragment was demonstrated, we sought to determine whether ribosome footprints from GFP-transfected HEK293T cells could be identified by this method.

[0057] We reasoned that transiently transfecting HEK293T cells with a GFP-overexpressing plasmid would have the advantages of (i) ensuring rapid identification of RPFs thanks to orthogonal reference sequences with well-defined open reading frames, and (ii) producing large amounts of recombinant protein with high footprint density on the GFP mRNA. The GFP protein appeared to be expressed at high levels 24 hours after transfection (Figure 3A). Cytoplasmic cell lysates were treated with RNase I to digest all RNA strands not protected by ribosomes to produce RPFs with 3'-P, which were then purified as described by Ingolia et al. 33 The RNA was purified according to the protocol described in the previous section. A total of 450 ng of purified RNA was used as input for library preparation, and the resulting cDNA library was sequenced on a MinION for 6 hours. The output produced sequences with a quality score of approximately 10, with less than 10% failed reads. BLASTn alignment of the MinKNOWN base-called reads to the reference GFP mRNA allowed for the detection of recurring GFP mRNA fragments. The length distribution of the GFP fragments ranged between 18 and 60 nt, with read accumulations of approximately 25 and 31 nt (Figure 3B), consistent with the standard RPF length distribution. 34 All GFP fragments mapped to the coding sequence, with no coverage of the 3' and 5' UTRs (Fig. 3C), suggesting that these GFP fragments are authentic RPFs and not footprints derived from non-ribosomal ribonucleoprotein complexes. 35 .

[0058] Furthermore, we observed that the resulting consensus sequences achieved excellent accuracy (96.5%) with 2 repeats and 100% consensus accuracy with 3 repeats, indicating that our library preparation strategy allows for the correct identification of RNA fragments contained within reads containing at least 3 repeats (Table 1; Figure 4).

[0059] JPEG0007733642000001.jpg54153

[0060] These results further confirm that CircAID-p-Seq produces ribosome profiling libraries suitable for the ONT platform and provide evidence that this protocol (CircAID-p-Seq) allows for the specific detection of actual ribosome footprints along transcripts.

[0061] material and method Ribosomal protected fragments and linkers The custom linker (having the nucleotide sequence set forth in SEQ ID NO: 3) was synthesized by IMMAGINA BioTechnology srl (Trento) and consisted of a 109-mer oligonucleotide with -OH groups at both ends.

[0062] Ribosome-protected fragments (RPFs) consisting of 30-mer oligonucleotides having the nucleotide sequence shown in SEQ ID NO: 1 with 5'-P and 3'-P were synthesized by Integrated DNA Technologies (Coralville) or produced in vitro.

[0063] In vitro produced RPFs were obtained from HEK293T (human embryonic kidney) cells (SIGMA, catalog number 12022001) transfected with a plasmid encoding GFP (pMAX_GFP™, Lonza catalog number V4XP-3024 - SEQ ID NO: 2). Cells were monitored for GFP expression 24 hours later by fluorescence microscopy (Olimpus DP70). Transfected cells were lysed by treatment with CHX (10 μg / mL, SIGMA catalog number 01810) for 5 minutes at 37°C. RPFs were produced by treating 0.3 AU 260 nM CHX-treated cell lysates with 2.25 U of RNAse I (Ambion, catalog number AM2295) in W-buffer (Immagina Biotechnology catalog number #RL001-4) for 45 minutes at room temperature (Clamer et al., 2018). 36 as described in

[0064] RNAse I digestion was stopped by adding 10 U of Superase Inhibitor (Thermo Scientific, Cat. No. AM2696) for 10 min on ice. After digestion, lysates were purified (Ingolia et al., 2009) 33 The samples were treated with 1% SDS (Sigma catalog no. 05030) and 0.1 mg proteinase K (Euroclone, catalog no. EMR022001) for 75 minutes at 37°C (as described in [Clamer et al., 2018]). Total RNA was extracted with acid-phenol:chloroform, pH 4.5 (Ambion, catalog no. AM9722). RNA was precipitated with isopropanol, air-dried, resuspended in 10 mM Tris-HCl (pH 8), and analyzed on a 15% TBE-urea polyacrylamide gel (Invitrogen, catalog no. EC6885BOX). 30-mer RPFs were size-selected and extracted from the gel (following the Ribolace protocol, Immagina Biotechnology) (Clamer et al., 2018). 36 .

[0065] Upon purification, the in vitro produced RPF fragments were subjected to 5' phosphorylation using T4 PNK 3'minus (NEB, Cat. No. M0236S) before capture with linker R.

[0066] RPF fragment-linker ligation The RPF fragment, phosphorylated at both the 5' and 3' ends, was ligated with Linker R (Immagina Biotechnology, catalog number #RLP001-1) using RtcB ligase (NEB, catalog number M0458S) according to the following reaction conditions: 90 pmol RPF, 30 pmol Linker R, 45 pmol RtcB ligase, 1X RtcB ligase buffer, 100 μM GTP, 1 mM MnCl2 (final volume 30 μL). The reaction was incubated at 37 °C for 2 h, and then the mixture was loaded onto a 15% acrylamide / 8 M urea precast gel (Invitrogen, catalog number EC6885BOX). The desired product (140 nt long) was purified by gel extraction to control the efficiency of the reaction. The gel purification step is not essential for the overall workflow.

[0067] Circularization and RNaseR treatment Circularization of the RtcB ligation product was carried out for 2 hours at 25°C in a total volume of 20 μL containing 10 U of T4 RNA Ligase 1 (NEB, Cat. No. M0204L), 1x Buffer T4 RNA Ligase, 20% PEG 8000, and 50 μM ATP. The circularization reaction was then incubated for 1 hour at 37°C with 20 U of RNase R (Lucigen, Cat. No. RNR07250) to remove any undesired products (i.e., linear RNA or contaminating products). The circularized RNA product was loaded onto a 15% acrylamide / 8M urea precast gel (Invitrogen, Cat. No. EC6885BOX) and purified by gel extraction to control the efficiency of the reaction. The gel purification step is not essential for the overall workflow.

[0068] Reverse transcription-rolling circle amplification (RT-RCA) and second strand synthesis RT-RCA was performed in 20 μL under the following conditions: 50 ng of circular RNA, 200 U of reverse transcriptase, 1× buffer RT, 0.5 mM dNTPs, 50 pmol of RT-RCA Rev primer (having the nucleotide sequence shown in SEQ ID NO: 5), 10% glycerol, using a primer annealing to the 3' region of the linker and Maxima H Minus reverse transcriptase (Thermo Fisher, catalog no. EP0752). The reaction was carried out at 42°C for 4 hours and then stopped at 70°C for 10 minutes. After cDNA synthesis, the circular RNA template was hydrolyzed by adding 0.1 N NaOH at 70°C for 10 minutes.

[0069] To generate the second strand from the single-stranded cDNA molecules, a single-cycle PCR was performed with Super AB Taq polymerase (AB Analitica catalog number 06-36-020) using an RT-RCA Fw primer (having the nucleotide sequence shown in SEQ ID NO: 4) annealing to the 5' region of the linker. The reaction contained 20 μl from the RT reaction, 1× buffer, 0.2 mM dNTPs, 2 mM MgCl2, 1.25 U Taq polymerase, and 50 pmol RT-RCA Fw primer (total volume 50 μl) and was subjected to the following program: initial denaturation at 95°C, followed by one cycle of 95°C for 30 seconds, 51°C for 30 seconds, and 70°C for 2 minutes. Double-stranded cDNA was purified using AMPure XP beads (Agencourt, catalog number A63881) according to the manufacturer's instructions.

[0070] Library preparation and nanopore sequencing The purified cDNA was prepared for nanopore sequencing. Briefly, the cDNA was subjected to end repair and dA tailing using the NEBNext End Repair / dA Tail Module (NEB, catalog no. E7546S) according to the manufacturer's instructions, and incubated at 20°C for 5 minutes and 65°C for 5 minutes. The reaction mixture was purified using AMPure XP beads (Agencourt). The ONT adapter mix was added according to the Direct cDNA Sequencing Kit protocol (SQK-DCS109, ONT), then loaded onto an R9.4 flow cell and sequenced using a MinION sequencer.

[0071] Data analysis For bioinformatics analysis of direct cDNA sequencing, all alignments to reference sequences were performed using BLAST-n or CLC Genomics Workbench (QIAGEN). To generate consensus sequences, single GFP repeats were aligned using Mesquite software, and the WebLogo online tool was used to generate the final consensus sequence.

[0072] Overview of the circAID-p-seq method Step 1. RPF phosphorylation. Upon selection and purification, RPFs bearing 3'P or 3'cP are subjected to 5' phosphorylation by T4 PNK 3'Minus according to the protocol shown in Table 2.

[0073] JPEG0007733642000002.jpg70153

[0074] The reaction is incubated at 37° C. for 1 hour and then purified through a Zymo column purification kit.

[0075] Step 2. RtcB ligation. The RPF from step 1, which is phosphorylated at both the 5' and 3' ends, is ligated to a 109 nt linker RNA molecule (Linker R) by RtcB ligase. RtcB ligase joins the 5' OH end of Linker R to the 3' P / 3' cP end of the RPF according to the protocol shown in Table 3.

[0076] JPEG0007733642000003.jpg76153

[0077] Incubate at 37 °C for 2 hours.

[0078] The reaction is loaded onto a 15% Tris-borate-EDTA (TBE)-urea acrylamide gel, and the ligation product is size-selected, gel-extracted, precipitated in isopropanol, and finally resuspended in 8 μL of water. The purified product (RPF-Linker R) is approximately 140 nt long and has 5'P and 3'OH termini. Such steps are not essential to the practice of the methods disclosed herein.

[0079] Step 3. Circularized 5'P-RPF-Linker-R-3'OH product The 5'P-RPF-linkerR-3'OH product is subjected to circularization by ligation of the 5'P and 3'OH ends with T4 RNA ligase 1. The reaction conditions are shown in Table 4.

[0080] JPEG0007733642000004.jpg51153

[0081] Incubation: 2 hours at 25°C.

[0082] Step 4. RNase R The reaction conditions are provided in Table 5.

[0083] JPEG0007733642000005.jpg48153

[0084] Incubate at 37 °C for 1 hour.

[0085] The reaction is loaded onto a 15% Tris-borate-EDTA (TBE)-urea acrylamide gel, and the circular RNA molecules are gel extracted (gel extraction is not an essential step for the practice of the methods disclosed herein). After isopropanol precipitation, the circular RNA is resuspended in 8 μL and quantified (QuBit quantification).

[0086] Step 5. Reverse transcription-rolling circle amplification (RT-RCA). For the generation of multimeric single-stranded cDNA, the reagents are mixed in the amounts shown in Table 6.

[0087] JPEG0007733642000006.jpg46153

[0088] Heat the circular RNA primer mix to 65°C for 5 minutes, then incubate on ice for at least 1 minute. Add the reagents to the annealed RNA in the amounts shown in Table 7.

[0089] JPEG0007733642000007.jpg46153

[0090] Incubate at 42°C for 4 hours, then add 0.1 N NaOH and heat the mixture to 70°C for 20 minutes. Finally, precipitate the reaction by adding 156 μL of nuclease-free water, 20 μL of 3M sodium acetate, 300 μL of isopropanol, and 2 μL of Glycoblue. After precipitation, resuspend in 20 μL of nuclease-free water.

[0091] Step 6. Synthesis of the second strand. To produce the second strand from the single stranded cDNA molecules (produced in step 5), one cycle of PCR is carried out under the reaction conditions provided in Tables 8 and 9.

[0092] JPEG0007733642000008.jpg63153

[0093] JPEG0007733642000009.jpg48153

[0094] The reaction is purified by adding 45 μL of AMPure XP beads (agencourt). The final product is eluted in a total volume of 25 μL of nuclease-free water.

[0095] Step 7. ONT library preparation. The purified double-stranded cDNA (see step 6) is used for ONT library preparation according to the protocol of the Direct-cDNA Sequencing Kit (SQK-DCS109), starting from the "End Prep Step".

[0096] JPEG0007733642000010.jpg229153JPEG0007733642000011.jpg224153JPEG0007733642000012.jpg225153

Claims

1. 1. A method for preparing cDNA molecules for sequencing of at least one RNA molecule contained in a biological sample, comprising: Steps below: (i) obtaining a biological sample containing at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end; (ii) phosphorylating the 5′ end of said at least one RNA molecule, thereby introducing a phosphate group at the 5′ end of said at least one RNA molecule to obtain at least one RNA molecule phosphorylated at both ends; (iii) ligating the 3′ end of at least one RNA molecule phosphorylated at both ends to the 5′ end of a random RNA linker having —OH groups at both ends to obtain at least one first ligation product; (iv) self-ligating the at least one first ligation product to form at least one circular RNA molecule, wherein the at least one circular RNA molecule is mixed with linear RNA molecules present in the biological sample; (v) digesting the linear RNA molecule; (vi) subjecting the at least one circular RNA molecule to reverse transcription rolling circular amplification to obtain at least one single-stranded cDNA molecule having a sequence complementary to at least one copy, or 2 to 500 copies, of the at least one RNA molecule; Including, wherein the at least one single-stranded cDNA molecule having a sequence complementary to the sequence of the at least one copy of the at least one RNA molecule is suitable for sequencing; method.

2. 10. The method of claim 1, The method comprises the further step (vii) of generating a complementary cDNA strand of the at least one single-stranded cDNA molecule to obtain at least one double-stranded cDNA molecule. method.

3. 3. The method of claim 1 or 2, The phosphorylation step (ii) is carried out using a phosphorylating enzyme selected from T4 PNK 3'minus, T4 PNK and a recombinant version of T4 PNK; method.

4. The method according to any one of claims 1 to 3, the ligation step (iii) is carried out using a first ligase enzyme selected from RtcB, Archease, Arabidopsis Thaliana tRNA ligase, and a eukaryotic tRNA ligase; method.

5. The method according to any one of claims 1 to 4, the self-ligation step (iv) is carried out using a second ligase enzyme selected from T4 Rnl1, T4 Rnl2, T4 Rnl2tr, T4 Rnl2 K227Q, Mth Rnl and an ATP-independent ligase that catalyzes intramolecular ligation; method.

6. The method according to any one of claims 1 to 5, The digestion step (v) is carried out using a 3'-5' exoribonuclease or a 5'-3' exoribonuclease; method.

7. The method according to any one of claims 1 to 6, The reverse transcription rolling circular amplification step (vi) is carried out using a reverse transcriptase selected from modified MMLV-RT (Moloney Murine Leukemia Virus Reverse Transcriptase) and AMV-RT (Avian Myeoloblastosis Virus Reverse Transcriptase); method.

8. The method according to any one of claims 2 to 6, The step (vi) of generating the complementary cDNA strand is carried out using a DNA polymerase enzyme selected from Taq polymerase having 5'-3' exonuclease activity; method.

9. The method according to any one of claims 1 to 8, The random RNA linker has a length comprised between 50 and 500 nucleotides. method.

10. The method according to any one of claims 1 to 9, the random RNA linker has a minimum free energy comprised between -3 and -150 kcal / mol; method.

11. The method according to any one of claims 1 to 10, the at least one RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end is produced by treating the biological sample with an endoribonuclease, exoribonuclease, ribozyme, or toxin capable of cleaving mRNA, tRNA, snRNA, snoRNA, Y RNA, lncRNA, piRNA, siRNA, viral RNA, or rRNA; method.

12. A kit comprising random RNA linkers, a first ligase enzyme, an exoribonuclease, and optionally a second ligase enzyme, for use in a method according to any one of claims 1 to 11, (i) the random RNA linker has —OH groups at both ends; (ii) the first ligase enzyme is suitable for ligating the 3' end of an RNA molecule having a phosphate group or a 2',3'-cyclic phosphate group at its 3' end and a phosphate group at its 5' end to the 5' end of the random RNA linker; (iii) the exoribonuclease is suitable for enzymatically digesting linear RNA molecules; and (iv) the second ligase enzyme is suitable for circularizing the ligation product obtained by ligating the random RNA linker to the RNA molecule; kit.

13. 13. The kit of claim 12, the first ligase enzyme is selected from RtcB, Archease, Arabidopsis Thaliana tRNA ligase, and eukaryotic tRNA ligase; kit.

14. 14. The kit according to claim 12 or 13, the second ligase enzyme is selected from T4 Rnl1, T4 Rnl2, T4 Rnl2tr, T4 Rnl2 K227Q, Mth Rnl and an ATP-independent ligase that catalyzes intramolecular ligation; kit.

15. The kit according to any one of claims 12 to 14, The random RNA linker has a length comprised between 50 and 500 nucleotides. kit.

16. The kit according to any one of claims 12 to 15, The random RNA linker has a minimum free energy comprised between -3 and -150 kcal / mol. kit.

17. The kit according to any one of claims 12 to 16, further comprising (i) a kinase, and / or (ii) an endoribonuclease, ribozyme, or toxin capable of cleaving mRNA, tRNA, snRNA, snoRNA, Y RNA, lncRNA, piRNA, siRNA, viral RNA, or rRNA; kit.

18. The kit according to any one of claims 12 to 17, The random RNA linker has the nucleotide sequence set forth in SEQ ID NO:

3. kit.

Citation Information

Patent Citations

  • Method of amplifying DNA from RNA in a sample

    US20140057322A1