Method for analyzing tRNA using direct sequencing
Through nanopore-based methods, using splint double-stranded oligonucleotides and reverse transcription technology, the rapid and accurate quantification of tRNA abundance and modification in biological samples was achieved, solving the problems of low efficiency and no multiple analysis in the prior art, and achieving efficient and accurate tRNA detection.
Patent Information
- Application Number
- CN202380068770.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-27
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to detect and quantify the abundance and modification of tRNA in biological samples quickly and accurately, and the previous methods are inefficient, time-consuming, and do not allow multiple samples to be analyzed.
Using a nanopore-based approach, the quantification of tRNA abundance and modification was achieved by contacting RNA samples containing tRNA with pre-annealed splint double-stranded oligonucleotides, reverse transcription and linearization, followed by direct nanopore sequencing.
Fast and accurate measurement of tRNA abundance and modification is achieved, yield increases by at least 10 times, simplifying library preparation steps, allowing multiple samples to be analyzed, and faster than prior art.
Smart Images

Figure BDA0005328295450000121 
Figure BDA0005328295450000311 
Figure BDA0005328295450000321
Abstract
Description
Technical Field
[0001] The present invention provides a direct and rapid method for detecting and quantifying modifications present in transfer RNA (tRNA) in a biological sample. The present invention also provides a kit for performing the method. Background Art
[0002] tRNA is a highly modified small non-coding RNA that plays a key role in protein translation. The modification present in tRNA is the key to its stability, and as the recognition element (identity element) of correct aminoacylation, it ensures that the fidelity of the genetic code is maintained. Although tRNA plays a central role in the protein translation process, tRNA modification is reversible and can be dynamically regulated during environmental exposure, across cell cycle stages, and during tumorigenesis. Similarly, tRNA abundance is also dysregulated during environmental exposure and in several human diseases. Therefore, tRNA abundance and tRNA modification dynamics are quantified, which is crucial for understanding the function and regulation of tRNA, and for understanding how and why tRNA dysregulation is associated with a variety of human diseases.
[0003] Quantification of tRNA modifications is usually achieved by using mass spectrometry-based methods (LC-MS / MS), which are highly sensitive but do not provide positional information. On the other hand, quantification of tRNA abundance is achieved using next-generation sequencing (NGS)-based methods, which require initial conversion from tRNA molecules to cDNA. Therefore, NGS-based methods are blind to most tRNA modifications.
[0004] A promising alternative to characterizing the tRNAome using NGS-based techniques is the direct RNA nanopore sequencing (DRS) platform developed by Oxford Nanopore Technologies (ONT), which can sequence natural tRNA molecules and, in principle, can detect and measure both tRNA modifications and tRNA abundance without the need for reverse transcription or PCR. Previous work has demonstrated the feasibility of sequencing natural tRNAs using nanopores, using either solid-state nanopores or biological (from Oxford Nanopore Technologies) nanopores.
[0005] One limitation of the methods described in these articles is that tRNA reads cannot be base-called and therefore cannot be mapped.
[0006] Scientific article Thomas NK, Poodari VC, Jain M, Olsen HE, Akeson M, Abu-Shumays RL. Direct Nanopore Sequencing of Individual Full Length tRNA Strands. ACS Nano. 2021; 15 (10): 16642-16653. doi: 10.1021 / acsnano. 1c06488 provides a method for directly quantifying and sequencing a single tRNA. However, compared with the method of the present invention, this method is much more time-consuming and much less efficient. This document discloses a gel purification step, which increases the required input amount (gel purification has a low RNA recovery rate), leads to fragmentation of RNA molecules, and increases the complexity and time required for preparing the library. In addition, the oligonucleotide design of the document (5' and 3' oligonucleotides connected to tRNA molecules) is incompatible with the linearization of tRNA molecules, thus resulting in a decrease in sequencing yield, partly due to pore clogging and the need for gel-mediated tRNA selection, which promotes tRNA fragmentation. In addition, the yield of tRNA reads obtained per sequencing run is extremely low (about 50,000 reads, while 1-2M reads are typically obtained from a DRS run). Furthermore, this document does not disclose whether the exact in vivo tRNA abundance and / or tRNA modifications are recapitulated using this method. The present invention demonstrates that this method is not sufficient to correctly recapitulate the abundance of tRNAs present in a given sample. Finally, this previous method does not allow for sample multiplexing, limiting its applicability to samples where the amount of RNA is not limited.
[0007] The present invention relates to a nanopore-based method that allows rapid, accurate and direct measurement of tRNA abundance and modifications, has at least about 10 times more tRNA reads than previous work, reproduces tRNA modifications at the sites where they are modified, simplifies the library preparation steps and therefore the time required for library preparation, and allows multiplexed analysis of samples.
[0008] Overall, the method of the present invention is a sensitive and accurate method for quantification of tRNA abundance and modification profiles with single transcript resolution. The robust and simple library preparation workflow can be completed within one day, while sequencing can be completed within one to two days, which is faster than the methods disclosed in the prior art. Summary of the invention
[0009] In this specification, the terms "complementary" and "complementarity" are interchangeable and refer to the ability of polynucleotides to form base pairs with each other. Base pairs are usually formed by hydrogen bonds between nucleotide units in antiparallel polynucleotide chains or regions. Complementary polynucleotide chains or regions can be base paired in a Watson-Crick manner (e.g., A to T, A to U, C to G). 100% (or complete) complementarity refers to the situation where each nucleotide unit of a polynucleotide chain or region can form hydrogen bonds with each nucleotide unit of a second polynucleotide chain or region. Incomplete (or partial) complementarity refers to the situation where some (but not all) nucleotide units of two chains or two regions can form hydrogen bonds with each other and can be expressed as a percentage.
[0010] In this specification, the term "hybridization" is used to refer to a structure formed by two independent RNA strands, which form a double-stranded structure by base pairing from one strand to the other. These base pairs are considered to be GC, AU and GU. (A-adenine, C-cytosine, G-guanine, U-uracil). As in the case of complementarity, the hybridization can be complete or partial.
[0011] In the present specification, the expression "base recognition error" refers to mismatch, insertion and deletion generated in DRS.
[0012] In the present specification, the expressions "5' RNA adapter" are synonymous with "first splinted oligonucleotide" and "first splinted RNA oligonucleotide" and can be used interchangeably.
[0013] In the present specification, the expressions "3' RNA adaptor" and "second splint oligonucleotide" and "second splint RNA oligonucleotide" are synonymous, and when this oligonucleotide has at least 1 DNA nucleotide from its 3' end or terminus, it is called "second splint RNA:DNA oligonucleotide" and can be used interchangeably.
[0014] In the present specification, the expressions "ONT RTA adaptor A" and "first adaptor DNA nucleotide" are synonyms and can be used interchangeably.
[0015] In the present specification, the expressions "ONT RTA adaptor B" and "second adaptor DNA oligonucleotide" are synonyms and can be used interchangeably.
[0016] In this specification, the term "oligonucleotide" refers to RNA, DNA or RNA:DNA oligonucleotide.
[0017] In the present specification, unless otherwise explained, the term "nucleotide" may refer to both, ribonucleotide or deoxyribonucleotide.
[0018] A first object of the present invention relates to a method for quantifying tRNA abundance and tRNA modification, comprising the following steps:
[0019] a) contacting an RNA sample containing tRNA with the following substances in the presence of a reagent having ligation activity:
[0020] A pre-annealed splint double-stranded oligonucleotide comprising:
[0021] - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3' end,
[0022] - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region,
[0023] At the end of this step, the first splint oligonucleotide is aligned with the 5'
[0024] The ends are adjacent and complementarily anneal to the 3' end of the tRNA through the tRNA hybridization region, and complementarily anneal to the second splint oligonucleotide through the second splint oligonucleotide hybridization region. Figure 2 , Nano-tRNA seq, Box 1
[0025] b) contacting the product of step a) with the following substances in the presence of a reagent having ligation activity
[0026] A pre-annealed adapter DNA double-stranded oligonucleotide comprising:
[0027] - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and
[0028] - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to said second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide.
[0029] Such that at the end of this step, the first adaptor DNA oligonucleotide is adjacent to the 3' end of the terminal region of the second splint RNA oligonucleotide, and the second DNA oligonucleotide is complementarily annealed to both the second splint RNA oligonucleotide and the first adaptor DNA oligonucleotide. Figure 2 , Nano-tRNA seq, Box 2
[0030] c) performing reverse transcription to linearize the product of step b) to obtain a library ( Figure 2 , Nano-tRNA seq, Box 3)
[0031] d) performing direct nanopore sequencing on the library of step c) to obtain the abundance of tRNA and its modification in the sample ( Figure 2 , Nano-tRNA seq, Box 4).
[0032] In a specific embodiment, step d) is divided into the following sub-steps:
[0033] e) contacting the library of step c) with an oligonucleotide adaptor configured for nanopore direct sequencing in the presence of a reagent having ligation activity
[0034] f) loading the product of the previous step into a flow cell comprising a membrane having a nanopore therein, the nanopore providing a passage through the membrane, coupled to an electric current intensity, wherein the product of the previous step passes through the nanopore causing the electric current intensity to be perturbed, and
[0035] g) analyzing the sequence to obtain the abundance of tRNA and its modification in the sample.
[0036] tRNA is generally between 60-120 nucleotides, while messenger RNA (mRNA) is much longer; there are other small RNAs in the cell, but tRNA usually accounts for the largest proportion of small RNA species. Therefore, in another specific embodiment, the RNA sample containing tRNA is such an RNA sample, wherein at least 10%, preferably at least 20%, more preferably at least 30%, even more preferably at least 40%, even more preferably at least 50%, even more preferably at least 60%, even more preferably at least 70%, even more preferably at least 80%, even more preferably at least 85%, even more preferably at least 90%, even more preferably at least 95%, even more preferably 100% of the RNA has a length of less than 300 nucleotides, preferably less than 200 nucleotides, more preferably less than 150 nucleotides, but in any case, equal to or greater than 40 nucleotides. Therefore, the RNA sample to be sequenced contains tRNA species.
[0037] In a specific embodiment, the RNA sample containing tRNA is a tRNA-enriched RNA sample.
[0038] A tRNA-enriched RNA sample is an RNA sample wherein at least 1% of the RNA molecules of the sample are tRNA, preferably at least 2%, more preferably at least 3%, even more preferably at least 5%, even more preferably at least 7%, even more preferably at least 10% of the RNA molecules of the sample are tRNA, even more preferably at least 20% of the RNA molecules of the sample are tRNA, even more preferably at least 30% of the RNA molecules of the sample are tRNA, even more preferably at least 40% of the RNA molecules of the sample are tRNA, even more preferably at least 50% of the RNA molecules of the sample are tRNA, even more preferably at least 60% of the RNA molecules of the sample are tRNA, even more preferably at least 70% of the RNA molecules of the sample are tRNA, even more preferably at least 80% of the RNA molecules of the sample are tRNA, even more preferably at least 90% of the RNA molecules of the sample are tRNA, and even more preferably 100% of the RNA molecules of the sample are tRNA.
[0039] In a specific embodiment, the reagent with ligation activity can be an enzyme or a chemical reagent, preferably an enzyme. In other aspects, the connection is carried out by enzymatic ligation. Suitable reagents (e.g., ligase and corresponding buffer, etc.) and kits for carrying out enzymatic ligation are known and available, for example, the instant sticky-end ligase premix (Instant Sticky-end Ligase Master Mix) available from New England Biolabs (Ipswich, MA). Available ligases include, for example, T4 RNA ligase 2 (NEB#M0239) from NEB. The conditions suitable for carrying out the ligation will vary depending on the type of ligase used. Information about such conditions is convenient and available.
[0040] However, for the purpose of this specific ligation, the concentration of commercially available ligases is generally very low. Therefore, in a preferred embodiment, the ligase is E. coli T4 Rnl2 (gp24.1), whose nucleotide sequence is SEQ ID N°1 and the codon-optimized sequence SEQ ID N°2, and the concentration ranges from 1 to 20 mg / ml, preferably 2 to 11 mg / ml, preferably 2 to 10 mg / ml, and more preferably 4 to 9 mg / ml. The concentration of the ligase is important for reducing the time required for the efficient ligation of the tRNA and the pre-annealed splint double-stranded oligonucleotide while keeping the reaction volume to a minimum.
[0041] In another specific embodiment, step a) is at room temperature, i.e., between 17°C and 29°C, preferably between 18°C and 28°C, more preferably between 19°C and 27°C, even more preferably between 20°C and 26°C, for 10 min to 6 hours, preferably 20 min to 5 hours, more preferably 30 min to 4 hours, even more preferably 45 min to 4 hours, even more preferably 1 hour to 3 hours.
[0042] In a preferred embodiment, step a) is from 1 hour to 3 hours at a temperature between 20°C and 26°C.
[0043] In another specific embodiment, step a) is performed at a temperature between 2°C and 6°C for 8 to 16 hours.
[0044] At a temperature between 1°C and 15°C, preferably between 1.5°C and 12°C, more preferably between 2°C and 9°C, even more preferably between 2°C and 6°C, it is 4 hours to 24 hours, preferably 5 hours to 22 hours, more preferably 6 hours to 20 hours, even more preferably 7 hours to 18 hours, even more preferably 8 hours to 16 hours.
[0045] In another preferred embodiment, step a) is performed at a temperature between 2°C and 6°C for 8 to 16 hours.
[0046] In another specific embodiment, a crowding agent is added during the ligation process to maximize its efficiency, and the crowding agent is selected from the group consisting of poly(ethylene glycol) (PEG) 8000, PEG 6000 or any other commercial crowding agent and combinations thereof. In a specific embodiment, the crowding agent is PEG8000, about 10% w / v, or about 20% w / v, or about 30% w / v, or about 40% w / v, or about 50% w / v, preferably between 15% w / v and 25% w / v.
[0047] In a specific embodiment, prior to step a), the RNA composition comprising tRNA may be contacted with a reagent comprising a deacylation activity. This step has the advantage of increasing the ligation yield of the next step, since it prevents ligation of the second splint oligonucleotide.
[0048] The deacylation step was optimized from the original deacylation protocol published by Dittmar KA, Goodenbour JM, Pan T (2006) Tissue-Specific Differences in Human Transfer RNA Expression. PLoS Genet 2(12):e221. https: / / doi.org / 10.1371 / journal.pgen.0020221 (see Materials and Methods, tRNA Microarray section, second paragraph) by reducing the reaction volume (finally 10uL RNA in dH20 and 95ul Tris-HCl pH 9.0), omitting the neutralization step, and increasing the volume of ethanol added during column clean-up to collect the deacylated RNA composition containing tRNA.
[0049] The RNA sample containing tRNA used in the methods described herein can be from any source, including humans, animals, plants, bacteria, and fungi or yeast. For example, body fluids (blood or plasma), tissue samples, organs, organelles, or single cells obtained using methods known in the art. Commercial kits are available for isolating RNA, including, for example, RNeasy kits (Qiagen). Methods, reagents, and kits for isolating, purifying, and / or concentrating RNA from a source of interest are known in the art and are commercially available. For example, kits for isolating RNA from a source of interest include nucleic acid isolation / purification kits from Qiagen, Inc. (Germantown, Mc); nucleic acid isolation / purification kits from Life Technologies, Inc. (Carlsbad, CA); Nucleic acid isolation / purification kits. In certain aspects, nucleic acids are isolated from fixed biological samples (e.g., formalin-fixed paraffin-embedded (FFPE) tissues). Commercially available kits such as those from Qiagen, Inc. (Germantown, Me.) can be used. DNA / RNA FFPE kit and Life Technologies, Inc. (Carlsbad, CA) Total Nucleic Acid Isolation Kit for FFPE, isolates RNA from FFPE tissues.
[0050] In a specific embodiment, the RNA composition comprising tRNA is treated with a deoxyribonuclease prior to the method of the invention.
[0051] Prior to step a) and step b) of the method of the present invention, both the splint oligonucleotide and the adapter oligonucleotide need to be pre-annealed in order to obtain a double-stranded oligonucleotide.
[0052] For pre-annealing of the splint oligonucleotides, both the first splint oligonucleotide and the second splint oligonucleotide are mixed in a molar ratio of about 3:1, 2.5:1, 2:1, 1.5:1 or about 1:1 molar ratio first splint oligonucleotide: second splint oligonucleotide. In a specific embodiment, the first splint oligonucleotide and the second splint oligonucleotide are mixed in a molar ratio of about 1:1, 1:1.5, 1:2, 1:2.5 or about 1:3 first splint oligonucleotide: second splint oligonucleotide, preferably 1:1 molar ratio.
[0053] For pre-annealing of the adapter oligonucleotides, both the first adapter DNA oligonucleotide and the second adapter DNA oligonucleotide are mixed at a molar ratio of about 3: 1, 2.5: 1, 2: 1, 1.5: 1, 1.2: 1 or about 1: 1 first adapter DNA oligonucleotide: DNA oligonucleotide. In a specific embodiment, the first adapter DNA oligonucleotide and the second adapter DNA oligonucleotide are mixed at a molar ratio of about 1: 1, 1: 1.2, 1: 1.5, 1:2, 1:2.5 or about 1:3 first adapter DNA oligonucleotide: DNA oligonucleotide, preferably a molar ratio of 1: 1.
[0054] The conditions for annealing can be general conditions used in the art according to conventional protocols known to technicians. In a specific embodiment, the above-mentioned splint oligonucleotide or adapter oligonucleotide is annealed in about 8-12mM Tris-HCL (pH 7.5), 45-55mM NaCl and The adapters were pre-annealed in a solution of ribonuclease inhibitor (Promega, N251A) at a final concentration of about 50 ng / μL and heated to 65°C to 85°C for about 15 s, and then cooled to 18°C to 27°C at a rate of about 0.1°C / s to allow the adapters to hybridize.
[0055] In a specific embodiment, the nucleotides of the first splint oligonucleotide are selected from RNA, DNA or a combination thereof, preferably RNA. When the nucleotides are RNA nucleotides, in this specification, the oligonucleotide may be referred to as a first splint RNA oligonucleotide.
[0056] In another specific embodiment, the tRNA hybridizing region of the first splint oligonucleotide has at least 3 nucleotides (preferably RNA nucleotides) that match the CCA overhang of the tRNA, ie, UGG plus a random nucleotide (N) in the 3' end.
[0057] In a specific embodiment, the 3' terminal sequence of the first splint oligonucleotide is 5'-rUrGrG(rN)-3', N is any ribonucleotide selected from rA, rU, rG or rC.
[0058] In another specific embodiment, the nucleotides of the second splint oligonucleotide are selected from RNA, DNA or a combination thereof.
[0059] In a specific embodiment, the nucleotide of the second splint oligonucleotide is RNA. Therefore, in this specification, the oligonucleotide can be referred to as the second splint RNA oligonucleotide.
[0060] In specific embodiments, the second splint oligonucleotide is a RNA:DNA oligonucleotide.
[0061] In a preferred embodiment, the nucleotides of the second splint oligonucleotide are both RNA and DNA; specifically, there is at least 1 terminal DNA base (preferably 3 terminal DNA bases) starting from its 3' end. Therefore, in this specification, this oligonucleotide can be referred to as a second splint RNA:DNA oligonucleotide. The disadvantage of the second splint DNA oligonucleotide is that the alignment step after sequencing will be worse. The disadvantage of the second splint RNA oligonucleotide is that it will self-hybridize in the first connection step, thereby reducing the yield of step a).
[0062] The second splint RNA:DNA oligonucleotide may have at least 1 terminal DNA nucleotide, preferably 2 or 3 terminal DNA nucleotides, preferably 10 or 11, or 12 or 13 or 14 or 15 or even up to 20 DNA nucleotides starting from the 3' end of the oligonucleotide. The technical effect of these DNA nucleotides is to prevent self-ligation, thereby increasing the overall ligation yield.
[0063] The terminal DNA nucleotide is selected from the group consisting of A, T, C, G and combinations thereof.
[0064] In another specific embodiment, the second adaptor DNA oligonucleotide hybridization region of the second splint oligonucleotide may be any sequence of at least 10 nucleotides, preferably a poly A tail of at least 10 A nucleotides. This overhang of at least 10 nucleotides ensures that the library (i.e., the RNA sample containing tRNAs with the first splint oligonucleotide and the second splint oligonucleotide) can be coupled with the default ONT oligonucleotides known in the prior art for direct RNA library preparation (RTA ligation step).
[0065] A size of at least 10 nucleotides, preferably 20 nucleotides in the adapter design improves the mappability of the reads. The size must be as long as possible to improve the alignment of the tRNAs, and it must be as small as possible to remove it by cleanup; any commercial RNA cleanup kit can be used, preferably a kit that retains any size greater than 17 nt, such as the Zymo cleanup kit. For this purpose, bead cleanup can be used as an alternative to these cleanup kits, which gives more flexibility in terms of RNA loss during the cleanup step.
[0066] In a specific embodiment, the terminal DNA nucleotide(s) are part of the second adaptor DNA oligonucleotide hybridization region of the second splint RNA:DNA oligonucleotide. For example, it may be one, two or three "A"s of the poly A tail of 10 A nucleotides.
[0067] The first splint oligonucleotide and the second splint oligonucleotide are designed such that when the first splint oligonucleotide hybridization region and the second splint oligonucleotide hybridization region of the oligonucleotides hybridize to their respective regions, one end of the second splint oligonucleotide is adjacent to the 3' end of the tRNA, while the opposite end of the second splint RNA:DNA oligonucleotide is adjacent to the first adaptor DNA oligonucleotide. These adjacent ends can be covalently linked (e.g., by enzymatic ligation, e.g., by T4 RNA ligase 2), and can then be used in downstream applications of interest (e.g., PCR amplification, nanopore direct sequencing, next generation sequencing, and / or the like) facilitated by one or more sequences in the utilized oligonucleotides.
[0068] When the first splint oligonucleotide and the second splint oligonucleotide are both RNA nucleotides, step a) and step b) can be performed simultaneously, and the reagent having ligation activity can be any commercial T4 RNA ligase 2, or preferably SEQ ID No 1, more preferably SEQ ID No 2, or a variant having at least 80% identity, more preferably 85% identity, even more preferably at least 90% identity and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity relative to the polynucleotide sequence of SEQ ID No 1 or SEQ ID No 2. In any case, the concentration of the ligase must be 2 to 11 mg / ml, preferably 2 to 10 mg / ml, more preferably 4 to 9 mg / ml.
[0069] When the first splint oligonucleotide is RNA and the second splint oligonucleotide is RNA:DNA, step b) needs to be performed after step a), and the reagent with ligation activity used in step a) can be any commercial T4 RNA ligase 2, or preferably SEQ ID No 1, more preferably SEQ ID No 2, or a variant having at least 80% identity, more preferably 85% identity, even more preferably at least 90% identity and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity relative to the SEQ ID No 1 or SEQ ID No 2 polynucleotide sequence. In any case, the concentration of the enzyme must be 2 to 11 mg / ml, preferably 2 to 10 mg / ml, more preferably 4 to 9 mg / ml. Any commercial DNA ligase can be used in step b).
[0070] In another specific embodiment, the second splint oligonucleotide hybridization region of the second adaptor DNA oligonucleotide can be any sequence as long as it is complementary to the second adaptor DNA oligonucleotide hybridization region of the second splint oligonucleotide. In a specific embodiment, it has a poly-T tail of at least 10 nucleotides, preferably 10 T nucleotides.
[0071] Generally, it is important to keep these oligonucleotides as short as possible to reduce costs, but large enough to align correctly in direct RNA sequencing methods such as nanopore direct sequencing. RNA lengths below 150 nucleotides are poorly detected by this technology, while RNA lengths below 60 nucleotides are not even detected.
[0072] In another specific embodiment, the size of the first splint oligonucleotide is between 15 and 40 nucleotides, preferably between 20 and 30 nucleotides.
[0073] In another specific embodiment, the size of the second splint oligonucleotide is between 25 and 50 nucleotides, preferably between 28 and 35 nucleotides. A size of 24 nt or larger is required for efficient base calling and alignment to capture a maximum number of tRNA reads.
[0074] For optimal read alignability, the second splint oligonucleotide should be 25 nt RNA or more, and in a preferred embodiment it is 30 nt RNA to provide a buffer for potential base recognition problems associated with the 10X poly A tail homopolymer at the 3' end of the oligonucleotide.
[0075] In a specific embodiment, the second RNA:DNA splint oligonucleotide has 3 nt DNA bases from the 3' end to reduce self-ligation. Even with only 1 DNA nucleotide self-ligation will still be avoided. The second RNA splint oligonucleotide will not avoid self-ligation, but it will avoid a clean-up step.
[0076] Since the first splint oligonucleotide cannot be base recognized, it is not possible to confirm the minimum length requirement as for the second splint oligonucleotide, but it should be similar to the 3' oligonucleotide, which means that there should be 20-25 nt of RNA buffer at the 5' end. In a preferred embodiment, the 5' first splint oligonucleotide is 24 nt in length.
[0077] In another specific embodiment, the size of the first adaptor DNA oligonucleotide is between 15 and 40 nucleotides, preferably between 20 and 35 nucleotides.
[0078] In another specific embodiment, the size of the second adaptor DNA oligonucleotide is between 25 and 60 nucleotides, preferably between 30 and 50 nucleotides.
[0079] Table 1 shows specific examples of oligonucleotides that can be used in the methods of the present invention.
[0080] Table 1: Splint oligonucleotide sequences and adapter oligonucleotide sequences
[0081]
[0082] The underlined nucleotides are DNA
[0083] N can be any nucleotide selected from A, C, T, G
[0084] In a specific embodiment, the first RNA splint oligonucleotide is SEQ ID No 3 and the second splint oligonucleotide is SEQ ID No 4 or SEQ ID No 5
[0085] Or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity to SEQ ID No 3, SEQ ID No 4, SEQ ID No 5.
[0086] In a specific embodiment, one or more of these oligonucleotides are modified to have a 5' monophosphate group, because the oligonucleotide has a 5' triphosphate group after being chemically synthesized, preferably the second splint oligonucleotide and the first adapter DNA oligonucleotide. The skilled person will know how to perform such modification using common sense.
[0087] In a specific embodiment of the invention, one or more oligonucleotides of the invention have a randomly shuffled string of nucleotides which serves as a unique identifier for the sample to allow multiplex analysis, i.e., analyzing several samples together simultaneously in steps e) and f) of the method of the invention.
[0088] This barcode or unique identifier uniquely identifies the first adapter DNA oligonucleotide, the second adapter DNA oligonucleotide, or both, and / or uniquely identifies the sample source of the RNA containing the tRNA being sequenced, so as to enable sample multiplexing by labeling each molecule from a given sample with: a specific barcode or "tag"; a barcode sequencing primer binding domain (a domain to which a primer for sequencing the barcode binds); a molecular identification domain (e.g., a molecular index tag, such as a random tag of 4, 6, or other number of nucleotides) for uniquely labeling a molecule of interest, for example, to determine expression levels based on the number of instances in which a unique tag is sequenced; the complement of any such domain; or any combination thereof. In certain aspects, the barcode domain (e.g., sample index tag) and the molecular recognition domain (e.g., molecular index tag) can be contained in the same nucleic acid.
[0089] In a specific embodiment, the barcode or unique identifier is in both the first adapter DNA oligonucleotide and the second adapter DNA oligonucleotide, being complementary between the oligonucleotides. Since the nanopore direct sequencing will only sequence the nucleic acid strand bound to the first DNA oligonucleotide (and not its complementary strand), it is the barcode (specific sequence) contained in the oligonucleotides that will later be used to demultiplex the signal in step f) of the method of the invention.
[0090] Although barcodes for similar methods have been shown before (e.g. Smith MA, Ersavas T, Ferguson JM et al. Molecular barcoding of native RNAs using nanopore sequencing and deep learning. Genome Res. 2020; 30(9): 1345-1353. doi: 10.1101 / gr.260836.120), the present invention expands the repertoire of barcodes and also includes the possibility of longer sequences. Importantly, the barcodes of the present invention cannot be subsequently used for data splitting using the DeePlexiCon algorithm used in the publication.
[0091] In this specific embodiment, the oligonucleotide hybridizing regions of both DNA adaptor oligonucleotides can be replaced with the barcode sequence. It is therefore critical that the barcodes are complementary and have the same size.
[0092] An example of a first adapter DNA oligonucleotide for a multiplex analysis run is: N(15-45)TAGTAGGTTC. SEQ ID No 8 comprises the last 10 nucleotides TAGTAGGTTC of the first adapter DNA oligonucleotide for a multiplex run, which are also the last 10 nucleotides of SEQ ID No 6
[0093] Wherein "N" is a nucleotide selected from the group consisting of A, C, T, G, and will contain the barcode.
[0094] An example of a second adapter DNA oligonucleotide for a multiplex analysis run is: GAGGCGAGCGGTCAATTTTN(15-45)TTTTTTTTTT. SEQ ID 9 comprises the first 19 nucleotides GAGGCGAGCGGTCAATTTT of the second adapter DNA oligonucleotide for a multiplex analysis run, which are also the first 19 nucleotides of SEQ ID No 7
[0095] Wherein "N" is a nucleotide selected from the group consisting of A, C, T, G, and will contain the barcode.
[0096] The exemplary string of 10 "T"s mentioned above is the second splint RNA:DNA oligonucleotide hybridizing region, which can be any other sequence as long as it is complementary to the second adaptor DNA oligonucleotide hybridizing region of the second splint RNA:DNA oligonucleotide.
[0097] In a specific embodiment, the barcodes are:
[0098] >Barcode_2SEQID No 10: GTGATTCTCGTCTTTCTGCG
[0099] >Barcode_3SEQID No 11: GTACTTTTCTCTTTGCGCGG
[0100] >Barcode_4SEQID No 12: GGTCTTCGCTCGGTCTTATT
[0101] >Barcode_23SEQ ID No 13: GCAGCTGGATGCGAGAGGGAAGCTGGTTGAAATAATA
[0102] >Barcode_48SEQ ID No 14:GTTCCGCGCACGCTTTTCTCGTTTTGTCTTTTACACT
[0103] >Barcode_82SEQ ID No 15: GGGAAAGAGCGGTACACTCTCCGAGGGAGAGGCCCAC
[0104] The reverse transcription of step c) allows the product to be linearized, and this linearization reduces the pore clogging that occurs when sequencing structured RNA molecules during nanopore direct sequencing, thereby increasing the sequencing yield by 50% compared to other methods without a linearization step but with a gel purification step. In addition, gel purification is labor-intensive, requires a lot of time, and leads to fragmentation of tRNA.
[0105] The conditions for the reverse transcription are known to any person skilled in the art; for example, a temperature between 55°C and 65°C for a duration of 45 to 75 min is used. Any commercial reverse transcriptase can be used. For example, Maxima reverse transcriptase from ThermoFisher, SuperScript TM II reverse transcriptase or SuperScript TM IV reverse transcriptase.
[0106] The helicase protein is located in one strand of the double-stranded sequencing adapter RNA oligonucleotide, allowing this specific RNA strand to be translocated at a constant rate by the nanopore in the next step and thus sequenced.
[0107] In step d), the sequencing of the linearized product of step c) can be performed by any direct sequencing method including nanopores, such as Oxford Nanopore technology. The nanopore direct sequencing and the materials and protocols for performing the sequencing are known in the art. For example, in EP0815438B1.
[0108] In a specific embodiment, the oligonucleotide adapter configured for nanopore direct sequencing is a double-stranded sequencing adapter DNA oligonucleotide, the helicase protein binds to one of the strands, and has a complementary strand, a first DNA adapter oligonucleotide hybridization region (see Figure 2 , Nano-tRNA seq, Box 3).
[0109] In a specific embodiment, the nanopore direct sequencing comprises a membrane, which can be a solid-state membrane or a biological membrane.
[0110] Any known nanopore direct sequencing method or product can be used to perform step f) of the present invention, such as the method or product disclosed in EP0815438B1 or EP1238275B1.
[0111] The analysis or execution algorithm used in step g) can be any commercial execution algorithm known to the skilled person, among the many execution algorithms known in the art to be suitable for nanopore direct RNA sequencing.
[0112] The first step is to extract the reads. This step can be done by commercial software, such as MinKNOW or any software configured to analyze the sequencing results of the nanopore direct sequencing.
[0113] The next step in the analysis is base calling, which can be accomplished by a skilled artisan using several implementation algorithms known in the art, such as Guppy or Bonito.
[0114] The final step of the analysis is alignment, which can be performed by several known implementation algorithms, such as Minimap2 or BWA, which are universal sequence alignment programs that align nucleic acid sequences to large reference databases.
[0115] In a specific embodiment, the implementation algorithm is configured to capture (and sequence) more tRNAs in a quantitative manner compared to the default parameters or adjusted parameters of MinKnow.
[0116] The method of the present invention, combined with known algorithms for analyzing the nanopore sequencing step, provides 10-12 times more reads, compared with other methods in the prior art (such as Thomas NK, Poodari VC, Jain M, Olsen HE, Akeson M, Abu-Shumays RL. Direct Nanopore Sequencing of Individual Full Length tRNA Strands. ACS Nano. 2021; 15(10): 16642-16653. doi: 10.1021 / acsnano.1c06488), and the number of aligned tRNA reads in these reads is about 5 times higher, and the recovery rate of shorter tRNA molecules is improved, otherwise these shorter tRNA molecules will be lost in large quantities, resulting in a biased representation of tRNA abundance.
[0117] In another specific embodiment, the implementation algorithm in MinKNOW is configured to capture (and sequence) more tRNAs in a quantitative manner.
[0118] The current version of MinKNOW (July 21, 2022), as well as previous versions of MinKNOW, are optimized to capture long RNA reads (typically >200 nucleotides). It detects valid reads in two steps. The molecules that pass through the pore are designated as:
[0119] Adaptors, for up to 5 seconds (can be shorter if the signal variability of the molecule is greater than the variability in the adaptor definition), and only thereafter
[0120] If the chain definition is met for at least 2 seconds, it is considered a chain
[0121] Only reads that meet the strand criteria are reported in the Fast5 file. Therefore, a read must stay in the well for 2-7 seconds before it is reported. This corresponds to approximately 150-490 bases.
[0122] In another specific embodiment, a custom sequencing script was prepared in the commercial software MinKNOW version 1.11.5. for performing sequencing, which is the standard software driving the nanopore sequencing device.
[0123] The parameters of the software (preferably MinKNOW) configured to analyze the sequencing results of the nanopore direct sequencing are adjusted to a maximum of 1 second for the adapter and a maximum of 2 seconds for the strand. By reducing both times, the chance of RNA molecules being classified as reads increases and thus reported in the Fast5 file.
[0124] In a specific embodiment, the execution algorithm uses the (Burrow-Wheeler Aligner) BWA configuration, and its parameters are selected from the group consisting of:
[0125] i) bwa mem-W13-k6-xont2d,
[0126] ii) bwa mem-W13-k6-xont2d-T20,
[0127] iii) bwamem-W13-k6-xont2d-T10,
[0128] iv) bwamem-W9-k5-xont2d-T10 and
[0129] v)bwasw-z10-a2-b1-q2-r1.
[0130] Preferred is bwa mem-W13-k6-xont2d-T20.
[0131] In another specific embodiment, the execution algorithm for the multiplex analysis run includes minimap2 with the following adjusted parameters: -xmap-ont-k6-w3-n1-m10-s13-A1-B1-O1-E1. These adjusted configurations allow for higher sensitivity and better alignment.
[0132] In summary, the present invention provides a simple, cost-effective, high-throughput and reproducible method to accurately reproduce tRNA abundance and simultaneously capture tRNA modification changes using native tRNA nanopore sequencing, providing a novel framework for studying the tRNA group at single-molecule resolution while retaining all tRNA modification information.
[0133] Another object of the present invention relates to a kit for quantifying tRNA abundance and tRNA modification in an RNA sample containing tRNA.
[0134] This kit comprises reagents disclosed by the methods described herein for quantifying tRNA abundance and tRNA modifications in a tRNA-containing RNA sample.
[0135] In a specific embodiment, the kit comprises:
[0136] - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3' end,
[0137] - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region,
[0138] - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and
[0139] - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to said second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide.
[0140] In a specific embodiment, the nucleotides of the first splint oligonucleotide are selected from RNA, DNA or a combination thereof, preferably RNA having a 3' terminal sequence of 5'-rUrGrG(rN)-3'.
[0141] In a specific embodiment, the nucleotides of the second splint oligonucleotide are RNA. In another specific embodiment, the nucleotides of the second splint oligonucleotide are both RNA and DNA, specifically, there is at least one terminal DNA base starting from its 3' end.
[0142] The second splint RNA:DNA oligonucleotide may have at least 1 terminal DNA nucleotide, preferably 2 or 3 terminal DNA nucleotides, even up to 10 or 15 DNA nucleotides starting from the 3' end of the oligonucleotide.
[0143] In another specific embodiment, the sequences of these oligonucleotides are sequences described in Table 1, or variants having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity with respect to one or more of the sequences described in Table 1.
[0144] In another specific embodiment, the kit further comprises a double-stranded sequencing adapter DNA oligonucleotide, the helicase protein binds to one of the strands and has a complementary strand, the first DNA adapter oligonucleotide hybridization region.
[0145] The kit further comprises one or more of the following reagents: a reagent having ligation activity, a reagent for performing reverse transcription, a solution for performing deacylation, and a buffer and solution required for reverse transcription, ligation and / or deacylation reactions.
[0146] In a specific embodiment, the agent having ligation activity is an enzyme, preferably a DNA ligase, an RNA ligase or both. In a preferred embodiment, the RNA ligase is T4 RNA ligase 2. In a more preferred embodiment, the nucleotide sequence of the T4 RNA ligase 2 is SEQ ID No 1 or SEQ ID No 2, or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity with respect to SEQ ID No 1 or SEQ ID No 2.
[0147] The kit may further comprise reagents and solutions for obtaining an RNA sample containing tRNA and / or enriching the amount of tRNA in such an RNA sample.
[0148] The kit further comprises instructions for carrying out the method of the invention and the required plastics.
[0149] The components of the present kit may be present in separate containers, or multiple components may be present in a single container. Suitable containers include a single tube (e.g., a vial), one or more wells of a plate (e.g., a 96-well plate, a 384-well plate, etc.), or the like.
[0150] These instructions for carrying out the method of the present invention using any one of the reagents of the above-mentioned test kit can be recorded on a suitable recording medium. For example, these instructions can be printed on a substrate, such as paper or plastics, etc. Therefore, these instructions can be present in these test kits as package inserts, present in the label of the container of the test kit or its component (that is, associated with packaging or sub-packaging, etc.). In other embodiments, these instructions exist as electronic storage data files present in suitable computer-readable storage media (for example, portable flash drive, DVD, CD-ROM, disk, etc.). In some other embodiments, actual instructions are not present in the test kit, but provide a way to obtain these instructions from a remote source (for example, via the Internet). An example of this embodiment is a test kit comprising a website, at which these instructions can be viewed and / or can be downloaded from the website. As these instructions, the way to obtain these instructions is recorded on a suitable substrate.
[0151] Each of the terms "comprising", "consisting essentially of", and "consisting of" may be replaced with one of the other two terms. The terms "a" or "an" may refer to one or more of the elements it modifies (e.g., "an agent" may represent one or more agents), unless the context clearly describes one of these elements or more than one of these elements. The term "about" as used herein refers to a value within 10% of the underlying parameter (i.e., plus or minus 10%; for example, a weight of "about 100 grams" may include weights between 90 grams and 110 grams). Using the term "about" at the beginning of a listing of values modifies each value (e.g., "about 1, 2, and 3" means "about 1, about 2, and about 3"). When describing a listing of values, the listing includes all intermediate values and all fractional values thereof (e.g., a listing of values "80%, 85%, or 90%" includes an intermediate value of 86% and a fractional value of 86.4%). When a list of values is followed by the term "or more", the term "or more" applies to each value listed (e.g., "80%, 90%, 95%, or more" or "80%, 90%, 95% or more" or "80%, 90% or 95% or more" means "80% or more, 90% or more, or 95% or more"). When describing a list of values, the list includes all ranges between any two of the listed values (e.g., a list of "80%, 90% or 95%" includes ranges of "80% to 90%, "80% to 95%, and "90% to 95%"). Certain implementations of this technology are set forth in the Examples below.
[0152] The present invention also discloses the following clauses:
[0153] 1. A method for quantifying tRNA abundance and tRNA modification, comprising the following steps:
[0154] a) contacting an RNA sample containing tRNA with the following substances in the presence of a reagent having ligation activity:
[0155] A pre-annealed splint double-stranded oligonucleotide comprising:
[0156] - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3' end,
[0157] - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region,
[0158] At the end of this step, the first splint oligonucleotide is aligned with the 5'
[0159] The ends of the tRNA are adjacent to each other and are complementary annealed to the 3' end of the tRNA through the tRNA hybridization region, and are complementary annealed to the second splint oligonucleotide through the second splint oligonucleotide hybridization region,
[0160] b) contacting the product of step a) with the following substances in the presence of a reagent having ligation activity
[0161] A pre-annealed adapter DNA double-stranded oligonucleotide comprising:
[0162] - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and
[0163] - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to said second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide,
[0164] such that at the end of this step, the first adaptor DNA oligonucleotide is adjacent to the 3' end of the terminal region of the second splint RNA oligonucleotide and the second DNA oligonucleotide is complementarily annealed to both the second splint RNA oligonucleotide and the first adaptor DNA oligonucleotide,
[0165] c) performing reverse transcription to linearize the product of step b) to obtain a library, and
[0166] d) performing nanopore direct sequencing using the library of step c) to obtain the abundance of tRNA and its modification in the sample.
[0167] 2. The method according to the preceding clause, wherein step d) comprises:
[0168] - contacting the library of step c) with an oligonucleotide adaptor configured for nanopore direct sequencing in the presence of an agent having ligation activity,
[0169] - loading the product of the previous step into a flow cell comprising a membrane in which a nanopore is present, the nanopore providing a passage through the membrane, coupled to an electric current intensity, wherein the product of the previous step passes through the nanopore, causing a perturbation in the electric current intensity, and
[0170] - analyzing said sequences to obtain the abundance of tRNA and its modifications in said sample.
[0171] 3. The method according to any of the preceding clauses, wherein at least 85% of the RNA molecules in the RNA sample containing tRNA have a size between 300 nucleotides and 40 nucleotides.
[0172] 4. The method according to any of the preceding clauses, wherein at least 10% of the RNA molecules in the RNA sample containing tRNA are tRNA.
[0173] 5. A method according to any of the preceding clauses, wherein the ligating agent in step a) is an enzyme or a chemical reagent, preferably an enzyme, more preferably Escherichia coli T4 RNA ligase 2, whose polynucleotide sequence is SEQ ID N°1 or SEQ ID N°2 or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity with SEQ ID No 1 or SEQ ID No 2.
[0174] 6. A method according to any of the preceding clauses, wherein the concentration of E. coli T4 RNA ligase 2 SEQ ID N°1 or SEQ ID N°2 is between 1 and 20 mg / ml.
[0175] 7. The process according to any of the preceding clauses, wherein the duration of step a) is from 10 min to 6 hours at a temperature between 17°C and 29°C.
[0176] 8. The process according to any of the preceding clauses, wherein the duration of step a) is from 4 hours to 24 hours at a temperature between 1°C and 15°C.
[0177] 9. A method according to any of the preceding clauses comprising between 15% w / v and 25% w / v polyethylene glycol.
[0178] 10. The method according to any of the preceding clauses, wherein the RNA sample containing tRNA is subjected to a deacylation reaction prior to step a).
[0179] 11. A method according to any of the preceding clauses, wherein the RNA sample is obtained from any source, including humans, animals, plants, bacteria and fungi or yeast.
[0180] 12. The method according to any one of the preceding clauses, wherein the RNA composition containing tRNA is treated with deoxyribonuclease prior to step a).
[0181] 13. The method according to any of the preceding clauses, wherein the first splint oligonucleotide and the second splint oligonucleotide are annealed prior to step a) at a molar ratio selected from 3:1 to 1:3, preferably 1:1.
[0182] 14. The method according to any of the preceding clauses, wherein the first adaptor DNA oligonucleotide and the second DNA oligonucleotide are annealed before step a) at a molar ratio of 3:1 to 1:3, preferably 1.2:1.
[0183] 15. The method according to any of the preceding clauses, wherein the first splint oligonucleotide is selected from RNA, DNA or a combination thereof, preferably RNA.
[0184] 16. The method according to any of the preceding clauses, wherein the tRNA hybridizing region of the first splint oligonucleotide is 5'-rUrGrG(rN)-3'.
[0185] 17. The method according to any of the preceding clauses, wherein the second splint oligonucleotide is selected from RNA, DNA or a combination thereof, preferably RNA:DNA.
[0186] 18. A method according to any of the preceding clauses, wherein the nucleotides of the second splint oligonucleotide are both RNA and DNA, starting from its 3' end are at least 1 terminal DNA base, preferably 2 or 3 terminal DNA nucleotides, preferably 10 or even up to 15 DNA nucleotides starting from the 3' end of the oligonucleotide.
[0187] 19. The method according to any of the preceding clauses, wherein the second RNA:DNA splint oligonucleotide has 3 nt DNA bases starting from the 3' end.
[0188] 20. A method according to any of the preceding clauses, wherein the terminal DNA nucleotide(s) are selected from the group consisting of A, T, C, G and combinations thereof.
[0189] 21. A method according to any of the preceding clauses, wherein the second adaptor DNA oligonucleotide hybridization region of the second splint oligonucleotide is any sequence of at least 10 nucleotides, preferably 20 nucleotides, as long as it is complementary to the second adaptor DNA oligonucleotide hybridization region of the second splint oligonucleotide, preferably a poly A tail of at least 10 A nucleotides.
[0190] 22. The method according to any of the preceding clauses, wherein the size of the first splint oligonucleotide is between 15 and 40 nucleotides, preferably between 20 and 30 nucleotides.
[0191] 23. A method according to any of the preceding clauses, wherein when the first splint oligonucleotide and the second splint oligonucleotide are RNA nucleotides, step a) and step b) can be performed simultaneously, and the reagent having ligation activity is commercial T4 RNA ligase 2, or SEQ ID No 1, or SEQ ID No 2, or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity relative to SEQ ID No 1 or SEQ ID No 2.
[0192] 24. The method according to any one of the preceding clauses 1 to 22, wherein when the first splint oligonucleotide is RNA and the second splint oligonucleotide is RNA:DNA, step b) needs to be performed after step a), and the reagent with ligation activity is T4 RNA ligase 2 or SEQ ID No 1 or SEQ ID No 2 for step a), and is any commercial DNA ligase for step b).
[0193] 25. The method according to any of the preceding clauses, wherein the second splint oligonucleotide hybridizing region of the second adaptor DNA oligonucleotide and the second adaptor DNA oligonucleotide hybridizing region of the second splint oligonucleotide are complementary sequences.
[0194] 26. The method according to any of the preceding clauses, wherein the first splint oligonucleotide is between 15 and 40 nucleotides, preferably between 20 and 30 nucleotides.
[0195] 27. The method according to any of the preceding clauses, wherein the second splint oligonucleotide is between 25 and 50 nucleotides, preferably between 28 and 35 nucleotides.
[0196] 28. A method according to any of the preceding clauses, wherein the oligonucleotides are selected from one or more of the group consisting of SEQ ID No 3, SEQ ID No 4, SEQ ID No 5, SEQ ID No 6 and SEQ ID No 7 or variants having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity to SEQ ID No 3, SEQ ID No 4, SEQ ID No 5, SEQ ID No 6 and / or SEQ ID No 7.
[0197] 29. A method according to any one of the preceding clauses 2 to 28, wherein the parameters of the software configured to analyze the sequencing results of the nanopore direct sequencing are adjusted to a definition of at most 1 second for adapters and a maximum of 2 seconds for chains.
[0198] 30. The method according to any one of the preceding clauses 2 to 29, wherein the parameters used to analyze the nanopore direct sequencing results are configured using BWA, and the parameters are selected from the group consisting of:
[0199] i) bwa mem-W13-k6-xont2d,
[0200] ii) bwa mem-W13-k6-xont2d-T20,
[0201] iii) bwamem-W13-k6-xont2d-T10,
[0202] iv) bwamem-W9-k5-xont2d-T10 and
[0203] v)bwasw-z10-a2-b1-q2-r1,
[0204] Preferred is bwa mem-W13-k6-xont2d-T20.
[0205] 31. A method according to any one of the preceding clauses 1 to 27, wherein the hybridization region of the first DNA adaptor oligonucleotide and the second DNA adaptor oligonucleotide are both random nucleotide sequences between 15 and 45 nucleotides, which are unique for each pair of the first DNA adaptor oligonucleotide and the second DNA adaptor oligonucleotide.
[0206] 32. A method according to any of the preceding clauses, wherein more than one tRNA-containing RNA sample is processed simultaneously by step d).
[0207] 33. The method according to any of the preceding clauses, wherein the hybridizing region of the first DNA adaptor oligonucleotide is SEQ ID No 8 and the hybridizing region of the second DNA adaptor oligonucleotide is SEQ ID No 9.
[0208] 34. The method according to any of the preceding clauses 31 to 33, wherein the implemented algorithm for analyzing the nanopore direct sequencing results is minimap2 with the following adjusted parameters: -xmap-ont-k6-w3-n1-m10-s13-A1-B1-O1-E1.
[0209] 35. A kit for performing the method according to any one of clauses 1 to 34.
[0210] 36. A kit according to the preceding clause, comprising:
[0211] - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3' end,
[0212] - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region,
[0213] - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and
[0214] - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to said second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide.
[0215] 37. The kit according to the preceding clauses 35 to 36, wherein the first splint oligonucleotide is an RNA oligonucleotide.
[0216] 38. The kit according to any one of the preceding clauses 35 to 37, wherein the first RNA splint oligonucleotide is SEQ ID No 3 and the second splint oligonucleotide is SEQ ID No 4 or SEQ ID No 5
[0217] Or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity to SEQ ID No 3, SEQ ID No 4, SEQ ID No 5.
[0218] 39. The kit according to any of the preceding clauses 35 to 37, wherein the second splint oligonucleotide is a RNA:DNA oligonucleotide having at least 1 terminal DNA base at its 3' end.
[0219] 40. A kit according to any one of the preceding clauses 35 to 39, comprising a T4 RNA ligase 2 having the following oligonucleotide sequence: SEQ ID No 1 or SEQ ID No 2 or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity to SEQ ID No 1 or SEQ ID No 2. BRIEF DESCRIPTION OF THE DRAWINGS
[0220] Figure 1 .Tapestation gel of small RNA extracted from yeast strains.
[0221] Figure 2 Comparison of tested strategies for sequencing tRNA molecules using nanopore DRS. (A) Schematic diagram of three different library preparations tested for sequencing tRNA molecules, namely strategy A, strategy B and the method of the present invention.
[0222] Figure 3 .The method of the present invention (Nano-tRNA seq) can effectively sequence both in vitro transcribed tRNA populations and natural tRNA populations (A) TBE-UREA gel shows the effect of reaction duration and addition of 20% PEG8000 on ligation efficiency (ON = overnight). (B) IGV snapshots of Nano-tRNAseq aligned reads from synthetic in vitro transcribed tRNA (upper column) or biological tRNA (lower column). Positions with mismatch frequencies greater than 0.2 are colored, while those positions showing mismatch frequencies less than 0.2 are shown in gray. (C) Scatter plot of tRNA abundance showing the reproducibility of Nano-tRNAseq when sequencing biological replicates of WT S. cerevisiae tRNA. The strength of the correlation is indicated by the Spearman correlation coefficient (ρ).
[0223] Figure 4 The choice of alignment software and parameters drastically affects the number of aligned tRNA reads. (A, B) Aligned to in vitro transcribed D. melanogaster mitochondrial tRNA Ala(UGC) (A) and Saccharomyces cerevisiae tRNA Phe (B) IGV snapshots of reads aligned using different alignment algorithms (minimap2, bwa-mem, bwa-sw) and parameters. The 5' RNA adapter region and the 3' RNA adapter region attached to the end of the tRNA molecule are included in the reference for alignment and are represented by orange and red bars, respectively. Positions with mismatch frequencies greater than 0.2 are colored, while those positions showing mismatch frequencies below 0.2 are shown in gray. (C) Bar graph depicting the effect of algorithm and parameter selection on the relative proportions of uniquely aligned reads (green) and misaligned reads (purple, reads aligned to the antisense strand are used as proxies to assess misalignments) (Table 5). (D) Proportion of aligned reads, and (E) alignment identity for each individual template from the bar graph in (C), using minimap2 or bwa-mem (W13-k6-T20). We should note that the Saccharomyces cerevisiae tRNA was not calculated. Phe(F) A bar graph showing the effect of trimming the length of the 5' RNA adapter (blue) and 3' RNA adapter (red) on the alignability of tRNA reads (Table 7). The conditions used by Nano-tRNAseq are shown in grey, while the effect of not using RNA adapters is shown in black.
[0224] Figure 5 Adjustment of MinKNOW parameters increases the number of tRNA reads sequenced and aligned. (A) Graph showing the conceptual differences between default and custom MinKNOW read classifications. (B) Histogram of sequencing yields based on base-called and uniquely aligned reads obtained with the default custom configuration (Table 8). (C) Histogram of relative fold changes in uniquely aligned reads relative to tRNA length (Table 9). (D) Histogram of read counts and alignment lengths for in vitro transcribed (IVT) tRNA reads captured with the default and custom configurations. (E) IVT tRNA molecules recovered with default and custom settings for Drosophila melanogaster mitochondrial tRNA Ala(UCG) and S. pneumoniae tRNA Ser(UGA) and native Saccharomyces cerevisiae tRNA Phe (Table 10) Histogram of relative proportions of reads, where the dashed line indicates the expected proportion. (F) Expected log read counts versus observed log read counts for 9 IVTtRNA molecules captured using a custom MinKNOW configuration (Table 11). Spearman correlation (ρ) is shown.
[0225] Figure 6.Comparison of the activity of different reverse transcriptases on tRNA linearization. (A) Strategy used to test the reverse transcription activity of different enzymes. Starting from in vitro transcribed (IVT) tRNA or native tRNA (1), the tRNA is polyadenylated (2) and annealed with an oligodT adaptor (3), which is used to initiate cDNA synthesis using different RT enzymes and conditions. The RNA strand of the linearized product (4) is digested, leaving the cDNA strand (5), which is examined by TapeStation. (B) Depicted are TapeStation spectra of the original poly A (pA) tRNA product (blue) and the cDNA product (orange) produced by reverse transcription of the template using different reverse transcriptases and incubation conditions. Truncated cDNA products are shown in gray triangles. The 25 nt peak present in all samples corresponds to the loading size marker. The upper column is IVT tRNA, while the lower column is commercial Saccharomyces cerevisiae tRNAPhe. (C) Helicase speed over time (events / s roughly corresponds to nt / s sequenced) for wild-type (WT) Saccharomyces cerevisiae total tRNA sequenced with or without reverse transcription (RT) and classified using the default or custom MinKNOW configuration. (D) Histogram showing the fold change in reads that were base called and uniquely mapped in the presence of RT compared to the absence of RT when WT Saccharomyces cerevisiae total tRNA was linearized.
[0226] Figure 7 Nano-tRNAseq enables quantification of tRNA abundance and RNA modifications and captures modification interdependencies. (A) tRNAs from WT and Pus4 KO Saccharomyces cerevisiae strains Ala(AGC) IGV track. Positions with mismatch frequencies greater than 0.2 are colored, while positions with mismatch frequencies less than 0.2 are shown in gray. Below are magnified IGV snapshots of the Ψ55 region of selected tRNAs, where the upper rows correspond to WT biological replicates and the lower rows are Pus4 KO biological replicates. (B) Comparison of mismatch frequencies of known Ψ sites in Saccharomyces cerevisiae WT versus Pus4 KO tRNA molecules. Each data point represents a known tRNA Ψ site, with red indicating the Ψ55 site and the black outline indicating sites with a total base recognition error ≥ 0.25 compared to WT, which serves as an indicator of Ψ modification frequency. (C) Heat map of total base recognition errors in Pus4 KO relative to WT for each nucleotide (x-axis) and for each tRNA isoacceptor (y-axis, arranged in descending order from most abundant to least abundant). Nucleotides with higher total base misidentification frequencies in the KO strain relative to the WT are shown in red tones, while those with lower total base misidentification frequencies in the KO are shown in blue tones, such as for the Pus4 target Ψ55 (green arrow) and in m 5U54 and m 1 A58 (pink arrow) indicates that these RNA modifications are downregulated in the Pus4 KO strain relative to the WT. (D) Schematic representation of the T-loop of the tRNA targeted by the Pus4 enzyme. The nucleotides with dashed outlines represent the Pus4 binding motif (RRUUCNA), and Ψ55 added by Pus4 is highlighted in green. RNA modifications m5U54 and m 1 A58 is highlighted in pink. (E) LC-MS / MS validation of RNA modification levels in tRNA from Saccharomyces cerevisiae cultured under standard conditions (WT), exposed to oxidative stress or heat stress, and Pus4 KO. Columns represent mean ± SEM of n = 3 biological replicates per condition. P values were determined using one-way analysis of variance (ANOVA) with Tukey correction for multiple comparisons and significance compared to WT. *P < 0.05, **P < 0.01, ***P < 0.001.
[0227] Figure 8 .Nano-tRNAseq tRNA abundance and modification quantification is highly reproducible. (A) Scatter plot of tRNA abundance in the Saccharomyces cerevisiae Pus4 knockout (KO) across biological replicates is shown. See also Table 14. Each dot represents a tRNA alloacceptor and is colored based on the alloacceptor type. Volcano plot of differential expression of pus4KO versus WT. Differentially expressed tRNAs were defined as having a corrected -log10 P value < 0.01 and an absolute log2 fold change greater than 0.6. (B) Pus4 KO biological replicate 1 (e.g. Figure 7 Figure 2 shows a heat map of the summed base miscall frequencies of the nucleotides of the WT (as in C) and replicate 2. Nucleotides with higher summed base miscall frequencies relative to WT are represented in red tones, while nucleotides with lower summed base miscall frequencies are represented in blue tones.
[0228] Fig. 9.Characterization of tRNA abundance and modification dynamics upon exposure to stress reveals that the CCA tail is deadenylated during oxidative stress. (A) Scatter plot of tRNA abundance in biological replicates of Saccharomyces cerevisiae subjected to heat stress (45°C for 1 hour) and oxidative stress (2 mM H202 for 1 hour). Each point represents a tRNA alloacceptor and is colored based on the alloacceptor type. The strength of the correlation is indicated by the Spearman correlation coefficient (ρ). See also Table 14. Volcano plot of differential expression of heat stress or oxidative stress against WT. See also Tables 15-16. Differentially expressed tRNAs are defined as those with a corrected log10 P value <0.01 and an absolute log2 fold change greater than 0.6. (B) Heat map of total and base misidentification errors in oxidative stress relative to WT for each nucleotide (x-axis) and for each tRNA (y-axis, arranged in descending order from most abundant to least abundant). Nucleotides with lower total base misidentification frequencies relative to WT are indicated in blue tones, while nucleotides with higher total base misidentification frequencies are indicated in red tones, as seen for the terminal A at position 76 (black arrow). (C) Schematic representation of a common S. cerevisiae cytoplasmic tRNA in its usual secondary structure, with the terminal A nucleotide of the CCA tail highlighted in red. (D) Zoomed-in snapshot of the IGV track with a close-up of the terminal A (black arrow). (E) Bar graph of the frequency of missing terminal A bases for each S. cerevisiae tRNA isoacceptor under oxidative stress (red), Pus4 knockout (KO) (orange) or heat stress (purple), or under WT conditions (blue).
[0229] Fig.10 . Nanopore signal from channel 300 is shown. Reads acquired using MinKNOW with "default" parameters (upper column) and using "custom" parameters optimized to acquire tRNA reads (lower column) are compared. In each column, the current intensity (in pA y-axis) captured at a specific nanopore channel (channel 300) for a given duration (shown in seconds, x-axis) is depicted.
[0230] Fig.11 .Snapshot of MInKNOW illustrating the "reload scripts" button.
[0231] Fig.12 .A snapshot of MInKNOW once the alternative configuration has been enabled (in this case, "FLO-MIN106-SHORT").
[0232] Fig.13 .Describes how to save a snapshot of MinKNOW for batch files.
[0233] Fig.14.Comparison of ligation efficiency of three ligases at different concentrations.
[0234] Fig.15 .Gel images of new oligonucleotides versus old oligonucleotides and ligation time. DETAILED DESCRIPTION
[0235] Example
[0236] Materials and methods
[0237] Preparation of tRNA for in vitro transcription (IVT)
[0238] A total of 9 unmodified in vitro transcribed tRNAs (see Table 2) were prepared as previously described [Saint-Léger A, Bello C, Dans PD, Torres AG, Novoa EM, Camacho N et al. Saturation of recognition elements blocks evolution of new tRNA identities. Sci Adv. 2016; 2: e1501860]. Briefly, each tRNA was assembled using six DNA oligonucleotides that were first annealed and then ligated between the HindIII and BamHI restriction sites of plasmid pUC19. BstNI linearized plasmids were used for in vitro transcription with T7 RNA polymerase according to standard and known protocols [Sampson JR, Uhlenbeck OC. Biochemical and physical characterization of an unmodified yeast phenylalanine transfer RNA transcribed in vitro. Proc Natl Acad Sci US A. 1988; 85: 1033–1037.]. Transcripts were separated by 8M urea / 10% polyacrylamide gel electrophoresis. tRNAs were identified by UV shadowing, electroeluted and ethanol precipitated. The tRNA pellet was resuspended in RNase-free water and the integrity of the IVT tRNA product was checked using 8M urea / 15% polyacrylamide gel and stored until further use.
[0239] Table 2 IVT tRNA and commercial tRNA references. Each IVT tRNA and commercial Saccharomyces cerevisiae tRNA with splint oligonucleotides Phe References
[0240]
[0241] Nucleotides in bold correspond to the first splint RNA oligonucleotide and the second splint RNA:DNA oligonucleotide.
[0242] Removal of the 5' triphosphate group of IVT tRNA
[0243] The 5' triphosphate groups were converted to 5' monophosphate groups by incubating 1 μL RppH enzyme (NEB, M0356S) per 100 ng input IVT tRNA with 1X Thermopol buffer (NEB, B9004S) in a total reaction volume of 30 μL at 37°C for 2 h. The reaction was inactivated by adding 0.6 μL 500 mM EDTA and incubating at 65°C for 5 min, and then cleaned up using the Zymo RNA Clean and Concentrator-5 kit (Zymo, R1016) according to the manufacturer's instructions to retain RNA ≥17 nt. This step is only required for IVT RNAs because they have three phosphate groups at their 5' end. In vivo tRNAs have only one phosphate group, so this step is not required for RNA samples containing tRNA.
[0244] Yeast strains and culture
[0245] Saccharomyces cerevisiae parent strain (BY4741) and Pus4 knockout strain (BY4741 MATa pus4::KAN) were obtained from the Yeast Knockout Collection (Dharmacon) and grown overnight at 30°C under standard conditions in 4 mL yeast extract peptone dextrose (YPD) medium (1% yeast extract, 2% bacterial peptone and 2% dextrose). The next day, the culture was diluted to 0.0001 OD600 in 200 mL YPD and grown overnight at 30°C with shaking (250 rpm). When the culture reached the mid-exponential growth phase (between OD600 0.5), the WT culture was divided into 3×50 mL subcultures, which were then incubated at 30°C (control), 45°C (heat stress) or in 2 mM H2O2 (oxidative stress) for 1 h. The Pus4 culture was divided into 1×50 mL culture and incubated at 30°C. After incubation, the culture was quickly transferred to pre-cooled 50 mL Falcon tubes and centrifuged at 3000 g for 5 min at 4° C., followed by two washes with water and then the pellet was snap frozen at −80° C. Repeats were performed on consecutive days.
[0246] Extraction of RNA from yeast cultures
[0247] The quick-frozen yeast pellet was resuspended in 660 μL of 425-600 μm glass beads (Sigma-Aldrich, G8772) with 340 μL of acid-washed and autoclaved Reagent (Thermo Fisher Scientific, 15596018). Cells were disrupted by vortexing seven times for 15 s at top speed and cooling the samples on ice for 30 s between cycles. The samples were then incubated at room temperature for 5 min and 200 μL of chloroform was added. After briefly vortexing the suspension, the samples were incubated at room temperature for 5 min. The samples were then centrifuged at 14,000 g for 15 min at 4°C and the upper aqueous phase was transferred to a new tube. To precipitate RNA, 1 volume of molecular grade isopropanol and 1 μL of GlycoBlue TM Coprecipitant (Thermo Fisher Scientific, AM9515), mixed by inversion, and incubated at room temperature for 10 min. The sample was centrifuged at 14,000g for 15 min at 4°C, and the precipitate was washed with ice-cold 70% ethanol. After air drying on the bench for 5 min, the precipitate was resuspended in nuclease-free water and the RNA purity was measured using a NanoDrop1000 spectrophotometer. The sample was treated with Turbo deoxyribonuclease (Thermo Fisher Scientific, No. AM2238) and then cleaned up using the Zymo RNA Clean and Concentrator-5 kit (Zymo, R1016) according to the manufacturer's instructions to retain RNA ≤200 nt. Briefly, 1 volume of RNA binding buffer was mixed with 1 volume of 100% ethanol. 2 volumes of this RNA binding buffer and ethanol solution were added to the reaction, transferred to the Zymo-IC column, and spun at ≥12,000g for 1 min at room temperature. Add 1 volume of 100% ethanol to the flow-through containing the 17-200 nt fraction and transfer it to a new Zymo-IC column and spin at ≥12,000g for 1 min at room temperature. Add 400 μL of RNA preparation buffer to the column and spin at ≥12,000g for 1 min at room temperature, then add 800 μL of RNA wash buffer and spin at >12,000g for 2 min at room temperature, transfer to a fresh collection tube, and spin for 1 min. Elute RNA in nuclease-free water. Figure 1 ) RNA concentration was determined using Qubit fluorescence quantification, RNA purity was measured using a NanoDrop 1000 spectrophotometer, and RNA electropherograms were obtained using an Agilent 4200 TapeStation RNA HS Screentape assay.
[0248] RNA extraction and size selection of small RNAs from human cell lines
[0249] Resuspend the snap-frozen pellets from different cell lines in 500 μL Reagent (Thermo Fisher Scientific, 15596018) and incubate at room temperature for 5 min, and add 100 μL chloroform. After briefly vortexing the suspension, the sample was incubated at room temperature for 3 min. The sample was then centrifuged at 16,000 g for 15 min at 4 ° C, and the upper aqueous phase was transferred to a new tube. To precipitate RNA, 0.7 volumes of molecular grade isopropanol and 1 μL GlycoBlue TM Coprecipitant (ThermoFisher Scientific, AM9515), mixed by inversion, and incubated at room temperature for 15 min. The sample was centrifuged at 12,000g for 30 min at 4 ° C, and the precipitate was washed with ice-cold 70% ethanol. After centrifugation at 7,500g for 5 min, air-dried on the bench for 8 min, the precipitate was resuspended in nuclease-free water, and RNA purity was measured using a NanoDrop 1000 spectrophotometer. The sample was treated with Turbo deoxyribonuclease (Thermo Fisher Scientific, No. AM2238) and then cleaned up using the Qiagen RNeasy MinElute cleanup kit (Qiagen, No. 74204) according to the manufacturer's instructions to retain RNA ≤200 nt, but some steps were added to simultaneously retain RNA >200 nt. Briefly, 350 μL of RLT buffer was added to the sample, 250 μL of 100% ethanol, and the sample was then transferred to an RNeasy MinElute spin column and centrifuged at >8,000 g for 40 s at room temperature to retain RNA >200 nt. 450 μL of 100% ethanol was added to the flow-through containing the 17-200 nt fraction, transferred to a new RNeasy MinElute spin column, and centrifuged at >8,000 g for 40 s at room temperature. 500 μL of RPE wash buffer was added to all columns: those retaining RNA >200 nt and those retaining RNA ≤200 nt, and spun at ≥10,000 g for 40 s at room temperature, then 500 μL of 80% ethanol was added and the column was first spun at >10,000 g for 40 s at room temperature and then for 2 min to dry any ethanol that might have remained. RNA was eluted in 14 μL of nuclease-free water. RNA concentration and purity were determined using a NanoDrop 1000 spectrophotometer, and RNA electropherograms were obtained using an Agilent 4200 TapeStation RNA HSS Creentape assay.
[0250] tRNA deacylation
[0251] Purified tRNA in vivo can be aminoacylated, which can inhibit 3' ligation and poly A tailing reactions. Commercial Saccharomyces cerevisiae tRNAPhe (Sigma Aldrich, R4018), commercial Saccharomyces cerevisiae total tRNA (Sigma Aldrich, AM7119), and tRNA purified from Saccharomyces cerevisiae BY4741 WT culture and Pus4 knockout culture were resuspended in 10 μL of nuclease-free water and deacylated in 95 μL of 100 mM Tris-HCl (pH 9.0) for 30 min at 37°C. Deacylated tRNA was recovered using the Zymo RNA Clean and Concentrator-5 kit (Zymo, R1016) according to the manufacturer's instructions to retain RNA ≥ 17 nt, and confirmed using the Agilent 4200 TapeStation RNA HS Screentape assay.
[0252] Nanopore direct tRNA sequencing library preparation (Nano-tRNAseq)
[0253] The tRNA library was prepared using the SQK-RNA002 kit (Oxford Nanopore Technologies) with some protocol modifications as described herein. All oligonucleotides used in this study were obtained from Integrated DNA Technologies (IDT) (see Table 1 for sequences). The 5' RNA splint adapter first splint RNA oligonucleotide SEQ ID NO 3 was designed to be complementary to the 3' NCCA overhang of the mature t RNA, while the 3' splint RNA:DNA adapter: second splint oligonucleotide SEQ ID NO 5 was designed to be complementary to the rest of the 5' RNA splint adapter (first splint RNA oligonucleotide) with a short poly A segment for annealing the ONT RTA adapter (second adapter DNA oligonucleotide) thereto ( Figure 2 ). In 10mM Tris-HCl (pH 7.5), 50mM NaCl and 1μL The 5' RNA splint adapter and 3' RNA splint adapter were prepared in a 1:1 molar ratio in a solution of ribonuclease inhibitor (Promega, N251A) to a final concentration of 50 ng / μL and heated to 75°C for 15 s and cooled to 25°C at a rate of 0.1°C / s to allow these adapters to hybridize. DNA oligonucleotides with the same sequence as the ONT RTA adapter were ordered from IDT and annealed in the same manner as the 5' splint adapter and 3' splint adapter. Deacylated tRNAs were ligated to the pre-annealed 5' splint adapter and 3' splint adapter at a molar ratio of 1.2:1 (assuming an average tRNA length of 90 nt). The ligation was performed at room temperature for 2 h in a total reaction volume of 50 μL using 20% PEG 8000 (NEB, B10048), 1X T4 RNA Ligase 2 Buffer (NEB, B0239S), 4 μL 6 mg / mL recombinant E. coli T4 RNA2 Ligase (made in-house, see below), and 1 μL RNase inhibitor (Promega, N251A). Then add 2 volumes of room temperature equilibrated Ampure RNAClean XP beads (Beckman-Coulter, A63987) to the reaction and pipette gently up and down and incubate for 15 minutes at room temperature on a Hula mixer. Wash the beads with freshly prepared 70% ethanol and leave to air dry. Elute the sample by resuspending the beads in nuclease-free water and incubating for 10 minutes at room temperature on a Hula mixer.
[0254] RNA concentration was determined using RNA HS Qubit fluorescence quantification. 200 ng of 5' and 3' tRNAs were connected to pre-annealed RTA adapters at a molar ratio of 1:2 (approximately 4.3 pmol tRNA to 8.6 pmol RTA adapter). The connection was performed at room temperature for 30 min in a total reaction volume of 15 μL using 1X Quick Ligation Reaction Buffer (NEB, B6058S), 1.5 μL T4 DNA ligase (NEB, M0202M, 2,000,000 units / mL) and 0.5 μL RNase inhibitor (Promega, N251A).
[0255] After ligation, 13 μL of nuclease-free water, 2 μL of 10 mM dNTPs, 8 μL of 5X Maxima H Minus reverse transcriptase buffer, and 2 μL of Maxima H Minus reverse transcriptase (Life Technologies, EP0751) reverse transcription premix were added directly to the reaction, mixed thoroughly by pipetting, and incubated at 60°C for 1 h, at 85°C for 5 min, and then brought to 4°C. As described for the ligation reaction, the linearized tRNA was cleaned up using 2X Ampure RNAClean XP beads. The RNA concentration was not measured during the treatment step because it was below the detectable range of the Qubit RNA HS assay. Finally, the ONT·RMX sequencing adapter was ligated at room temperature for 30 min in a total reaction volume of 40 μL using 1X Rapid Ligation Reaction Buffer (NEB, B6058S), 3 μL T4 DNA Ligase (NEB, M0202M, 2,000,000 units / mL), and 6 μL RMX adapter. Then add 2 volumes of Ampure RNAClean XP beads and mix them into the reaction by gently pipetting up and down, and incubate for 10 min at room temperature on a Hula mixer. The sample was washed twice with 150 μL WSB (wash buffer), where the precipitate was resuspended by flicking the tube. The sample was eluted in 20 μL ELB (elution buffer) and incubated for 10 minutes at room temperature on a Hula mixer. The final library was prepared by adding 17.5 μL nuclease-free water and 37.5 μL vortexed RRB, and kept on ice until loading. According to the standard ONT SQK-RNA002 protocol, the MinION flow cell (FLO-MIN-106) that will perform direct nanopore RNA was quality checked, pretreated (primed) and loaded.
[0256] Alternative (failed) nanopore direct tRNA sequencing library strategy
[0257] The alternative (failed) tRNA DRS pools tested were prepared using the SQK-RNA002 kit (Oxford Nanopore Technologies) with some protocol modifications as described here for the following library preparation strategy ( Figure 2 ).
[0258] Strategy A: Deacylated tRNA was polyadenylated using E. coli poly(A) polymerase (NEB, M0276S) at 37°C for 30 min following the manufacturer's instructions. 5' RNA splint adapters used in Nano-tRNAseq and all library preparation strategies as described were ligated to the poly(A)-tailed tRNA at a 2:1 molar ratio. Reactions were performed overnight at 4°C in a total reaction volume of 50 μL using 20% PEG 8000, 1X T4 RNA ligase 2 buffer, 4 μL 6 mg / mL recombinant E. coli T4 RNA 2 ligase, 1 μL RNaseOUT TM (Invitrogen, 18080051). Then add 1.8 volumes of Ampure RNAClean XP beads and mix into the reaction by gently pipetting up and down, and incubate for 15 minutes at room temperature on a Hula mixer. Wash the beads with freshly prepared 70% ethanol and leave to air dry. For elution, resuspend the beads in nuclease-free water and incubate for 10 minutes at room temperature on a Hula mixer. Use Qubit fluorescence quantification to determine RNA concentration. The connection of RTA and RMX adapters, the final library preparation steps, and the flow cell QC and loading are as described above.
[0259] Strategy B: As described above, the 5' splint RNA adaptor (first splint RNA oligonucleotide SEQ ID No 3) and the ONTRTA adaptor oligo A (first adaptor DNA oligonucleotide) were annealed at a molar ratio of 1: 1. The annealed 5' splint RNA adaptor and 3' splint DNA adaptor were ligated to 5' monophosphate, deacylated tRNA and cleaned up using the same protocol as in strategy A. Ligation of the RMX adaptor, final library preparation steps, and flow cell QC and loading were as described in Nano-tRNAseq.
[0260] Expression of Recombinant Protein of Escherichia coli T4 RNA Ligase 2
[0261] Unlike T4 DNA ligase, commercial T4 RNA ligase 2 is only available in low concentrations. To improve the efficiency of our library preparation ligations, we produced highly concentrated recombinant E. coli T4 RNA ligase 2 in-house. Specifically, the codon-optimized sequence of E. coli T4 RNA ligase 2 SEQ ID No 2 (T4RNL2) ORF DNA was ordered from IDT Synthesis and cloned in frame into the expression plasmid pETM14 with the coding sequence of a hexahistidine tag followed by a 3C precision cleavage recognition sequence. Protein expression and purification were performed at the Protein Technology Department of the Center for Gene Regulation according to previously described procedures [Bullard DR, Bowater RP. Direct comparison of nick-joining activity of the nucleic acid ligases from bacteriophage T4. Biochem J. 2006; 398: 135–144.] (For ligation tests, see Fig.14 and Fig.15 ). For long-term storage at -80°C, add glycerol to a final concentration of 10%.
[0262] Conditions for the production of T4 RNA ligase 2
[0263] The protein expression conditions of SEQ_ID No 1 or SEQ ID No 2 are as follows: for BL21 (DE3) in 1L2TY growth medium; grow in a temperature range of 32°C to 40°C, preferably about 37°C, until OD600 0.4; and induce protein expression with 0.4mM IPTG; at 37°C for about 3h. The maximum protein amount obtained per 1L is about 7mg.
[0264] Suggested protein purification information, although any other known protein purification protocol may be used
[0265] Cell lysate of yeast expressing SEQ ID No 1 or SEQ ID No 2 as mentioned above
[0266] - Add 50 mL of buffer containing about 50 mM TRIS pH 7.4, about 200 mM NaCl, about 50 mM imidazole, and about 10% glycerol, plus protease inhibitors (completely EDTA-free); Triton X100 and PMSF
[0267] - Resuspend the pellet; lyse the cells using a French press and clarify by ultracentrifugation
[0268] purification
[0269] -Affinity column: HiTrap 5mL Ni 2+ column
[0270] -SEC Hiload 16 / 60 200
[0271] The concentration of the ready-to-use ligase T4 RNA ligase SEQ_ID NO 1 or SEQ ID NO is about 2 mg / ml to 11 mg / ml, preferably about 9 mg / mL, in 10 mM Tris-HCl, 50 mM KCl, 35 mM (NH4)2SO4, 0.1 mM DTT, 0.1 mM EDTA, 50% glycerol pH 7.5 at 25° C. However, any other solution well known in the art that does not interfere with the ligase activity may be used.
[0272] Gel purification and LC-MS / MS of tRNA
[0273] Gel-purified tRNA was used only for LC-MS / MS. 5 μg of the 17-200 nt fraction of each sample was prepared in 2X RNA loading dye (NEB, B0363A) and commercial Saccharomyces cerevisiae tRNA as a marker. Phe and total tRNA and heat denatured at 94°C for 5 min. The run samples were loaded onto a 15% 7M TBE-urea gel (Life Technologies, EC6885BOX), leaving an empty lane between each sample to avoid cross contamination, and run at 100 V in 1X TBE until the bromophenol blue marker was three quarters of the way down the gel. The gel was washed with 1X SYBR TM Gold (Invitrogen, S11494) was post-stained for 5 min in the dark. The gel was transferred to a replicator transparent film (Niceday, 607510) and underlit with UV. The gel region corresponding to tRNA (approximately 70 to 110 nt) was cut out with a sterile scalpel and transferred to a Zymo-Spin from a ZR small RNA PAGE recovery kit (Zymo, R1070). TM IV column. tRNA was extracted from the gel according to the manufacturer's instructions and the extracted tRNA profile was confirmed using an Agilent 4200 TapeStation RNA HS Screentape assay. 500 ng of gel-purified tRNA was digested with a nucleoside digestion mix (NEB, M0649) at 37°C for 1 h according to the manufacturer's instructions. HyperSep TMSpinTip columns (ThermoFisher, 60109-404) desalt the nucleoside digestion solution; first, wash the column with 40 μL 60% acetonitrile by centrifugation at 100g for 10 min, then wash the column with 40 μL 0.1% formic acid by centrifugation at 100g for 5 min. The digested sample is mixed with 30 μL formic acid, added to the column, and collected in a fresh collection tube by centrifugation at 100g for 10 min. The effluent is reapplied to the column and centrifuged at 100g for 10 min. The sample now bound to the column is washed with 40 μL 0.1% formic acid by centrifugation at 100g for 5 min. Replace the collection tube to prepare for elution, and add 40 μL 60% acetonitrile to the column, and elute the sample by centrifugation at 100g for 5 min. LC-MS / MS of Saccharomyces cerevisiae tRNA modifications was performed by the CRG / UPF Proteomics Facility; Briefly, 125 ng of each digested and desalted sample was analyzed by LC-MS.MS using a 40 min gradient on an Orbitrap XL. As a quality control, ribonucleoside standards were run between samples to prevent carryover and assess instrument performance. Heat stress replicate 2 had a profile with an altered chromatogram with significantly less Ψ than all other samples and was therefore discarded from analysis.
[0274] tRNA reverse transcription optimization
[0275] As described in Strategy A, IVT tRNA and commercial Saccharomyces cerevisiae tRNA were Phe Poly A tailing was performed. Importantly, only synthetic RNA had a 5' monophosphate group. For the SuperScript II reverse transcription test, 100ng of poly A tailed RNA, 1μL 100μm 3'RT test adapter (SEQ ID No 26 / 5Phos / ACT TGC CTG TCG CTT CTT TTT TTT TTT TTT TTT TTTTVN, "N" can be any nucleotide selected from A, C, T, G, and "V" can be any nucleotide selected from A, C and G), 1μL 10mM dNTP (Promega, M750B) were mixed, the total reaction volume was 12μL, incubated at 65°C for 5min, and then cooled on ice. Then, 4μL 5X first strand buffer (ThermoFisher, 18064014) or 5X first strand (FS) buffer supplemented with 65mM MnCl2, 1μL 0.1M DTT, 1μL RNaseOUT TM 、1μL SuperScript TMII reverse transcriptase (ThermoFisher, 18064014), and the reaction was incubated at 42°C for 1 hour, inactivated by heating at 70°C for 15 minutes, and then digested with ribonuclease. For the SuperScript IV reverse transcription test, 100 ng of poly A-tailed RNA, 1 μL of 100 μm 3'RT test adapter, and 1 μL of 10 mM dNTP were mixed in a total reaction volume of 12 μL, incubated at 65°C for 5 minutes, and then cooled on ice. Then, 4 μL of 5X SuperScript TM IV RT buffer (ThermoFisher, 18090010), 1 μL 0.1M DTT, 1 μL RNaseOUT TM 、1μL SuperScript TM IV reverse transcriptase (ThermoFisher, 18090010), and the reaction was incubated at 55°C or 60°C for 1 hour, inactivated by heating at 85°C for 5 minutes, and then digested with ribonuclease. For the TGIRT reverse transcription test, 100 ng of poly A-tailed RNA, 1 μL of 100 μm 3'RT test adapter, 4 μL of 5X TGIRT RT buffer, 1 μL of 0.1 M DTT, 1 μL of TGIRT TM -III(InGex,TGIRT50), 1μL RNaseOUT TM Mix in a total reaction volume of 19 μL and incubate at room temperature for 30 min. Then, add 1 μL of 10 mM dNTPs and incubate the reaction at 60°C for 1 hour, inactivate by heating at 75°C for 15 minutes, and then perform RNase digestion. For the Maxima reverse transcription test, mix 100 ng of poly A-tailed RNA, 1 μL of 100 μm 3'RT test adapter, and 1 μL of 10 mM dNTPs in a total reaction volume of 12 μL, incubate at 65°C for 5 min, and then cool on ice. Then, add 4 μL of 5X Maxima RT buffer, 1 μL of RNaseOUT TM , 1 μL Maxima H Minus reverse transcriptase (ThermoFisher, EP0751), and the reaction was incubated at 55°C or 60°C for 1 hour, inactivated by heating at 85°C for 5 minutes, and then subjected to ribonuclease digestion. The reasoning behind the higher incubation temperature is that it better unwinds the tRNA structure and allows access to the reverse transcriptase. After reverse transcription, 1.5 μL RNase Cocktail TMEnzyme mix (ThermoFisher, AM2286) was added to the reaction and incubated at 37°C for 10 min to digest the RNA. The reaction was cleaned up using 1.5X Ampure XP beads as described, and tRNA cDNA and input poly A tRNA were run on the TapeStation using the RNA HS assay. Although we cannot be sure that cDNA runs at the same speed as RNA on the TapeStation, we infer that we should still be able to observe the truncated cDNA template as a peak in the electropherogram.
[0276] Saccharomyces cerevisiae tRNA reference set
[0277] Reference sequences of mature Saccharomyces cerevisiae tRNAs were retrieved from GtRNAdb2 [Chan PP, Lowe TM. GtRNAdb 2.0: an expanded database of transfer RNA genes identified in complete and draft genomes. Nucleic Acids Res. 2016; 44: D184–9.]. GtRNAdb2 reports 275 tRNA sequences annotated in the Saccharomyces cerevisiae genome, but only 43 unique isoacceptor-anticodon pairs (including Und-NNN). Most tRNA isoacceptors (i.e., with the same anticodon) have multiple copies, such as Asp-GTC and Gly-GCC, which each have 16 copies. Most of these copies are identical, with only 55 unique mature tRNA sequences. The remaining 12 sequences are highly similar copies with 95-99% identity. For example, Asp-GTC-1 and Asp-GTC-2 have 96.9% identity. To facilitate reliable alignment and accurate tRNA quantification, we decided to keep only a single representative sequence for each isoacceptor-anticodon pair. We selected the first annotated sequence (i.e., Gly-GCC-1-1) for each isoacceptor-anticodon pair in GtRNAdb2.
[0278] Base calling and alignment of tRNA reads
[0279] Reads were base called using Guppy basecaller v3.6.1 in high accuracy mode. All Us were converted to Ts before alignment. Base called reads were aligned using minimap2 v2.17-r941 with recommended parameters (-xmap-ont) or sensitive parameters (-ax map-ont-k5), or BWA v0.7.17-r1188. For BWA, two modes (MEM and SW) were tested, and array parameters were called as follows (ordered from the most stringent to the least stringent settings): i) bwamem-W13-k6-xont2d, ii) bwa mem-W13-k6-xont2d-T20, iii) bwamem-W13-k6-xont2d-T10, iv) bwamem-W9-k5-xont2d-T10 and v) bwasw-z10-a2-b1-q2-r1. Reads aligned to the reverse strand (antisense strand) were designated as "wrongly aligned". By comparing the number of uniquely aligned reads and the number of wrongly aligned reads (Table 5), we selected the best performing algorithm and parameters (bwa mem-W13-k6-xont2d-T20). We should note that when aligning tRNA reads, the sequences of the 5' RNA adapter and 3' RNA adapter were included in the respective references. The effect of the presence and length of the 5' RNA adapter and 3' RNA adapter on the read alignability was investigated by shortening the respective adapter sequences from the alignment reference in steps of 5 nucleotides (Table 7).
[0280] Analysis of tRNA abundance
[0281] Only unique (alignment quality above 0) primary alignments were considered. Differentially expressed tRNAs were inferred using DESeq2. Volcano plots were generated using the EnhancedVolcano package [Blighe, Rana, Lewis. EnhancedVolcano: Publication-ready volcano plots with enhanced colouring and labeling. R package version.]. Differentially expressed tRNAs were defined as those with a corrected P value < 0.01 and an absolute log2 expression fold change greater than 0.6.
[0282] Analysis of differential tRNA modifications
[0283] Differential tRNA modifications were measured using differential base call errors (mismatches, insertions, deletions) for each tRNA nucleotide. The sum of base call errors was calculated by subtracting the frequency of the reference base from 1. The frequency of the reference base was equal to the number of reads with base calls equal to the reference base divided by the coverage depth at this position. Only uniquely aligned reads (primary alignments with alignment quality above 0) were considered for downstream analysis.
[0284] Adjusting MinKNOW parameters to capture small RNA molecules
[0285] The MinKNOW configuration can be adjusted as follows to enhance the capture of reads of small RNA molecules such as tRNAs from direct RNA nanopore sequencing runs (if these parameters are not adjusted, MinKNOW will incorrectly discard small RNA reads (or tRNA reads) as "adapter-only" reads, and these reads will never be stored in the FAST5 file or FASTQ file). Thus, this step enhances the recovery of small RNA or tRNA populations. We should note that the "default" MinKNOW parameters will still capture small RNAs or tRNAs, but this will reduce the number of reads by about 10 times (see Fig.10 ), and biased representation of tRNAs (longer small RNAs or tRNAs will be over-represented compared to short tRNAs, see Figure 5 E).
[0286] For all 512 channels, sequencing runs were performed with the option "bulk dump" raw files turned on. Sequencing simulations were performed with default and custom MinKNOW configurations. By default, MinKNOW defines the duration of adapters as at most 5 seconds, while chains (actual reads) are at least 2 seconds. Therefore, an RNA molecule must stay in the well for up to 7 seconds to be classified and reported as an actual read. The average speed of the motor protein (RNA helicase) used in the DRS experiment is 70nt per second, so 7 seconds corresponds to approximately 490nt. This definition makes sense for sequencing long molecules because it filters out adapter-only reads. However, for sequencing short RNAs, it would be reasonable to shorten the definitions of both adapters and chains. We evaluated several configurations that shortened the duration of adapters to 1 second and shortened chains to: 1, 2, 3, or 4 seconds. Subsequently, the number of reported reads, base call reads, aligned reads, and unique aligned reads generated by the default and custom MinKNOW configurations were compared. We observed that using 1s definition for adapters and 2s for strands resulted in the highest number of aligned reads and uniquely aligned reads (Table 8). Therefore, these settings will be used throughout the manuscript unless otherwise stated.
[0287] Step 1: Set up an alternative MinKNOW configuration file
[0288] In order to set up an alternative MinKNOW configuration, you must use the terminal to copy and enable the configuration file via rsync. This can be done with the following command (requires root privileges):
[0289] $rsync-a conf / package / flow_cells.toml / opt / ont / minknow / conf / package
[0290] $rsync-a conf / package / sequencing / *.toml / opt / ont / minknow / conf / package / sequencing
[0291] Step 2: Reload the script in MinKNOW
[0292] Each time the script is reloaded, MinKNOW reads the configuration from:
[0293] -flowcells.toml (usually found in opt / ont / minknow / conf / packages / flowcells.toml)
[0294] -sequencing*.toml (usually found in / opt / minknow / conf / packages / sequencing / sequencing_*.toml)
[0295] In order to reload the script, you just need to click on the "Reload Script" button like Fig.11 As shown:
[0296] You should now see the alternative configurations in the Flow Cell Type field (see Fig.12 ). MinKNOW with alternative configurations (such as "FLO-MIN106_short") can now be run, which will allow the capture of short tRNAs during sequencing runs without the need to save bulk FAST5 files.
[0297] Step 3. Start sequencing using a new MinKNOW configuration (or once you have saved a "batch" from a previous run). "fast5" can simulate sequencing run):
[0298] You should run MinKNOW as you would a normal sequencing run, but this time select the alternative configuration that has been set up (FLO-MIN106-Short). We should note that, so far, this can only be used with "bulk" dumps (raw current intensity files) that have been previously saved from your sequencing run. To do this, you must ensure that the batch file is saved during your sequencing run (this is optional in MinKNOW, you must click the option). The batch file should be specified in a configuration file, under the "Custom Settings" section of the ".tomml" file. An example of this file is embedded below:
[0299] ```bash
[0300] ###############################
[0301] #Sequencing Feature Settings#
[0302] ###############################
[0303] #basic_settings#
[0304] [custom_settings]
[0305] enable_elative_unblock_voltage=true
[0306] unblock_vo1tage_gap=480
[0307] run_time=7200#172800#(seconds)1hr=3600
[0308] start_bias_voltage=-180
[0309] #UI parameters
[0310] translocation_speed_min=50
[0311] translocation_speed_max=75
[0312] q_score_min=7
[0313] simulation=" / path / to / bulk_file.fast5"
[0314] ```
[0315] Additional information: How to save batch files
[0316] This is an option given by the MinKNOW software when you set up the parameters in a run. You must check this option and specify how many wells / channels you want to save the current intensity information from (see Fig.13 ). The 72h batch file for all holes can reach 250TM300Gb.
[0317] To enable batch file saving, the user should check the option "Output batch file" as shown above, which will then save the continuous raw current intensities from the sequencing run. This batch file can then be reprocessed using the alternative MinKNOW configuration using the steps above.
[0318] Comparison with published datasets
[0319] The estimated values of Saccharomyces cerevisiae tRNA expression obtained by the methods of the present invention were compared with those obtained by the Illumina-based orthogonal tRNA sequencing methods ARM-seq [Cozen AE, Quartley E, HolmesAD, Hrabeta-Robinson E, PhizickyEM, Lowe TM. ARM-seq: AlkB-facilitated RNA methylation sequencing reveals a complex landscape of modified tRNA fragments. Nat Methods. 2015; 12: 879–884.], Hydro-tRNAseq [Gogakos T, Brown M, Garzia A, Meyer C, Hafner M, TuschlT. Characterizing Expression and Processing of Precursor and Mature Human tRNAs by Hydro-tRNAseq and PAR-CLIP. Cell Rep. 2017; 20: 1463–1475.] and mim-tRNAseq [Behrens A, Rodschinka G, Nedialkova The published estimates are reported per tRNA isoacceptor-anticodon pair and include the same references as those used in this work, with the exception of Hydro-tRNAseq, which omits two references (Leu-GAG and iMet-CAT) and reports five additional references (Leu-AAG, Leu-CAG, Ala-CGC, Pro-CGG, and Arg-TCG). Pairwise comparisons with Hydro-tRNAseq exclude these references.
[0320] result
[0321] Standard nanopore DRS of tRNA molecules results in low sequencing yield and poor 5' end coverage
[0322] Nanopore direct sequencing or nanopore direct RNA sequencing (DRS) is a well-established long-read sequencing technology for studying RNA molecules, typically polyadenylated mRNAs. DRS is inefficient in capturing RNA molecules shorter than 200nt and is generally considered incapable of capturing sequences shorter than about 100nt, limiting its applicability in studying short RNA populations such as tRNAs. In addition, the first approximately 15 nucleotides at the 5' end of an RNA molecule are typically lost in a DRS run [Workman RE, Tang AD, Tang PS, Jain M, Tyson JR. Nanopore native RNA sequencing of a human poly(A) transcriptome. Nature. 2019. Available from: https: / / www.nature.com / articles / s41592-019-0617-2] because this portion cannot be adequately base-recognized when the 5' end of the molecule leaves the helicase due to the increased RNA translocation rate. To overcome these limitations, we reasoned that extension of both the 5′ and 3′ ends of tRNAs would lead to improved sequencing of tRNA molecules, since these molecules would now exceed this ~100 nt threshold in addition to capturing sequence and modification information at the 5′ tRNA end.
[0323] We first attempted a modified tRNA DRS approach in which a 5′ RNA adaptor (first splint RNA oligonucleotide) complementary to the 3′ CCA overhang present in the mature tRNA molecule was ligated to the 5′ end of a tRNA that had been previously polyadenylated in vitro (Strategy A, see Figure 2). A panel of nine synthetic in vitro transcribed (IVT) tRNAs (Table 2) were sequenced using this strategy. However, this approach produced a much lower sequencing yield (56,002 reads) than is typically expected from a DRS run (approximately 1-2 million reads). Furthermore, using minimap2 with the recommended DRS alignment parameters, only 7.5% of the reads were uniquely aligned to the tRNA reference set (Table 4). Relaxation of alignment parameters (-ax map-ont-k5) has previously been shown to improve the comparability of highly modified RNAs [Begik O, Lucas MC, Pryszcz LP, Ramirez JM, Medina R, Milenkovic I et al. Quantitative profiling of pseudouridylation dynamics innative RNAs with nanopore sequencing. Nat Biotechnol. 2021. doi: 10.1038 / s41587-021-00915-6], but it did not produce a significant increase in the number of aligned tRNA reads (Table 4). In addition, we observed poor coverage of the 5' end of the tRNA molecule, indicating that the efficiency of 5' RNA adapter ligation is low, which may be due to the 3' poly A tail annealing or ligase accessibility of the 5' RNA adapter. Space interference.
[0324] Based on these results, we modified the library preparation protocol to replace the poly(A) tail with a 3' DNA adapter complementary to the 5' RNA adapter (Strategy B), allowing the two oligonucleotides to be preannealed and ligated to the tRNA without hindrance from the poly(A) tail ( Figure 2 ). However, this strategy also produced a very low number of sequenced reads (63,502 reads) and a low percentage of uniquely aligned reads (6.5%) (Table 4). Despite the low number of aligned reads in strategy B, there was a slight improvement in coverage of the 5' ends of tRNA molecules.
[0325] Ligase comparison
[0326] Analysis of commercial ligases used in the first step of Nano-tRNAseq library preparation
[0327] The correct ligase for the first adapter ligation step of Nano-tRNAseq is T4 RNA Ligase 2. However, commercial T4 RNA Ligase 2 (NEB or Enzymatics) is not concentrated (10,000 U / mL and 30,000 U / mL, respectively). In contrast, T4 DNA Ligase (NEB) is sold in a concentrated form (2,000,000 U / mL). For this reason, although T4 DNA Ligase is less efficient at ligating this reaction, its increased concentration results in more ligations being achieved in an equivalent volume. Fig.14 We can see that at equal units T4 RNA Ligase 2 is more efficient for this reaction but will require excessive volumes for library preparation.
[0328] Comparison of T4 RNA Ligase 2 (in-house) and Concentrated T4 DNA Ligase (NEB)
[0329] Here, for the first Nano-tRNAseq library preparation reaction, we compared the performance of in-house T4 RNA ligase 2 and concentrated commercial T4 DNA ligase (NEB). Fig.15 ).
[0330] Oligonucleotides
[0331] SEQ ID No 27EN-003FAM:56-FAM-rArArArArArArArArArArA
[0332] SEQ ID No 6EN-010FAM_OligoA_ONT_shuffle1: / 5Phos / GGCTTCTTCTTGCTCTTAGGTAGTAGGTTC / 36-FAM /
[0333] SEQ ID No 7EN-014_oligoB_ONT_polydT:5'-GAGGCGAGCGGTCAATTTTCCTAAGAGCAAGAAGAAGCCTT TTTTTTTT-3'
[0334] Solution: Formamide stop solution: 95% formamide, 20mM EDTA, pH 8.0, 0.03% BPB and 0.03% xylene cyanol FF; Rapid DNA ligase (1X): 50mM Tris-HCl 10mM MgCl2 5mM DTT 1mM ATP at 25°C pH 7.6
[0335] Enzymes: Commercial T4 DNA ligase (NEB), Glycerol-free T4 RNA ligase 2, SEQ ID No 2
[0336] reaction
[0337] Pre-annealed oligonucleotides EN010:EN014: 72.7 pmol, 4.378 × 10 13 Copy number
[0338] EN-003FAM: 37.6pmol, 2.264×10 13 Copy number
[0339] Ligation was performed for 10 min at RT. Run on a 15% TBE-urea gel.
[0340] Different incubation times (ligation gel and / or flongle)
[0341] Oligonucleotides
[0342] 5' Oligonucleotide
[0343] ML-047_ONT_tRNA_oligoB_RNA–24nt
[0344] SEQ ID NO 3rCrCrU rArArG rArGrC rArArG rArArG rArArGrCrCrU rGrGrN
[0345] Old (OLD) oligonucleotide = 3' RNA oligonucleotide
[0346] ML-065_ONT_tRNA_3'Adapter_Long_RNA(Old)–30nt
[0347] SEQ ID NO 4
[0348] / 5Phos / rGrGrCrUrUrCrUrUrCrUrUrGrCrUrCrUrUrArGrGrArArArArArArArArA
[0349] NEW oligonucleotide = 3' RNA-DNA oligonucleotide -> oligonucleotide used in the final Nano-tRNAseq library
[0350] ML-068_ONT_tRNA_3'Adapter_Long_RNA+DNA(NEW)–33nt
[0351] SEQ ID NO 5
[0352] / 5Phos / rGrGrCrUrUrCrUrUrCrUrUrGrCrUrCrUrUrArGrGrArArArArArArArArAAAA
[0353] Flongle Run
[0354] WT yeast
[0355] 1. Overnight Ligation of Old RNA Oligonucleotides
[0356] Run ID: RNA261266
[0357] Liquidity pool ID: AEY401
[0358] 2. Ligation of old oligonucleotides 2h
[0359] Run ID: RNA037486
[0360] Flow cell ID: AJG115
[0361] 3. 2h Ligation of New RNA-DNA Oligonucleotides
[0362] Run ID: RNA816787
[0363] Flow cell ID: AJI950
[0364] Commercial yeast
[0365] 1. Old oligonucleotide ON ligated RNA345135
[0366] 2. New oligonucleotide 2h ligation RNA980403
[0367] Table 3 Results using different oligonucleotides and ligation times
[0368]
[0369]
[0370] Double ligation of RNA adaptors to the ends of tRNA molecules improves tRNA base recognition
[0371] Based on the results of strategy A and strategy B, we reasonably believed that padding the 5' tRNA end and the 3' tRNA end with RNA adapters that can be accurately base-recognized and aligned would enable us to capture the entire tRNA sequence. This method is also called Nano-tRNAseq ( Figure 2 ), and was most successful in using Nanopore DRS to sequence, base-identify, and align in vitro and native tRNA molecules (Table 3).
[0372] The 5' RNA adapter complementary to the CCA overhang of the mature tRNA was pre-annealed to the 3' RNA adapter containing three DNA bases at the 3' end ( Figure 2 ). Surprisingly, we observed that 3' adapters containing only RNA resulted in increased self-ligation (Table 4), a problem we subsequently mitigated by adding DNA bases to the end of the adapter. Next, the pre-annealed 5' RNA splint adapter and 3' RNA:DNA splint adapter were ligated to the deacylated tRNA by T4 RNA ligase 2. Different ligation times and the addition of molecular crowding agents were tested to ensure that conditions were selected that maximized ligation efficiency. Addition of 20% PEG8000 improved ligation efficiency, while running the reaction at 4°C overnight or at room temperature for 2 hours also improved ligation efficiency compared to 20 minutes at room temperature ( Figure 3 A). Since there was little difference between overnight ligation and 2-hour ligation, the latter was chosen as it allowed us to complete the library preparation protocol in one day. Subsequently, ONT RTA Adaptor A and ONT RTA Adaptor B were pre-annealed and ligated using T4 DNA ligase. This approach generated 208,000 base-call reads (Table 4), significantly increasing the sequence output by 3-4 times relative to previous strategies, and also improving the 5' and 3' coverage of both synthetic and biological tRNAs ( Figure 3 In addition, we found that the abundance recapitulated using Nano-tRNAseq was highly reproducible (ρ = 0.984) when sequencing both in vitro tRNA datasets and natural tRNA datasets ( Figure 3 C).
[0373] Alignment parameters and adapter length significantly affect tRNA read alignability
[0374] Due to the short and highly modified nature of natural tRNA reads, their alignment is challenging. In fact, natural tRNAs contain a large proportion of mismatched bases, which often originate from inaccurate base calls of modified bases in the DRS dataset [90,94,97]. As a result of these misidentified bases, the commonly used long read mapper minimap2 with recommended settings (-x map-ont) only aligns a portion of the reads (2.56%). Figure 2 A- Figure 2 C, see also Table 4).
[0375] To improve the alignability of Nano-tRNAseq reads, we first examined whether short read alignment algorithms, such as bwa [Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009; 25: 1754–1760.], could lead to an increase in the proportion of aligned reads. We used DNA from a Drosophila melanogaster mit tRNA construct containing three different tRNA constructs (IVT, Drosophila melanogaster mit tRNA, and Drosophila mit tRNA). Ala(UGC) 、IVT Streptococcus pneumoniae tRNA Ser (UGA) and native Saccharomyces cerevisiae tRNA Phe ) of a Nano-tRNAseq run, we found that the bwa-mem aligner outperformed minimap2 in terms of the proportion of aligned reads. Figure 4 Although the more relaxed bwa-mem configuration aligned more reads, this came at the expense of increased misalignments (antisense strand alignment reads were used as an indicator of misalignments) ( Figure 4 C, see also Table 5). When using bwa-mem-W13-k6-xont2d-T20, a sweet spot was found that aligned 54.63% of the reads with very few misalignments (0.19%). The contrast was even more stark when comparing the alignment capabilities of bwa-mem to minimap2 in natural tRNA molecules; although minimap2 aligned the IVT tRNA, it failed to align a single biological tRNA read ( Figure 4 For reads aligning to IVT tRNA, the alignment identities were comparable to those of minimap2 but slightly lower than those typically obtained in nanopore DRS runs, suggesting that even in the absence of modifications, short reads can lead to reduced base calling accuracy ( Figure 4 E). Notably, Saccharomyces cerevisiae tRNA Phe The alignment consistency of tRNA (about 74.5%) is lower than that of synthetic tRNA (81.8%), which may be due to the endogenous modification of Saccharomyces cerevisiae tRNA. Phe There are base modifications on Figure 4 of E, see also Table 6).
[0376] We then quantified whether the alignability of Nano-tRNAseq reads was affected by the length of the 5' and 3' RNA adapters. To simulate different RNA adapter lengths, we removed one or both of these adapters from the reference sequence. We found that the absence of these two RNA adapters had only a modest effect on the alignability of reads derived from IVT tRNAs (a decrease of 6-11%) ( Figure 4 F, see also Table 7), while the number of reads from natural Saccharomyces cerevisiae tRNA aligned to the reference was reduced by 55%. Therefore, even without RNA adapters in the reference sequence, short and unmodified sequences can be efficiently aligned, while short and modified reads benefit greatly from adapter extension, demonstrating that extending molecules with RNA adapters is critical for guiding the correct alignment of "mismatch"-enriched short reads (such as those derived from natural tRNAs). For natural Saccharomyces cerevisiae tRNA Phe , the read segment alignability reaches a plateau when the 3' adapter length is 25 nt ( Figure 4 F), which demonstrates that the 30nt RNA portion of the 3' RNA:DNA we used in Nano-tRNAseq is fully sufficient to achieve optimal read segment alignability.
[0377]
[0378]
[0379]
[0380]
[0381] Customized MinKNOW parameters lead to 12-fold increase in sequencing yield
[0382] A surprising feature of our initial tRNA sequencing runs was the low amount of sequenced reads. Although pore blockage caused by tRNA structure may partially explain the low sequencing yield, we also noticed that the current MinKNOW software (July 22, 2022 and earlier versions) classified a high proportion of reads as "adapter-only" reads in real time. Therefore, we hypothesized that due to the short signal length of tRNA reads, a considerable portion of tRNA reads may be discarded by the MinKNOW software because they resemble "adapter-only" reads.
[0383] The MinKNOW software is responsible for analyzing the continuous current (signal intensity) measured at each hole, reporting the signal area corresponding to the "read segment" in the FAST5 file, and then performing base recognition to generate a FASTQ file. We note that by default, MinKNOW reports read segments that last at least 2 seconds, which is roughly equivalent to 140nt RNA molecules (assuming that the helicase processing rate in DRS is constant at 70nt / s). Considering that a typical tRNA molecule is about 73nt, this would imply that even after the double RNA adapter is connected (24 and 30 RNA nucleotides are added to the 5' end and 3' end of the tRNA molecule, respectively), the size of the connected tRNA molecule will still be below the threshold, which may cause the tRNA read segment to be incorrectly designated as a "only adapter" read segment. In order to alleviate this problem, we tested whether the alternative MinKNOW configuration would improve the classification of tRNA read segments and thus increase sequencing yield. To this end, bulk dump files were saved during sequencing runs and reprocessed using different MinKNOW configurations (Table 8).
[0384]
[0385]
[0386] By reducing the MinKNOW chain minimum duration to 1 second and the adapter maximum duration to 2 seconds (a configuration referred to herein as “custom”), we captured approximately 12 times more base-calling tRNA reads and approximately 4.5 times more uniquely aligned tRNA reads compared to the default MinKNOW configuration ( Figure 5 AB, see also Table 8). Notably, we found that the default MinKNOW configuration not only resulted in low sequencing output, but actually led to a significant bias in the relative abundance of tRNA molecules. Specifically, we found that shorter tRNAs were better represented in our custom configuration ( Figure 5 9), which indicates that the default MinKNOW configuration discards shorter tRNA molecules and prefers to capture longer tRNA molecules, such as tRNA molecules with variable arms (e.g., tRNA Leu tRNA Arg tRNA Ser ). In addition, the relative proportions of tRNA reads recapitulated using the custom configuration are much better than the default settings ( Figure 5 E), and the tRNA abundance reported using the custom settings (i.e., aligned tRNA reads) correlated very well with the expected value (ρ = 0.94) ( Figure 5 F, see also Tables 10-11).
[0387] Table 9. Fold changes of IVT tRNAs of different lengths when using default and custom MinKNOW read capture configurations.
[0388]
[0389] Table 10. Ratios of IVT tRNA and S. cerevisiae tRNAPhe when using default and custom MinKNOW read capture configurations.
[0390]
[0391] Table 11. Ratio of expected to observed IVT tRNAs when using a custom MinKNOW read capture configuration.
[0392]
[0393] Reverse transcription of tRNA increases sequencing yield and reduces pore clogging
[0394] Next, we questioned whether removing tRNA structure using reverse transcription (RT) would further improve the sequencing yield of the methods of the invention. Linear tRNA molecules could (i) reduce clogging of the pore, allowing more reads to be sequenced and maintaining the integrity of the flow cell for longer, and / or (ii) stabilize the rate of translocation of tRNA across the pore, improving the accuracy of the base recognition algorithm. We should note, however, that due to the tight secondary and tertiary structure of tRNAs and their abundance of modifications that disrupt Watson-Crick base pairing, it is difficult to completely and accurately reverse transcribe tRNAs. To examine whether tRNA linearization might improve sequencing yield, we tested a range of commercial reverse transcriptases and incubation conditions on both IVT tRNA and native tRNA and examined their cDNA output ( Figure 6 We found that at about 60°C, both Maxima and SuperScript IV provided the best performance in generating full-length cDNA products and chose to use Maxima at about 60°C in our subsequent tRNA sequencing experiments ( Figure 6 B).
[0395] Next, we examined whether linearization of tRNAs would increase our sequencing yield. To this end, total tRNAs from Saccharomyces cerevisiae were sequenced using Nano-tRNAseq with and without a reverse transcription step. Notably, the default MinKNOW configuration without reverse transcription conditions resulted in more reads than with reverse transcription conditions. We hypothesized that these results might be explained by the fact that nonlinearized tRNAs are highly structured, resulting in slower helicase processing of these molecules ( Figure 6C), and ultimately make them more likely to be classified as "reads" by the default MinKNOW configuration. Using a custom MinKNOW configuration ( Figure 5 BE), the number of base-calling reads with reverse transcription was actually about 1.5 times higher than without reverse transcription ( Figure 6 D, see also Table 12). Similarly, with reverse transcription, the number of reads uniquely aligned to tRNA increased by about 1.5 times, and the relative abundance of tRNA isoacceptors was not affected by the linearization step. Overall, linearization of tRNA molecules improves sequencing yield by increasing the helicase translocation rate.
[0396] Table 12. Base calling and alignment statistics for S. cerevisiae total tRNA sequenced with and without reverse transcription (RT).
[0397]
[0398]
[0399]
[0400] Accurate recapitulation of tRNA abundance using Nano-tRNAseq
[0401] Our results show that using IVT tRNA as input results in observed tRNA abundances that are highly similar to expected values (ρ = 0.94) when using Nano-tRNAseq with optimized alignment settings and a custom MinKNOW configuration ( Figure 5 F). We then wondered whether the observed abundances by Nano-tRNAseq correlated well with those by Illumina-based methods. To this end, we compared the S. cerevisiae tRNA abundances predicted using Nano-tRNAseq with those reported using three different Illumina-based methods: ARM-seq, Hydro-tRNAseq, and mim-tRNAseq. In ARM-seq, tRNAs are pretreated with the demethylase Escherichia coli AlkB, which removes m1A, m3C, and a portion of m1G modifications. Hydro-tRNAseq relies on partial alkaline RNA hydrolysis, which generates fragments suitable for sequencing. In the case of mim-tRNAseq, the authors improved the efficiency of cDNA synthesis by optimizing TGIRT reverse transcription conditions and allowing position-specific mismatch tolerance during read alignment.
[0402] Nano-tRNAseq correlates best with Illumina-based methods that address the presence of RT truncating modifications, namely ARM-seq (ρ = 0.555) and mim-tRNAseq (ρ = 0.525), and worst with Hydro-tRNAseq (ρ = 0.182). The low correlation with Hydro-tRNAseq may be due to the fact that (i) fragments containing such modifications are particularly short and unlikely to be PCR amplified, and (ii) alignment of fragmented samples is challenging and may lead to spurious tRNA counts. Overall, the correlation of Illumina-based methods with Nano-tRNAseq is generally low, which is not surprising given the huge differences in library preparation and analysis. We should note that Illumina-based tRNA sequencing methods also showed only moderate correlation with each other (ρ = 0.283-0.616). In any case, nanopore sequencing has the advantage of being cheaper and faster than NGS-based methods, and because it sequences native RNA molecules directly, it circumvents the need to remove modifications that interfere with reverse transcription. Additionally, it does not require PCR amplification, which is known to introduce unwanted variation in sequencing.
[0403] Nano-tRNAseq can be used to quantify differential tRNA modifications
[0404] Previous work has shown that base call errors or mismatches with a reference can be used to detect RNA modifications [Stephenson W, Razaghi R, Busan S, Weeks KM, Timp W, Smibert P. Direct detection of RNA modifications and structure using single molecule nanopore sequencing. doi: 10.1101 / 2020.05.31.126763. Leger A, Amaral PP, Pandolfini L, Capitanchik C, Capraro F, Miano V et al. RNA modifications detection by comparative Nanopore direct RNA sequencing. Nat Commun. 2021; 12: 7198. Pratanwanich PN, Yao F, Chen Y, Koh CWQ, Wan YK, Hendra C et al. Identification of differential RNA modifications from nanopore direct RNA sequencing with xPore. Nature Biotechnology.2021.pp.1394–1402.doi:10.1038 / s41587-021-00949-w]. Consistent with these observations, the tRNA Phe The mismatch errors shown are much more numerous than those seen in synthetic IVT tRNA ( Figure 3 B). Upon closer inspection, the positions of many of these mismatches coincide with specific RNA modifications, which can affect the signal at single-base resolution and may also affect the signal of adjacent bases. These findings suggest that structure or short read length are not the primary factors causing errors in base recognition of native tRNA reads, but rather the presence of RNA modifications.
[0405] Based on this premise, we aimed to verify that the mispairing we observed in natural tRNAs is indeed the result of RNA modification. To this end, we sequenced tRNAs from WT and Pus4-deficient Saccharomyces cerevisiae strains. Pus4 is the enzyme responsible for the synthesis of Ψ55 from U55 in the T-loop of tRNA. Upon knockout of Pus4, we observed a significant loss of the characteristic U to C mispairing of Ψ at position 55 in tRNA, while other known Ψ sites not reported to be catalyzed by Pus4 were unaffected ( Figure 7In the absence of Pus4, only Ψ55 of tRNATrp(CCA) remained unchanged, however we note that this tRNA does not contain the typical Pus4 motif RRUUCNA and is therefore likely not targeted by this enzyme. Despite the loss of Ψ55 in Pus4-deficient S. cerevisiae, we observed only modest changes in the expression of tRNA isoacceptors ( Figure 8 A).
[0406] Table 13. Sum of base recognition errors at known Ψ sites in WT and Pus4 KO Saccharomyces cerevisiae tRNA
[0407]
[0408]
[0409]
[0410] Nano-tRNAseq identifies tRNA modification interdependencies
[0411] tRNA modifications are introduced in a defined sequential order, and the temporal order is controlled by cross-talk between modification events and RNA modifying enzymes. Barraud et al., 2019 used time-resolved nuclear magnetic resonance (NMR) to monitor tRNA maturation and reported a robust modification hierarchy in the T-loop of Saccharomyces cerevisiae tRNAPhe, in which Ψ55 modifies m5U54 and m 1 The introduction of A58 has a positive impact, while m 5 U54 vs m 1 To explore whether our approach could capture the effects of Ψ55 loss on other modifications, we plotted the total base recognition errors (base mismatches, insertions, and deletions) for each Pus4 knockout tRNA isoacceptor relative to the WT for each nucleotide position and tRNA molecule reference ( Figure 7 C, see also Figure 8 In addition to the reduction in mismatch frequency at position 55, we also observed reductions at positions 54 and 57-59 ( Figure 7 LC-MS / MS was used to confirm that this observation was due to m 5 U54 and m 1 The actual reduction of A58 modification ( Figure 7 In addition, the IGV tracks show that Ψ mismatch errors are usually limited to a single base ( Figure 7 A), and m 1 A and m 5U mismatch error signals spill over to adjacent bases. Although Nano-tRNAseq does not allow us to dissect the specific order of modification circuits, it can reveal the interdependence of RNA modifications at different tRNA sites in the same transcript and quantify site-specific changes in RNA modifications across tRNA isoacceptors in a high-throughput manner with greater sensitivity than orthogonal methods while having the benefit of measuring tRNA abundance. Nano-tRNAseq was demonstrated to be able to recapitulate the loss of Ψ55 catalyzed by Pus4 and the subsequent m 1 A58 and m 5 The known relationship between the loss of U54 ( Figure 7 With the generation of yeast knockout strains for each tRNA modification enzyme nearing completion, the methods of the present invention provide an excellent opportunity to characterize tRNA modification circuits in a holistic manner, providing valuable insights into how these processes are regulated and impact health and disease.
[0412] Nano-tRNAseq reveals tRNA deadenylation during oxidative stress but not heat stress
[0413] tRNA modifying enzymes are dysregulated in a variety of human diseases and are altered under certain cellular conditions. Indeed, previous studies have provided evidence that tRNA modification profiles are reprogrammed under stress conditions, such as increased temperature and oxidative stress. However, chromatography-mass spectrometry-based methods do not provide information about which tRNA isoacceptor the modification originates from, nor do they provide information about the sequence context. To examine changes in tRNA abundance and modifications at single-nucleotide and single-tRNA isoacceptor resolution, we sequenced tRNAs from WT Saccharomyces cerevisiae exposed to heat or oxidative stress (see Methods). Although we found that tRNA abundance was highly reproducible across biological replicates (ρ = 0.981-0.993)( Figure 3 C is WT, Fig. 9 A is heat stress and oxidative stress, see also Table 14), but compared with WT, we only observed modest differences in the expression of tRNA under heat stress or oxidative stress ( Fig. 9 A, see also Tables 15 and 16). Furthermore, when we plotted the total and base misidentification frequencies relative to WT for each nucleotide position and tRNA isoacceptor, we observed no discernible changes in the RNA modification profile under any of the conditions tested ( Fig. 9 Our LC-MS / MS results confirmed this finding ( Figure 7 In contrast, we observed a large increase in the frequency of base misidentification at the last nucleotide, position 76 ( Fig. 9 The nucleotide corresponding to the terminal A of the CCA tail ( Fig. 9 The IGV track shows that the coverage of this terminal A relative to its neighboring bases is reduced ( Fig. 9 We then calculated the frequency of terminal A deletions for each tRNA isoacceptor and found that the frequency of deletions was significantly higher in tRNAs subjected to oxidative stress compared to WT, Pus4 KO, and heat stress, indicating that deadenylation of terminal A is a phenomenon specific to oxidative stress ( Fig. 9 This result is consistent with a previous study that also found that the terminal A of the 3'CCA tail is rapidly removed during oxidative stress, which is reversible and can be quickly repaired, making the tRNA chargeable again, representing a rapid mechanism to inhibit and reactivate translation at a low metabolic cost. The method of the present invention can demonstrate this rapid and dynamic inhibition of translation by tRNA deadenylation by quantifying the level of terminal A deadenylation at tRNA isoacceptor resolution.
[0414] Table 14. Read counts for each tRNA isoacceptor for WT, stressed, and Pus4 KO S. cerevisiae tRNAs.
[0415] Read segment (#)
[0416]
[0417]
[0418] Table 15. Differential expression analysis of tRNAs in oxidatively stressed Saccharomyces cerevisiae.
[0419]
[0420]
[0421] Differentially expressed tRNAs were inferred using DESeq2. Base mean (average normalized counts for all samples), log2 fold change (log2 fold change between groups), lfcSE (standard error of the log2 fold change), stat (Walk statistic), p value (Wald test p-value), p corrected (Benjamini-Hochberg corrected p value).
[0422] Table 16. Differential expression analysis of heat stressed Saccharomyces cerevisiae tRNAs.
[0423]
[0424]
[0425]
[0426] Differentially expressed tRNAs were inferred using DESeq2. Base mean (average normalized counts for all samples), log2 fold change (log2 fold change between groups), lfcSE (standard error of the log2 fold change), stat (Walk statistic), p value (Wald test p-value), p corrected (Benjamini-Hochberg corrected p value).
[0427] Custom MinKNOW configuration captures both shorter and longer molecules
[0428] The default MinKNOW configuration is not very good at detecting reads derived from short molecules such as tRNA. We developed an alternative MinKNOW configuration to capture about 10 times more reads from tRNA experiments. This is achieved by reporting molecules classified as adapters or strands as reads (see Methods).
[0429] Data decomposition of tRNA reads using nanopore direct RNA sequencing
[0430] The method of the present invention can generate millions of tRNA reads from a single sequencing reaction. The cost of library preparation and sequencing is expensive. For downstream analysis, tens to hundreds of thousands of reads will be sufficient. Therefore, for tRNA sequencing, it will be necessary to obtain a multiplex analysis method that allows sequencing of multiple samples in a single reaction.
[0431] The method for data splitting of tRNA reads consists of the following steps:
[0432] 1. Based on the alignment information, identify the "barcode" region by segmentation of the read segment
[0433] 2. Base identification of DNA barcode region
[0434] 3. Assigning barcodes based on alignment of base-call DNA barcode regions with a reference set of DNA barcodes
[0435] Step 1: Barcode identification (segmentation of reads based on alignment information)
[0436] In contrast to previous algorithms (e.g., DeepPlexiCon), the segmentation step does not rely on the presence of a poly A signal in the reads. Instead, we define here the end of the barcode as the signal corresponding to the last aligned base of the read.
[0437] In order to know the position of the last aligned base of a read position in signal space, the read must first be base called and aligned, and then the exact read start in signal space can be identified by checking which was the last aligned base reported by the basecaller. This information is extracted from the "move" table provided by the basecaller.
[0438] Step 2: DNA barcode base calling using Guppy
[0439] Next, we developed a custom function for base calling DNA adapters from direct RNA sequencing libraries. We should note that RNA reads (i.e., those produced by direct RNA sequencing) are typically base called using an RNA base calling model, and therefore, they remove DNA adapter and barcode regions that cannot be base called under the "RNA base calling model." Therefore, we used the DNA model trained by Bonito to base call the DNA adapter regions of reads sequenced using direct RNA sequencing. The base calling itself was performed using guppy, but a pre-trained DNA model was used as the DNA base calling model, which was trained in-house using Bonito (see below for details on training this DNA model).
[0440] Step 3: Barcode assignment by sequence alignment
[0441] Base call barcode sequences were aligned to the barcode reference sequence using minimap2 with the following parameters: -xmap-ont-k6-w3-n1-m10-s13-A1-B1-O1-E1. Only the best, unique (alignment quality above 0) alignments were reported. Only barcodes that were uniquely aligned to the barcode reference were kept as "true". Non-uniquely aligned barcodes were reported as unknown (bc_0).
[0442] In a specific embodiment, the software for performing steps 1, 2 and 3 in a separate program can be used. This integrated design is much faster than performing all operations in sequence. For example, 1CPU and 1GPU can be used to perform base recognition, alignment and data splitting on 4K reads in 2-5 seconds. In order to perform the same operation in sequence, it is necessary to store base recognition FAST5 reads, which takes up a lot of space, and once data splitting has been completed, in fact, any other steps do not need them. Briefly, the sequential steps performed by this software are:
[0443] -i) tRNA base recognition using guppy_baserecognition via the pyguppy client (part of step 1 above)
[0444] -ii) Align the tRNA sequence to the reference using minimap2 via mappy (part of step 1 above)
[0445] -iii) Perform barcode recognition in signal space (part of step 1 above)
[0446] -iv) Barcode base calling via our custom barcode base calling model (CTC-CRF model trained from bonito) (step 2 above)
[0447] -v) Perform barcode classification by mapping barcode sequences to barcode references using minimap2 via mappy (step 3 above)
[0448] Evaluating the performance of tRNA data splitting algorithms
[0449] We tested 7 independent Nano-tRNAseq runs, in which each human tRNA sample was sequenced with a separate barcode. Overall, the method achieved an accuracy of 0.979 and a recovery of 0.996. bc_82 performed the worst (accuracy of 0.958). The statistics for each barcode can be found in Table 17 below.
[0450] Table 17: Oligonucleotides used in multiplex assays
[0451]
[0452] Shown in bold – barcode
[0453] Shown underlined – oligodT overhang required for annealing to 3' adaptor annealed tRNA template via splint. Default ONT library preparation uses this overhang (originally intended for capturing mRNA).
[0454] Table 18. Validation of data splitting methods using an independent Nano-tRNAseq dataset
[0455]
[0456] The model achieved an accuracy of 0.978 and a recovery of over 0.999 on the validation set. Accuracy per barcode: bc_2 0.982, bc_3 0.985, bc_4 0.977, bc_23 0.98, bc_48 0.979, bc_82 0.962
[0457] Comparison with using DeePlexiCon for data decomposition of tRNA reads
[0458] DeePlexiCon is performed in 3 steps: i) barcode identification (segmentation) in signal space, ii) signal to image conversion (GASF) and barcode classification using convolutional neural networks (ResNet) commonly used in image analysis. Currently, DeePlexiCon allows up to 4 samples to be merged (pool) in one sequencing reaction, reducing sequencing costs by nearly 4 times. DeePlexiCon achieves 95% accuracy and 93% recovery (Smith MA, Ersavas T, Ferguson JM et al. Molecular barcoding of native RNAs using nanopore sequencing and deeplearning. Genome Res. 2020; 30 (9): 1345-1353. doi: 10.1101 / gr.260836.120). However, DeePlexiCon performs very poorly for tRNA samples, achieving only an accuracy of 0.696 (Table 19), making it completely unusable for tRNA multiplex analysis. The most likely explanation for this low performance is a segmentation problem caused by the very short poly A tails (10 bases) used in the Nano-RNA seq protocol. The segmentation method used in DeePlexiCon was developed for long mRNA reads and those reads that tend to have much longer poly A tails (usually in the range of hundreds of bases).
[0459] Table 19. Accuracy of data segmentation of tRNA reads using DeepPlexiCon
[0460]
[0461]
[0462] Training the Bonito DNA base recognition model
[0463] We rely on CTC-CRF models trained with bonito, using a chunksize of 3000 signal samples, using windows of 31 and stride of 10. CTC-CRF models consisting of 25 million (sup) or 500,000 (fast) parameters were trained with bonito for 50 epochs with a learning rate of 2e -3 , using a total of more than 500,000 blocks for the 6 barcodes. 3% of the dataset was used as a validation set. The best performance on the validation set was achieved after 32 training epochs, which we used as the final barcode base calling model.
Claims
1. A method for quantifying tRNA abundance and tRNA modification, comprising the following steps: a) contacting an RNA sample containing tRNA with the following substances in the presence of a reagent having ligation activity: A pre-annealed splinted double-stranded oligonucleotide comprising: - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3' end, - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region, such that at the end of said step, the first splint oligonucleotide is adjacent to the 5' end of the tRNA, and is complementarily annealed to the 3' end of the tRNA through the tRNA hybridization region, and is complementarily annealed to the second splint oligonucleotide through the second splint oligonucleotide hybridization region, b) contacting the product of step a) with a pre-annealed adapter DNA double-stranded oligonucleotide in the presence of a reagent having ligation activity, comprising: - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to the second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide, such that at the end of said step, the first adaptor DNA oligonucleotide is adjacent to the 3' end of the terminal region of the second splint RNA oligonucleotide and the second DNA oligonucleotide is complementarily annealed to both the second splint RNA oligonucleotide and the first adaptor DNA oligonucleotide, c) performing reverse transcription to linearize the product of step b) to obtain a library, and d) performing nanopore direct sequencing using the library of step c) to obtain the abundance of tRNA and its modification in the RNA sample.
2. The method according to the preceding claim, wherein step d) comprises: - contacting the library of step c) with an oligonucleotide adaptor configured for nanopore direct sequencing in the presence of an agent having ligation activity, - loading the product of the previous step into a flow cell, said flow cell comprising a membrane in which a nanopore is present, said nanopore providing a passage through said membrane, coupled to an electric current intensity, wherein the product of the previous step passes through said nanopore, resulting in a perturbation of said electric current intensity, and - Analyzing said sequences to obtain the abundance of tRNA and its modifications in said sample.
3. A method according to any one of the preceding claims, wherein the ligating agent of step a) is an enzyme, preferably Escherichia coli (E. coli) T4 RNA ligase 2 having SEQ ID N°2 or a variant having at least 80% identity, more preferably 85% identity, even more preferably at least 90% identity and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity relative to the SEQ ID No 2 polynucleotide sequence.
4. The process according to any one of the preceding claims, wherein the duration of step a) is from 1 hour to 3 hours at a temperature between 20°C and 26°C.
5. The method according to any of the preceding claims, wherein the second adaptor DNA oligonucleotide hybridizing region of the second splint oligonucleotide is any sequence of at least 10 nucleotides complementary to the second adaptor DNA oligonucleotide hybridizing region of the second splint oligonucleotide, preferably a poly A tail of at least 10 A nucleotides.
6. A method according to any one of the preceding claims, wherein the second splint oligonucleotide is a RNA:DNA oligonucleotide, preferably having between 1 and 15 DNA nucleotides from the 3' end of the oligonucleotide.
7. The method according to any one of the preceding claims 1 to 6, wherein the first RNA splint oligonucleotide is SEQ ID No 3 and the second splint oligonucleotide is SEQ ID No 4 or SEQ ID No 5 Or a variant having at least 80%, preferably at least 85%, more preferably at least 90% and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity to SEQ ID No 3, SEQ ID No 4, SEQ ID No 5.
8. The method according to any one of the preceding claims 2 to 7, wherein the parameters of the software configured to analyze the sequencing results of the nanopore direct sequencing are adjusted to a maximum of 1 second definition for adapters and a maximum of 2 seconds for strands.
9. The method according to any one of the preceding claims 2 to 8, wherein the parameters used to analyze the nanopore direct sequencing results are configured using BWA, and the parameters are selected from the group consisting of: i) bwa mem-W13-k6-xont2d, ii) bwa mem-W13-k6-xont2d-T20, iii) bwamem-W13-k6-xont2d-T10, iv) bwamem-W9-k5-xont2d-T10 and v)bwasw-z10-a2-b1-q2-r1, Preferred is bwa mem-W13-k6-xont2d-T20.
10. The method according to any one of the preceding claims 1 to 7, wherein the hybridization region of the first and second DNA adaptor oligonucleotides is a random nucleotide sequence of between 15 and 45 nucleotides that is unique to each pair of first and second DNA adaptor oligonucleotides.
11. The method according to the preceding claim, wherein more than one RNA sample containing tRNA is treated simultaneously by step d).
12. The method according to any of the preceding claims 10 or 11, wherein the algorithm implemented for analyzing the nanopore direct sequencing results is minimap2 with the adjusted parameters: -xmap-ont-k6-w3-n1-m10-s13-A1-B1-O1-E1.
13. A kit for carrying out the method of any one of claims 1 to 12, comprising: - a first splint oligonucleotide comprising a second splint oligonucleotide hybridizing region and a tRNA hybridizing region, said tRNA hybridizing region being located at its 3′ end, - a second splint oligonucleotide comprising a first splint oligonucleotide hybridizing region, a second adaptor DNA oligonucleotide hybridizing region, - a first adaptor DNA oligonucleotide comprising a second DNA adaptor oligonucleotide hybridizing region, and - a second adaptor DNA oligonucleotide comprising a first DNA adaptor oligonucleotide hybridizing region and a second splint oligonucleotide hybridizing region, said second splint oligonucleotide hybridizing region being complementary to said second adaptor DNA oligonucleotide hybridizing region of said second splint oligonucleotide.
14. Kit according to the preceding claim, wherein the second splint oligonucleotide is a RNA:DNA oligonucleotide having at least one terminal DNA base at its 3' end.
15. A kit according to any one of the preceding claims 13 to 14, comprising T4 RNA ligase 2 having an oligonucleotide sequence of SEQ ID N°1 or SEQ ID N°2, or a variant having at least 80% identity, more preferably 85% identity, even more preferably at least 90% identity and even more preferably 91% or 92% or 93% or 94% or 95% or 96% or 97% or 98% or even up to 99% identity relative to the sequence of SEQ ID No 1 or SEQ ID No 2.
Citation Information
Patent Citations
Characterization of individual polymer molecules based on monomer-interface interactions
EP0815438B1
Characterization of hybridized polymer molecules based on monomer-interface interactions
EP1238275B1