Method for direct RNA nanopore sequencing based on reverse transcriptase

WO2026201203A1PCT designated stage Publication Date: 2026-10-01GENEUS TECH CHENGDU CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/087218
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-28
Filing Date
2026-03-30
Publication Date
2026-10-01

Smart Images

  • Figure CN2026087218_01102026_PF_FP_ABST
    Figure CN2026087218_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an RNA sequencing method, which comprises performing sequencing-by-synthesis by means of using a reverse transcriptase or a mutant thereof. Further provided are modified reverse transcriptases or sequencing complexes for use in the method, and a sequencing kit comprising the reverse transcriptases or the sequencing complexes.
Need to check novelty before this filing date? Find Prior Art

Description

A direct RNA nanopore sequencing method based on reverse transcriptase

[0001] This application claims priority to Chinese patent application No. 202510386465.3, filed with the Chinese Patent Office on March 28, 2025, the entire contents of which are incorporated herein by reference. Technical Field

[0002] This disclosure relates to the field of nucleic acid sequencing, and in particular to a direct RNA nanopore sequencing method that applies an improved reverse transcriptase R2Bm to a "synthesis-as-sequencing" scheme. Background Technology

[0003] RNA sequencing has broad prospects for biological applications, including bacterial and viral identification, differential splicing analysis, and differential gene expression analysis. Currently, mainstream RNA sequencing routes all require reverse transcription of RNA into cDNA followed by PCR amplification before sequencing. This not only increases the sample preparation process but also introduces systematic errors, including reduced sample complexity and distortion of relative abundance. Furthermore, due to PCR amplification, if barcode library construction is not used and PolyT primers or random primers are used directly, strand-specific sequencing cannot be achieved, meaning it's impossible to determine whether the obtained sequence originates from RNA or its complementary sequence. If reverse transcription and PCR amplification can be avoided during sample preparation, these drawbacks can be circumvented, leading to the development of direct RNA sequencing technology.

[0004] Between 2009 and 2010, researchers achieved direct RNA sequencing based on the "synthesis-as-sequencing" approach on the Helicos and Illumina platforms (Ozsolak, F. et al. Direct RNA sequencing. Nature 461, 814-818, (2009); Mamanova, L. et al. FRT-seq: amplification-free, strand-specific transcriptome sequencing. Nat Methods 7, 130-132, (2010)). However, at that time, both RNA sequencing routes were similar to NGS sequencing and could not overcome the drawbacks of discontinuous reactions and short read lengths.

[0005] Besides sequencing-by-synthesis (SBS) approaches, Oxford Nanopore Technologies (ONT) offers a nanopore sequencing solution that can also achieve direct RNA sequencing (US Patent Publication US20170253923A1, METHOD FOR NANOPORE RNA CHARACTERISATION). This solution uses DNA helicase as the kinetic protein to guide RNA through nanopores and uses algorithms to analyze the pore current to infer the RNA sequence. While this solution achieves direct RNA sequencing, it requires two adapter ligations, involves multiple sample preparation steps, and has low accuracy in identifying homopolymers. Similarly, a preprint article published in 2023 (https: / / www.biorxiv.org / content / 10.1101 / 2023.04.05.535757v2) describes an RNA sequencing route using reverse transcriptase R2Bm as the kinetic protein to guide RNA through nanopores. This route also suffers from the problem of insufficiently intuitive signal delivery.

[0006] In 2016, Genia (now acquired by Roche) proposed a "synthesis-as-sequencing" scheme based on nanopore detection. Its working principle involves coupling DNA polymerase with nanopore proteins, then using nucleotides with current-blocking groups as substrates for the polymerase to synthesize new DNA strands. These current-blocking groups can be captured by the nanopore under the influence of an electric field, altering the current flowing through the nanopore. Different current-blocking groups result in different changes to the nanopore current. By differentiating four types of nucleotides using different current-blocking groups, DNA sequence determination can be achieved. Based on this model, combined with long-progressive DNA polymerases (such as phi29), real-time sequencing of single-molecule, long-read DNA can be achieved (Stranges PB et al. Design and characterization of a nanopore-coupled polymerase for single-molecule DNA sequencing by synthesis on an electrode array. Proc Natl Acad Sci US A. 2016 Nov 1;113(44):E6749-E6756.). In theory, by simply replacing DNA-dependent DNA polymerase with RNA-dependent DNA polymerase (i.e., reverse transcriptase), a similar nanopore platform can be used to achieve real-time direct determination of RNA sequences. As the driving core of this system, the reverse transcriptase needs to meet the following requirements: 1. Efficient utilization of modified nucleotides. To date, all single-molecule "synthesis-while-sequencing" schemes require special modification of the nucleotides involved in synthesis to obtain base signals. The polymerase's efficient utilization of modified nucleotides is a prerequisite for achieving a certain sequencing efficiency and obtaining sufficient read lengths. 2. Kinetic parameters matched to the detection performance of the nanopore. If the reverse transcription synthesis reaction is completed too quickly, the sequence signal may not have enough time to be captured, or it may be difficult to distinguish from interference signals, resulting in sequencing errors. Conversely, if the synthesis reaction is completed too slowly, it not only reduces sequencing efficiency but also easily causes nucleotides to spontaneously dissociate from the polymerase before the synthesis reaction is complete, generating invalid signals in the sequencing process and reducing sequencing accuracy. 3. Good thermal stability. In a sequencing system where a reverse transcriptase molecule corresponds to one RNA template, the reverse transcriptase must possess a certain degree of thermal stability to maintain its activity for a sufficient period under sequencing conditions. Premature denaturation and inactivation due to poor thermal stability are fundamental to the continuous operation of the sequencing platform and the acquisition of sufficient read lengths and sequencing data. 4. Other characteristics of reverse transcription include salt tolerance, fidelity, and operating temperature compatibility with the sequencing platform. Summary of the Invention

[0007] On the one hand, this paper utilizes RNA sequencing methods, including sequencing-by-synthesis using reverse transcriptase or its mutants.

[0008] In some implementations, the RNA sequencing is nanopore sequencing.

[0009] In some implementations, the RNA is linear or circular RNA.

[0010] In some embodiments, the reverse transcriptase is R2Bm or PsRT.

[0011] In some embodiments, the reverse transcriptase or a mutant thereof is coupled to the nanoporin via one, two or more sites.

[0012] In some embodiments, the coupling site is located at the C-terminus, N-terminus, and / or middle of the amino acid sequence of the reverse transcriptase or its mutant form.

[0013] In some embodiments, the nanoporin is NetB, αHL, or a mutant thereof.

[0014] In some embodiments, the reverse transcriptase R2Bm or its mutants are missing 1-288 amino acid residues at its N-terminus, for example, missing amino acids 1-20, 1-25, 1-50, 1-80, 1-100, 1-150, 1-200, 1-220, 1-230, 1-240, 1-250, 1-260, 1-270, and 1-280 at the N-terminus, such as missing amino acids 1-251, 1-252, 1-253, 1-254, 1-255, 1-256, 1-257, 1-258, 1-259, 1-260, 1-261, 1-262, 1-263, 1-264, 1-265, 1-266, 1-267, 1-268, 1-269, 1-270, and 1-274 at the N-terminus. Amino acids at positions 1-272, 1-273, 1-274, 1-275, 1-276, 1-277, 1-278, 1-279, 1-280, 1-281, 1-282, 1-283, 1-284, 1-285, 1-286, and 1-287, or amino acids at positions 10-200, 20-250, and 50-280 at the N-terminus, etc.

[0015] In some embodiments, the reverse transcriptase R2Bm mutant comprises any one or any combination of the following mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F.

[0016] In some embodiments, the reverse transcriptase or its mutant and the nanoporin are coupled by covalent or non-covalent bonds.

[0017] In some implementations, the covalent bond includes a SpyCatcher-SpyTag bond.

[0018] In some embodiments, the non-covalent bond includes avidin-biotin binding, Cl7-Im7 binding, or BimBinder-BimTag binding.

[0019] In some embodiments, the reverse transcriptase R2Bm or its mutant has SpyCatcher / SpyTag / BimBinder / Cl7 fused to its N / C ends, and at least one subunit of the nanoporin NetB or its mutant includes SpyTag / SpyCatcher / BimTag / Im7.

[0020] In some embodiments, SpyCatcher / SpyTag / BimBinder / Cl7 is fused to the N-terminus of the nanoporin NetB, or SpyTag is linked to the nanoporin NetB or its mutants via cysteine ​​residues contained in the nanoporin NetB or its mutants.

[0021] In some embodiments, the nanoporin NetB mutant includes the I27C, D96C, or N236C mutation.

[0022] In some embodiments, the reverse transcriptase R2Bm or its mutant is fused with SpyCatcher / SpyTag / BimBinder / Cl7 / antibiotin, and at least one subunit of the nanoporin NetB or its mutant includes SpyTag / SpyCatcher / BimTag / Im7.

[0023] In some embodiments, the Cl7 / avidin replaces positions 297-303, 670-676, or 376-383 of the reverse transcriptase R2Bm or its mutant; or is inserted at any position of 297-303, any position of 670-676, or any position of 376-383.

[0024] In some embodiments, the Im7 is fused to the N / C terminus of the nanoporin NetB or its mutant.

[0025] In some embodiments, the reverse transcriptase R2Bm or a mutant thereof is also fused with a solubilizing tag, such as TrxA or MBP.

[0026] In some embodiments, the reverse transcriptase is PsRT or a mutant thereof, and the nanoporin is αHL or a mutant thereof.

[0027] In some embodiments, the PsRT mutants include mutants D15R, L63R, K233A, A234D and / or E257R.

[0028] In some embodiments, the PsRT mutant is further fused with avidin at the N-terminus and SpyCatcher at the C-terminus.

[0029] In some embodiments, the nanoporin αHL is fused to its C-terminus with a SpyTag and includes an N17C mutation.

[0030] In some embodiments, the nanoporin αHL is linked to biotin at the N17C mutation site.

[0031] In some embodiments, the reverse transcriptase and nanoporin include a purification tag, such as HisTag, at the N-terminus or C-terminus.

[0032] In some embodiments, the amino acid sequence of the reverse transcriptase R2Bm is as shown in SEQ ID NO:1, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:1.

[0033] In some embodiments, the amino acid sequence of the reverse transcriptase R2Bm mutant is as shown in any one of SEQ ID NO:2-34, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:2-34.

[0034] In some embodiments, the amino acid sequence of at least one subunit of the nanoporin NetB mutant is as shown in any one of SEQ ID NO:36-45, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:36-45.

[0035] In some embodiments, the amino acid sequence of the PsRT mutant is as shown in SEQ ID NO:48, including an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:48.

[0036] In some embodiments, the amino acid sequence of the nanoporin αHL is as shown in SEQ ID NO:46, including an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:46.

[0037] In some implementations, the method is performed at room temperature.

[0038] In some implementations, the method is in Mn 2+ or Mg 2+ It is performed under the condition that it exists.

[0039] In some implementation schemes, Mn 2+ Concentrations of 0.2-1.2 mM, Mg 2+ The concentration is above 30mM.

[0040] In some implementations, the method is carried out in the presence of 350-650 mM KAc.

[0041] In some embodiments, the reverse transcriptase is capable of synthesizing DNA strands using modified nucleotides as substrates.

[0042] In some embodiments, the reverse transcriptase is R2Bm or a mutant thereof, and the structure of the modified nucleotide is shown below:

[0043] In some embodiments, the reverse transcriptase is PsRT or a mutant thereof, and the structure of the modified nucleotide is shown below:

[0044] On the other hand, this article provides a reverse transcriptase, which is a modified R2Bm, wherein the modification includes one or more of the following:

[0045] 1) Introduce any one or any combination of the following mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F;

[0046] 2) The N-terminus of 1-288 amino acid residues is missing;

[0047] 3) Introduce one or more coupling moieties that can bind to the coupling moieties on nanoporous proteins;

[0048] 4) Introduce solubilizing tags, such as TrxA or MBP;

[0049] 5) Introduce purification tags, such as HisTag.

[0050] In some implementations, the coupling portion is capable of forming covalent or non-covalent bonds with the nanoporous protein.

[0051] In some embodiments, the coupling portion is located at the C-terminus, N-terminus, and / or middle of the amino acid sequence of the reverse transcriptase or its mutant form.

[0052] In some implementations, the coupling portion is SpyCatcher or SpyTag, avidin or biotin, or Cl7 or Im7, BimBinder-BimTag.

[0053] In some implementations, the SpyCatcher, SpyTag, BimBinder, or Cl7 is located at the C or N end.

[0054] In some embodiments, the SpyCatcher, SpyTag, BimBinde, Cl7, or avidin replaces positions 297-303, 670-676, or 376-383; or is inserted at any position 297-303, any position 670-676, or any position 376-383; or replaces amino acids at positions 672-675, 378-381, or 301-304.

[0055] In some embodiments, the amino acid sequence of the R2Bm is as shown in SEQ ID NO:1, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:1.

[0056] In some embodiments, the amino acid sequence of the reverse transcriptase is as shown in any one of SEQ ID NO:2-34, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:2-34.

[0057] On the other hand, this article provides a reverse transcriptase, which is a modified PsRT, wherein the modification includes one or more of the following:

[0058] 1) Introduce mutations D15R, L63R, K233A, A234D and / or E257R;

[0059] 2) Introduce one or more coupling moieties capable of binding to coupling moieties on nanoporous proteins; and

[0060] 3) Introduce purification tags, such as HisTag.

[0061] In some implementations, the coupling portion includes avidin and / or SpyCatcher.

[0062] In some embodiments, the avidin is fused to the N-terminus of the reverse transcriptase, and the SpyCatcher is fused to the C-terminus of the reverse transcriptase.

[0063] In some embodiments, the amino acid sequence of the reverse transcriptase is as shown in SEQ ID NO:48, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:48.

[0064] On the other hand, this article provides sequencing conjugates, which include the aforementioned reverse transcriptase and the nanoporous protein conjugated therewith.

[0065] In some embodiments, the reverse transcriptase and the nanoporin bind through two or more coupling sites.

[0066] In some embodiments, the nanoporin is nanoporin NetB, αHL, or a mutant thereof.

[0067] On the other hand, this article provides an RNA sequencing kit that includes the aforementioned reverse transcriptase or sequencing conjugate.

[0068] In some implementations, the sequencing kit also includes modified nucleotides.

[0069] In some embodiments, the modified nucleotide is selected from one or more of the following compounds: Attached Figure Description

[0070] Figure 1 shows the synthetic route of the current blocking tag PA-15.

[0071] Figures 2A and 2B show the LC-MS detection spectra of the current blocking label PA-15.

[0072] Figure 3 shows the synthetic route of the PA-15 modified nucleotide dG5P-PA-15.

[0073] Figure 4 shows the synthetic route of the current blocking tag PA-16.

[0074] Figures 5A and 5B show the LC-MS detection spectra of the current blocking label PA-16.

[0075] Figure 6 shows the synthetic route of the PA-16 phase-modified nucleotide dC6P-PA-16.

[0076] Figure 7 shows the synthetic route of the current blocking tag PA-25.

[0077] Figures 8A and 8B show the LC-MS detection spectra of the current blocking label PA-25.

[0078] Figure 9 shows the synthetic route of the PA-25 modified nucleotide dT5P-PA-25.

[0079] Figure 10 shows the synthetic route of the current blocking tag PA-11.

[0080] Figures 11A and 11B show the LC-MS detection spectra of the current blocking tag PA-11.

[0081] Figure 12 shows the synthetic route of the PA-11 modified nucleotide dA5P-PA-11.

[0082] Figure 13 is a model diagram of a direct RNA sequencing biochemical system based on nanopore-based sequencing-by-synthesis.

[0083] Figure 14 shows the structure of the RNA7T template.

[0084] Figure 15A shows representative results of the elongation rate of R2Bm02 nucleotides modified with phosphate ends. Figure 15B shows representative results of the elongation rate of R2Bm02 mutant nucleotides modified with phosphate ends.

[0085] Figure 16 shows representative results of the heat resistance of R2Bm02.

[0086] Figure 17 shows representative results of the selectivity of R2Bm02 for catalytic metal ions.

[0087] Figure 18 shows representative results of the salt tolerance assessment of R2Bm02 at different KAc concentrations.

[0088] Figure 19 shows the structure of N-Maleimide-SpyTag.

[0089] Figure 20 shows the structure of Biotin-PEG3-Maleimide.

[0090] Figure 21 shows representative results of purity during the preparation of NetB01 monomer, NetB02 monomer conjugation modification, and NetB01+NetB02 conjugated protein.

[0091] Figure 22 shows representative results of purity during the preparation of PN from R2Bm08 mutant and NetB02 conjugate protein.

[0092] Figure 23 shows a schematic diagram of the sequencing apparatus of this disclosure.

[0093] Figure 24 shows the sequencing signal of linear RNA template T2 read by R2Bm08 coupled with NetB02 coupled with nanopore.

[0094] Figure 25 compares the sequencing results of R2Bm08 and R2Bm09 coupled with NetB02 nanopores for reading linear RNA template T2. (A) Sequencing results alignment of R2Bm08 coupled with NetB02 nanopores for reading linear RNA template T2; (B) DW average of the four base sequencing numbers of R2Bm08 coupled with NetB02 nanopores; (C) Sequencing results alignment of R2Bm09 coupled with NetB02 nanopores for reading linear RNA template T2; (D) DW average of the four base sequencing numbers of R2Bm09 coupled with NetB02 nanopores.

[0095] Figures 26A, 26B, and 26C together show the complete sequencing results alignment information for R2Bm08 coupled with NetB02 coupled nanopores.

[0096] Figure 27 shows the accuracy and read length distribution of reading circular RNA template T3 by R2Bm08 coupled with NetB02 coupled nanopores after 30 minutes.

[0097] Figure 28 shows the synthetic route of the DL-04 modified nucleotide dG5P-DL-04.

[0098] Figure 29 shows the LC-MS detection spectrum of the DL-04 modified nucleotide dG5P-DL-04.

[0099] Figure 30 shows the synthetic route of the HC-10 modified nucleotide dC6P-HC-10.

[0100] Figures 31A and 31B show the LC-MS detection patterns of the HC-10 modified nucleotide dC6P-HC-10.

[0101] Figure 32 shows the synthetic route of the HC-05 modified nucleotide dT5P-HC-05.

[0102] Figures 33A and 33B show the LC-MS detection patterns of the HC-05 modified nucleotide dT5P-HC-05.

[0103] Figure 34 shows the synthetic route of the HC-03 modified nucleotide dA5P-HC-03.

[0104] Figures 35A and 35B show the LC-MS detection patterns of HC-03 modified nucleotide dA5P-HC-03.

[0105] Figure 36 shows the sequencing signal of linear RNA template T1 read by PsRT01 coupled with αHL02 nanopore. Detailed Implementation

[0106] Unless otherwise stated, all technical and scientific terms used herein have the meanings commonly understood by one of ordinary skill in the art.

[0107] As used herein, the terms “about” or “around” mean a value that may deviate from the mentioned value by up to 1%, up to 5%, up to 10%, up to 15%, and in some cases up to 20%. The deviation range includes integer values ​​and, where applicable, non-integer values, forming a continuous range.

[0108] The term "or" refers to a single element among the listed optional elements, unless the context explicitly indicates otherwise. The term "and / or" refers to any one, any two, any three, any more, or all of the listed optional elements; when there is a " / " between two parallel elements, it indicates that the two elements are in an "and / or" relationship.

[0109] The terms “comprising,” “containing,” “having,” “including,” and similar expressions used herein do not exclude elements not listed. These terms also include cases where the text consists only of the listed elements.

[0110] The term "peptide," used interchangeably with "protein," refers to a biomolecule composed of amino acids linked together by peptide bonds. The term "peptide fragment" is used to indicate that it is part of a larger protein molecule, and can itself be a protein or even a fusion protein. These peptide fragments can be linked together through one or more coupling sites.

[0111] When referring to proteins, the term "mutant" refers to a variant of a parent protein (such as the original or native protein) that involves the substitution, addition, or deletion of one or more amino acids relative to the parent protein. "Amino acid substitution," also known as "amino acid replacement," in this context refers to the replacement of one amino acid residue (called the original amino acid residue) at a specific position in the amino acid sequence of a protein by another amino acid residue (the substituted amino acid residue). For example, if the second amino acid in the parent protein is glutamic acid (E), and the corresponding position in the mutant protein is alanine (A), then an amino acid substitution is considered to exist in the mutant protein: alanine replaces glutamic acid. For amino acid substitutions, the following nomenclature is used in this context: original amino acid, position, and substituted amino acid, and the amino acid name uses the single-letter abbreviations for amino acids as defined by IUPAC. For the example described above, it can be represented as E2A. "Amino acid deletion," also known as "amino acid deficiency," in this context refers to the absence of an amino acid residue at a specific position in the mutant protein compared to the parent protein. This position may appear as a gap when compared to the parent protein. "Amino acid addition," also known as "amino acid insertion," in this article refers to an amino acid residue at a certain position in the mutant that has no corresponding amino acid residue in the parent protein compared to the parent protein. In other words, amino acid insertion means that one or more amino acid residues are added when compared to the parent protein. The sequence identity between the mutant and the parent protein is usually above 80%, such as above 85%, 90%, 95%, 96%, 97%, 98%, or even above 99%.

[0112] The term "sequence identity" (also known as "sequence uniformity") refers to the degree of similarity between two amino acid sequences (e.g., a query sequence and a reference sequence), typically expressed as a percentage. Generally, sequence alignment is performed and gaps (if any) are introduced before calculating the percentage of similarity between two amino acid sequences. If amino acid residues or bases in the two sequences are the same at a given alignment position, the two sequences are considered to be identical or matched at that position; if amino acid residues or bases in the two sequences are different, they are considered to be inconsistent or mismatched at that position. In some algorithms, the number of matched positions is divided by the total number of positions in the alignment window to obtain sequence identity. In other algorithms, the number of gaps and / or gap length are also taken into account. Commonly used sequence alignment algorithms or software include EMBOSS, DANMAN, CLUSTALW, MAFFT, BLAST, MUSCLE, etc. For the purposes of this invention, the BLAST algorithm is used in some embodiments for amino acid sequence alignment with preset software parameters.

[0113] The term "coupling" refers to a connection formed between proteins or peptides in any manner, which can be covalent or non-covalent. Covalent coupling includes peptide bond connections or other forms of covalent bond connections. In some embodiments, two proteins or peptides are linked by peptide bonds, which can form between the main chains or between side chains (i.e., heteropeptide bonds, for example, formed via spytag-spycatcher pairs). Peptide bonds can form directly between the two proteins or peptides or through other short peptides. In other embodiments, coupling between two proteins or peptides can be achieved using click chemistry techniques. For example, using existing biochemical methods, a Cys mutation is introduced at a specific site on the first protein or peptide, and a DBCO (dibenzocyclooctylene) group, commonly used in click chemistry, is introduced onto the Cys group via maleimide mediation; simultaneously, another click chemistry active group, an N3 (azide) group, is introduced onto the second protein or peptide via a non-natural amino acid-mediated approach; then, under mild conditions, the paired click chemistry active groups can spontaneously form a stable covalent bond, achieving covalent coupling between the two. Non-covalent coupling includes protein or peptide linkages formed through antigen-antibody binding, ligand-receptor binding, enzymes and their inhibitors, etc. Non-covalent high-affinity protein interactions, such as the high-affinity protein pair of Cl7 and Im7, can serve as an effective coupling method to support the formation of relatively stable complexes between two protein molecules. A first protein or peptide terminated with Im7 and a second protein or peptide terminated with Cl7 can be readily obtained through fusion expression. The interaction between Cl7 and Im7 can facilitate non-covalent coupling between the first protein or peptide and the second protein or peptide. In other embodiments, coupling between the first protein or peptide and the second protein or peptide can also be based on the tight binding between biotin and avidin. Biotin and avidin tetramer (streptavidin) have extremely high affinity and are considered to have a near-semi-covalent interaction. In some embodiments, the selected avidin is the monomeric protein Monomeric Rhizavidin (hereinafter abbreviated as MR). Although the affinity between monomeric avidin and biotin is reduced compared to the complete tetramer of streptavidin, it is still sufficient to support efficient coupling at sub-nM concentrations. In the case of the reverse transcriptase and nanoporous protein involved in this invention, coupling between them can be achieved through any of the methods described above.

[0114] The term "coupling moiety" refers to a specific structure used to form a coupling relationship between two proteins or polypeptides. In some embodiments, the coupling relationship is formed through amino acid residues of the two proteins or polypeptides themselves, such as peptide bonds, in which case the coupling moiety is an amino acid residue. In other embodiments, the coupling relationship is formed by introducing a group or peptide for forming the coupling relationship into one or both proteins, such as introducing biotin and streptavidin into two proteins, respectively, in which case the coupling moiety can be considered as biotin or streptavidin. Since one, two, or more coupling moieties can be present or introduced on reverse transcriptase and nanoporin (or its subunits), one, two, or more coupling sites can be formed between reverse transcriptase and nanoporin (or its subunits). The term "sequencing conjugate" refers to a complex formed by the coupling of reverse transcriptase and nanoporin, which can be used for RNA sequencing.

[0115] The term "solution tag" refers to a specific polypeptide or protein sequence designed to improve the solubility of a target protein, reduce inclusion body formation, or promote proper protein folding. Solubility tags are typically expressed fused to the target protein. Examples of solubility tags include maltose-binding protein (MBP), thioredoxin (Trx), and the SUMO tag.

[0116] The term "purification tag" refers to a specific polypeptide or protein sequence used to efficiently separate and purify a target protein using methods such as affinity chromatography. Purification tags are typically expressed fused to the target protein. Examples of purification tags include His tags and GST tags. Some polypeptide or protein sequences possess both solubilizing and purification capabilities, such as GST tags and Strep tags.

[0117] When referring to the reverse transcriptase as an R2Bm mutant, the amino acid position is based on the wild-type reverse transcriptase, that is, the amino acid position is referenced to the sequence shown in SEQ ID NO:1, even though the reverse transcriptase as an R2Bm mutant may have an amino acid deletion at the N-terminus.

[0118] When describing the location of other proteins or sequences inserted into R2Bm or its mutants, it should be understood that it can be inserted at the first amino acid position of the region and between any two remaining adjacent amino acid residues. For example, for the spycatcher insertion at positions 297-303, it includes insertion at position 297 (i.e., between positions 296 and 297), between positions 297 and 298, between positions 298 and 299, between positions 299 and 300, between positions 301 and 302, and between positions 302 and 303.

[0119] To address the shortcomings of existing technologies, this disclosure, based on the "synthesis-while-sequencing" approach, achieves for the first time real-time direct sequencing of single RNA molecules using nanopores. In this disclosure, samples can be directly used for sequencing without reverse transcription and PCR amplification, avoiding the introduction of related systematic errors. Compared to other RNA direct sequencing schemes based on "synthesis-while-sequencing," this scheme offers longer read lengths and is performed in real-time, potentially allowing for indirect inference of template structure and epigenetic information through extension speed. Compared to ONT's RNA direct sequencing technology, this scheme has a simpler sample preparation process, generates single-base-resolved signals that can more accurately identify base homopolymers, and can also achieve higher single-sequence accuracy through rolling circle reverse transcription.

[0120] This disclosure combines protein nanopore detection with a "synthesis-while-sequencing" technique, coupling reverse transcriptase to a nanopore and using nucleotides with four different current-blocking tags as raw materials. During reverse transcription, the four blocking tags sequentially pass through the nanopore, generating different characteristic blocking currents. By recording and analyzing these current changes, the sequence of the RNA to be tested can be deduced.

[0121] Specifically, this disclosure will originate from the silkworm (Bombyx mori). The R2 non-LTR retrotransposon R2Bm of mori) is introduced as a retrotranscriptional element with any one or any combination of the following mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F, to increase its speed or signal hysteresis time (dwell time). This disclosure first investigated the sequencing-related biochemical properties of the full-length R2Bm (specific sequence SEQ ID NO:3) with a solubilizing tag and coupling protein. Experimental results show that the R2Bm mutant provided in this paper exhibits excellent overall performance in terms of modified nucleotide substrate utilization, catalytic kinetic parameters, thermal stability, operating temperature, strand substitution function, and salt tolerance, making it suitable for single-molecule RNA sequencing technology based on a "synthesis-as-sequencing" protocol.

[0122] To efficiently capture positively charged current-blocking tags to obtain sequencing signals, this disclosure conjugates a modified NetB nanoporin mutant (assembled from a standard monomer NetB01 with the specific sequence SEQ ID NO:35 and conjugate monomers with different conjugate tags) with an R2Bm mutant carrying the corresponding conjugate tag, and demonstrates its sequencing accuracy. Specific conjugation schemes are shown in Table 7.

[0123] The advantages of this invention include at least the following:

[0124] 1. No RT-PCR amplification is required, reducing RNA sequence bias and information loss in sequencing systems;

[0125] 2. Compared to existing RNA direct sequencing schemes based on "synthesis-as-sequencing", it has longer read lengths and retains reverse transcription speed information (which may be used to infer RNA structure or epigenetic information).

[0126] 3. Compared to ONT's direct RNA sequencing protocol, this protocol eliminates the need for adapter ligation, simplifying sample preparation. The single-base resolution signals generated by this protocol facilitate more accurate reading of base homopolymers.

[0127] 4. This disclosure is currently the only method that can directly sequence circular RNA, and the accuracy of sequencing results can be improved by combining rolling circle sequencing with sequence consistency correction.

[0128] R2Bm reverse transcriptase

[0129] This disclosure uses the R2Bm protein (GenBank: BAC01114.1) with the amino acid sequence SEQ ID NO:1 as a basis to optimize its coupling mode with nanopores and explore its potential application in a direct RNA sequencing platform based on "synthesis-as-sequencing". The R2Bm protein is derived from the R2 non-LTR retrotransposon of the silkworm (Bombyx mori), and its structure is deposited in the Protein Database (PDB) with the accession code 8gh6. Existing documents disclose its reverse transcription biological activity and its application in direct RNA sequencing routes where RNA passes directly through nanopores, but its application in direct RNA nanopore sequencing based on "synthesis-as-sequencing" has not been reported.

[0130] The modification of wild-type R2Bm protein in this disclosure includes the following aspects: (1) introducing the following single-point and arbitrary combination mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F, to increase its speed or signal hysteresis time (dwell time). (2) Fuse SpyCatcher, SpyTag, TrxA, MBP, avidin, BimBinder and other tag proteins to R2Bm protein to couple nanoporous proteins or increase their soluble expression level; (3) Remove some amino acid residues from the N-terminus of wild-type R2Bm in some mutants to obtain a more suitable coupling position.

[0131] Nucleic acid sequencing methods

[0132] Using the R2Bm reverse transcriptase disclosed herein, this disclosure provides a method for RNA sequencing.

[0133] In order to efficiently capture positively charged current blocking tags, this disclosure takes a NetB mutant with different coupling tags that has been modified accordingly as an example, and couples it with an R2Bm mutant with the corresponding coupling tag.

[0134] This disclosure describes the fusion expression of a coupling tag at a specific position on a NetB monomer or the coupling tag modified by Cys, as detailed in Table 7 (hereinafter referred to as the coupling monomer). The coupling monomer is mixed with a common monomer with the specific sequence SEQ ID NO:35 and assembled. After separation and purification, a 1:6 nanoporous protein containing one copy of the coupling monomer and six copies of the common mutant is obtained (hereinafter referred to as the "coupled nanopore"). By progressively mixing R2Bm mutants with different coupling tags, primer-bound test RNA, and coupled nanopores, a ternary complex formed by coupled nanopores, coupled reverse transcriptase, and test RNA is obtained (hereinafter referred to as the RPN complex).

[0135] The RPN complex is loaded into the sequencing device, and then four nucleotides with different current blocking tags are added to start sequencing data recording. The signal data is then processed to obtain the nucleic acid sequence information of the RNA to be tested.

[0136] Reagent kits or devices for nucleic acid replication, amplification or sequencing

[0137] Using the R2Bm001 reverse transcriptase disclosed herein, this disclosure provides kits and devices for RNA sequencing.

[0138] The test salt buffer used was Buffer-Seq (650 mM KAc, 20 mM Tris, pH 7.5). Four nucleotides modified with different positively charged current-blocking tags, each with a final concentration of 5 μM and 0.8 mM MnCl2, were dissolved in Buffer-Seq as sequencing reagents. The structures and synthetic routes of the four positively charged current-blocking tags and their corresponding modified nucleotides (referred to as "dN6P") are as follows:

[0139] Synthetic route of the current-blocking tag PA-15 (Figure 1)

[0140] LC-MS detection spectrum of current blocking tag PA-15 (Figure 2)

[0141] Synthetic route of the corresponding PA-15 modified nucleotide dG5P-PA-15 (Figure 3)

[0142] Synthetic route of the current-blocking tag PA-16 (Figure 4)

[0143] LC-MS detection spectrum of current blocking tag PA-16 (Figure 5)

[0144] Synthetic route of the corresponding PA-16 modified nucleotide dC6P-PA-16 (Figure 6)

[0145] Synthetic route of the current-blocking tag PA-25 (Figure 7)

[0146] LC-MS detection spectrum of current blocking tag PA-25 (Figure 8)

[0147] Synthetic route of the corresponding modified nucleotide dT5P-PA-25 of PA-25 (Figure 9):

[0148] Synthetic route of the current-blocking tag PA-11 (Figure 10)

[0149] LC-MS detection spectrum of current blocking tag PA-11 (Figure 11)

[0150] Synthetic route of the corresponding modified nucleotide dA5P-PA-11 of PA-11 (Figure 12)

[0151] The testing equipment was the G-seq-500 sequencer developed internally by our unit. A schematic diagram of its basic testing unit is shown in Figure 23.

[0152] This study confirms that reverse transcriptase and nanopore coupling can achieve real-time RNA sequencing at the single-molecule level; it discloses a complete biochemical protocol for good RNA sequencing, including RNA sample processing, selection and coupling schemes of reverse transcriptase and nanopores, and corresponding nucleotide reagents (reverse transcriptases can also utilize terminally modified nucleotides); and the selection, identification, and modification of reverse transcriptases suitable for single-molecule sequencing. For example:

[0153] 1) It was confirmed that reverse transcriptase and nanopore coupling can achieve real-time RNA sequencing at the single-molecule level in the "synthesis-while-sequencing" route.

[0154] 2) It was confirmed that reverse transcriptase R2Bm is suitable for direct RNA nanopore sequencing based on "synthesis-as-sequencing";

[0155] 3) A simple and efficient RNA sample processing and sequencing workflow was disclosed;

[0156] 4) The coupling method between R2Bm and NetB nanopores can bring good signal quality;

[0157] 5) Introducing the following single-point and arbitrary combination mutations into R2Bm: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, and L513F can increase its speed or signal dwell time (DW).

[0158] More specifically, in Example 1 below, the present disclosure obtained protein products of various R2Bm mutants by recombinant expression in Escherichia coli and purified them by affinity chromatography.

[0159] In Example 2, purified R2Bm02 was subjected to an extension experiment using RNA as a template, demonstrating that R2Bm possesses reverse transcriptase extension activity. This example also demonstrates that, in addition to extension activity, this reverse transcriptase also exhibits strand substitution functionality. During the extension of the product strand at the 3' end, it unwinds the original template-complementary double-stranded structure and replaces the original complementary strand with the newly generated product strand, thus forming a new "template-product strand" double-stranded structure. R2Bm exhibits good reverse transcription activity and strand substitution function at 30°C, indicating its compatibility with nanopore detection platforms. As membrane proteins, nanopores need to be embedded in a phospholipid bilayer to function properly, and such membranes typically require near room temperature to be stable. Therefore, this example demonstrates the good temperature compatibility of R2Bm with nanopore sequencing systems. Furthermore, Example 2 demonstrates that R2Bm and its mutants can be efficiently extended using four nucleotides with phosphate-terminated modifications, and that, relative to R2Bm02, they include any of the following mutations or any combination thereof, thereby improving the extension rate using the modified nucleotides: A531D, A533G, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, L513F, P456H, P459H, T349K, V434T, V453K, V453P, and Y350R. In this disclosure, the current-blocking groups modified on the nucleotides enter the nanopores under an electric field, generating base-type-specific blocking currents, which are the basis for the operation of the entire sequencing system. Therefore, the ability of reverse transcriptase R2Bm to modify bases is crucial to this invention.

[0160] In Example 3, this disclosure clarifies that the reverse transcriptase corresponding to R2Bm02 retains certain activity after being placed at 30°C for a period of time. This provides a guarantee for the long-term operation of the RNA sequencing system described in this disclosure, i.e., long-read sequencing.

[0161] In Example 4, this disclosure investigated the catalytic metal ion preference of R2Bm02. This reverse transcriptase in Mn 2 + or Mg 2+ It exhibits enzymatic activity in all its presence, and compared to Mg utilized under natural conditions... 2+ R2Bm in Mn 2+ Under these conditions, the enzyme activity is better than the former. Typically, the polymerase is coordinated with Mg... 2+ At this stage, it exhibits higher selectivity for nucleotide base types, higher fidelity, and fewer interference signals during sequencing, which is beneficial for achieving higher sequencing accuracy. However, in the case of coordinated Mn... 2+At this time, it has a faster catalytic speed and better progress. Considering the sequencing efficiency, this scheme mainly uses Mn 2+ As a catalytic metal ion, but also for Mg 2+ Catalytically advanced, more accurate direct RNA sequencing platforms offer the possibility.

[0162] Example 5 evaluated the salt tolerance of reverse transcriptase R2Bm02. In nanopore-based sequencing protocols, specific changes in current are the signal basis for sequencing. Therefore, a certain electrolyte concentration is required in the entire sequencing system to generate sufficient current intensity to support the signal-to-noise ratio for base identification. Generally, the electrolyte concentration supporting sequencing is higher than the ideal electrolyte concentration for polymerase activity. Example 5 demonstrates that R2Bm02 still retains a certain level of reverse transcriptase activity at high salt concentrations, and the degree of activity reduction is within an acceptable range.

[0163] In Example 6, this disclosure demonstrates the preparation process of NetB coupling monomers with different coupling tags and corresponding coupling nanopores.

[0164] Example 7 illustrates the RNA sequencing sample preparation process of the present invention. The sample preparation process of the present invention does not require adapter ligation and reverse transcription; the "template-primer-reverse transcriptase-nanopore" complex is purified by magnetic beads and can then be tested, making it simple and easy to implement.

[0165] Example 8 demonstrates the measured accuracy and corresponding DW signal of the R2Bm mutant reading linear RNA template T2. The results show that any single point or any combination of mutations on R2Bm—C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, L513F, N659K, R1027E, R961A, V434T, V453K, V453P, and Y350R—affects sequencing accuracy by influencing the DW signal.

[0166] Example 9 demonstrates the measured accuracy and read length of the "reverse transcriptase-nanopore" complex composed of R2Bm mutants with different N-terminal truncations and different coupling tags, coupled with corresponding nanopores, for reading circular RNA template T3. The results show that: 1. Both full-length and truncated R2Bm mutants can generate sequencing signals, with little difference in accuracy across multiple truncation sites; 2. Significant differences in sequencing signal quality generated by different coupling methods directly affect sequencing accuracy; 3. The coupling position of the R2Bm mutant on the NetB nanopore is also a key factor in improving sequencing accuracy. In Examples 8-9, the reverse transcriptase R2Bm identified in this disclosure has clearly demonstrated its application prospects for nanopore-based real-time direct RNA sequencing.

[0167] In Example 10, this disclosure demonstrates an experimental procedure and measured data for direct RNA sequencing using nucleotides modified with negatively charged current-blocking groups (nucleotides with positively charged current-blocking groups mentioned above) coupled to an αHL nanopore derived from the intron reverse transcriptase PsRT of Planococcus salinus type II (although positively charged current-blocking groups can also be used). This example demonstrates that the direct RNA sequencing system of this disclosure is not limited to specific reverse transcriptases, specific nanopores, or specific modified nucleotides.

[0168] In summary, a series of embodiments demonstrate that the R2Bm reverse transcriptase mutant identified in this invention exhibits excellent overall performance across various sequencing-related aspects. After coupling with nanopores, high-accuracy real-time direct RNA sequencing can be achieved, showing promising application prospects in the industrial application of RNA sequencing. Furthermore, this disclosure presents a complete direct RNA sequencing scheme based on reverse transcription and nanopore detection. This scheme is not limited to specific reverse transcriptases, coupling methods, nanopores, or modified nucleotides, and represents a universal direct RNA sequencing route.

[0169] The present invention will be further described in detail below through specific embodiments. These embodiments are descriptive examples and are not intended to be limiting. Based on the content and embodiments of the present invention, those skilled in the art may make reasonable and natural extensions to the specific implementations, which should not be considered as exceeding the scope of the present invention.

[0170] Example

[0171] Example 1. Preparation of R2Bm reverse transcriptase mutant

[0172] Mutant protein expression: Using pET22b as the expression vector, R2Bm protein mutants with different fusion tags were constructed. The vector was transformed into Escherichia coli BL21 strain, and single colonies were picked and inoculated into 4 mL of LB medium containing 100 μg / mL ampicillin. The culture was incubated at 37°C and 230 rpm for 16 h. Then, the culture was transferred to 200 mL of TB medium containing 100 μg / mL ampicillin and incubated at 37°C and 230 rpm for another 5 h. Finally, IPTG was added to a final concentration of 100 μM for induction, and the culture was incubated at 16°C and 200 rpm for 16 h.

[0173] Protein purification: The culture was transferred to a centrifuge flask and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells. The cell pellet was then resuspended in 60 mL Binding buffer 1 (50 mM Tris-HCl pH 7.5, 1 M NaCl, 10% glycerol, 0.1% TWEEN 20, 2 mM TCEP). A protease inhibitor (5 μg / mL leupeptin, 0.2 mM PMSF) was added. The cell resuspended cells were sonicated in an ice-water bath and then centrifuged at 15000 rpm, 4°C for 15 min to pellet the remaining cell debris. The supernatant was then transferred to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After loading, the column was washed with 50 mL of Washing buffer 1 (Binding buffer 1 with 50 mM imidazole added), and finally eluted with 3-10 mL of Elution buffer-1 (300 mM imidazole (pre-adjusted to pH 8.0), 50 mM Tris-HCl (pH 8.0), 200 mM NaCl, 10% glycerol, 25% trehalose, 2 mM TCEP) (the resulting product was used in Examples 2-5 below). The NI column-purified protein was further purified by molecular sieve chromatography (Superdex 200 Increase 10 / 300GL) (the resulting product was used in Examples 7-9 below). The final target protein was stored in buffer (50 mM Tris 7.5, 200 mM KCl, 5 mM DTT, 0.1% Tween-20, 20% Glycerol) at -80°C. The purity of the samples was verified by 12% SDS-PEGA gel electrophoresis.

[0174] Example 2. Elongation rate of reverse transcriptase R2Bm01 / 02 using phosphate-terminated nucleotides

[0175] Table 1. Reagent composition for the elongation rate test

[0176] Template preparation: RNA fluorescent template T1 (5'-cy3) (SEQ ID NO:54) at a final concentration of 500 nM was mixed with the corresponding DNA extension primer P1 (SEQ ID NO:55) and the fluorescence quenching primer Q1 (SEQ ID NO:56) in a ratio of 1:1.2:1.2. The mixture was then annealed in Buffer-Anneal (20 mM Tris, pH 8.0, 60 mM KCl) under the following conditions: 80 °C for 5 minutes, 52 °C for 1 minute, and 42 °C for 1 minute. The resulting complex was named RNA7T, and its schematic diagram is shown in Figure 14.

[0177] Nucleotide extension rate test: Referring to the formula in Table 1, RNA7T was mixed with R2Bm02 or R2Bm mutant at a 1:1 volume ratio and incubated at 25°C for 10 min. Tris, KAc, dN6P or dNTPs were added gradually, and Mn was added using a microplate reader (Synergy H1, Bio-Tek) via a sample pump. 2+ The detection was performed under excitation at 540 nm and emission at 570 nm. The elongation rate refers to the number of nucleotides incorporated into the newly synthesized strand per unit time; in this embodiment, it is quantified by the increase in fluorescence value (570 nm) per unit time. Before elongation, the fluorescent gene is quenched and cannot emit fluorescence because the fluorescent group (-Cy3) on the template is very close to the quencher group (-BHQ2) located on the pointer strand. With the elongation of the newly synthesized strand and the strand substitution effect of reverse transcriptase on the quencher primer, the fluorescent gene separates from the quencher group and emits fluorescence; therefore, the increase in fluorescence per unit time reflects the elongation rate of the newly synthesized strand.

[0178] Test Results: As shown in Figure 15A, R2Bm reverse transcriptase can bind to and elongate RNA templates in the presence of dNTPs or the modified nucleotide dN6P. After structural modification, R2Bm mutants, including any one or any combination of the following mutations, improved the elongation rate using the modified nucleotides: A531D, A533G, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, L513F, P456H, P459H, T349K, V434T, V453K, V453P, Y350R (see Figure 15B and Table 2). When the reaction system lacks nucleotides or reverse transcriptase, there is no fluorescent growth signal, indicating that the reverse transcriptase cannot extend.

[0179] Table 2. Speed ​​display of R2Bm mutant (time to complete extension)

[0180] Example 3. Thermostability assessment of reverse transcriptase R2Bm01 / 02

[0181] Table 3. Reagent composition for thermal stability assessment experiment

[0182] Referring to the formulation in Table 3, RNA7T and R2Bm were first mixed. The enzyme-template complex was heated using a PCR instrument and incubated at 25°C for 10 min, followed by heating at 30°C for 10, 30, 60, and 120 min, respectively. Then, the mixture was added to a black 96-well plate pre-contained with Tris, KAc, dN6P, and DTT. Mn was then added using a microplate reader via a sample pump. 2+ The thermal stability test was conducted at a temperature of 30℃, under conditions of 540nm excitation light and 570nm emission light.

[0183] The test results are shown in Figure 16. The activity of reverse transcriptase R2Bm02 decreased to some extent after incubation at 30℃ for different times, but remained within an acceptable range. This indicates that R2Bm02 reverse transcriptase is suitable for long-duration, long-read sequencing at 30℃.

[0184] Example 4. Selectivity of reverse transcriptase R2Bm01 / O2 for catalyzing metal ions

[0185] Table 4. Reagent composition for the metal ion selectivity assessment experiment

[0186] Referring to the formulation in Table 4, the reaction was carried out in a 96-well microplate. First, the template and enzyme were mixed and incubated at 25°C for 10 min. Then, a 96-well plate pre-contained with Tris, KAc, dN6P, DTT, etc., was added, and Mn was added to different final concentrations using a microplate reader via a sample pump. 2+ or Mg 2+ The detection was performed at 30℃, with excitation light at 540nm and emission light at 570nm.

[0187] The test results (Figure 17) show that the quaternary complex formed by enzyme-template-metal ion-modified nucleotide is a necessary condition for elongation, and the mutant is more sensitive to metal ions (Mn) 2+ / Mg 2+ The preference for Mn is reflected in the addition of Mn to the reaction solution. 2+ Or Mg 2+ At that time, whether the extension reaction can proceed normally, and the rate of extension. R2BmO2 in Mn 2+ Under certain conditions, normal elongation of modified bases is possible, with Mn at 0.4-1.2 mM. 2+ Within the concentration range, the elongation rate varies with Mn 2+ The increase in concentration slowed slightly. However, in Mg... 2+Under these conditions, the elongation rate is extremely low, only reaching 30 mM Mg. 2+ Significant extended fluorescence signals were only produced 3 minutes later under the given conditions, much slower than those produced by Mn. 2+ 10 seconds under the given conditions. Explanation of Mn 2+ More suitable for the RNA sequencing scheme disclosed herein, and also demonstrates the development of Mg-based RNA sequencing solutions. 2+ The possibility of higher accuracy sequencing systems.

[0188] Example 5: Salt tolerance assessment of reverse transcriptase R2Bm01 / 02

[0189] Table 5. Reagent composition for salt tolerance assessment experiment

[0190] Referring to the formulation in Table 5, first mix the template and enzyme, and incubate at 25°C for 10 min. Then, add the experimental group to PCR tubes pre-filled with interfering RNA, Tris, KAc, dN6P, and DTT, and incubate at 27°C for 30 min. The template-enzyme mixture for the control group is placed on ice. After incubation, transfer the components from the experimental group PCR tubes to 96-well microplates. The template-enzyme mixture for the control group is added in the same proportion to 96-well microplates pre-filled with interfering RNA, Tris, KAc, dN6P, and DTT. During testing, Mn is added via a sample pump using a microplate reader. 2+ The detection was performed at 30℃, with excitation light at 540nm and emission light at 570nm.

[0191] The test results are shown in Figure 18, which indicate that under the condition of interfering RNA, the enzyme activity of R2Bm02 decreased to basically the same extent when incubated at 27°C for 30 min under 350mM KAc or 650mM KAc. This shows that the decrease in enzyme activity of both enzymes is mainly due to insufficient thermal stability and is less affected by salt concentration. This indicates that R2Bm has good tolerance to high concentrations of KAc and is suitable for the sequencing scheme disclosed in this paper.

[0192] Example 6. Preparation of Coupled Nanopores

[0193] Preparation of NetB protein mutant common monomer (NetB01)

[0194] Mutant protein expression: pET26b was used as the expression vector. When constructing the NetB protein mutant expression vector, TEV and His-Tag tags were added to the C-terminus. The vector was transformed into Escherichia coli BL21 strain. Single clones were picked and inoculated into 10 mL of LB medium containing 50 μg / mL kanamycin and cultured at 37°C and 250 rpm for 16 h. Then the culture was transferred to 400 mL of TB self-induction medium containing 50 μg / mL kanamycin and cultured at 25°C and 250 rpm for another 16 h.

[0195] The culture was transferred to a centrifuge flask and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells. The cell pellet was then resuspended in 40 mL of Binding Buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol). The resuspended cell pellet was then sonicated in an ice-water bath, followed by centrifugation at 15000 rpm, 4°C for 15 min to pellet residual cell debris. The supernatant was added to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After addition, the column was washed with 10 mL of Washing Buffer 1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 30 mM imidazole). Finally, the target protein was eluted with 3–10 mL of Elution Buffer-1 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 300 mM imidazole).

[0196] After protein elution, the OD280 concentration was determined using a spectrophotometer. Then, TEV protease was added at a concentration of 100:1, and the mixture was incubated at 4°C for 16 h to remove the C-terminal His-Tag tag. The digested protein was dialyzed into 1 L of Binding Buffer-1 and dialyzed twice at room temperature, 1 h each time. Finally, the dialyzed sample was added again to a 1 mL pre-equilibrated Ni-agarose gel gravity column, and the flow-through sample was collected to obtain the NetB protein mutant. The sample was examined by 12% SDS-PEGA gel electrophoresis (Figure 22) and then used for further processing.

[0197] Preparation of NetB protein mutant conjugate monomers

[0198] The expression of the conjugated monomer is the same as that of the NetB protein mutant ordinary monomer.

[0199] After expression, the culture was transferred to a centrifuge flask and centrifuged at 4000 rpm, 4°C for 15 min to pellet the cells. The cell pellet was then resuspended in 40 mL of Binding Buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol). The resuspended cell pellet was then sonicated in an ice-water bath, followed by centrifugation at 15000 rpm, 4°C for 15 min to pellet residual cell debris. The supernatant was added to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After addition, the column was washed with 10 mL of Washing Buffer 2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 30 mM imidazole). Finally, the target protein was eluted with 3–10 mL of Elution Buffer-2 (50 mM PB (pH 8.0), 200 mM NaCl, 10% glycerol, 300 mM imidazole).

[0200] Coupling modification of NetB protein mutant coupling monomers

[0201] The mutant conjugate of NetB03-05 requires conjugation with N-Maleimide-SpyTag (as shown in Figure 19). Add 400 μM N-Maleimide-SpyTag to the eluted conjugate and incubate at room temperature for 3 hours. Finally, add 1 mM DTT to terminate the reaction. The mutant conjugate of NetB08 requires conjugation with Biotin-PEG3-Maleimide (as shown in Figure 20). Add 400 μM Biotin-PEG3-Maleimide to the eluted sample and incubate at room temperature for 3 hours. Finally, add 1 mM DTT to terminate the reaction. Samples were examined by 12% SDS-PEGA gel electrophoresis (Figure 21) and then used for further processing.

[0202] Preparation of NetB-conjugated nanoporous proteins

[0203] Based on the protein concentration detected by OD280, the ordinary monomer (NetB01) and the coupling monomer were mixed at a ratio of 6:1. DPHPC phospholipids were added to a final concentration of 0.75 mg / mL, and the mixture was allowed to stand at 37°C for 16 h. Then, it was heated at 55°C for 25 min, and β-OG (n-octyl-β-D-glucopyranoside) was added to a final concentration. After thorough dissolution of the phospholipids, Tween 20 was added to a final concentration. The mixture was then pipetted and centrifuged at 3500 rpm for 3 min to remove the precipitate. The supernatant was collected for later use. The supernatant was then added to a 1 mL pre-equilibrated Ni-agarose gel gravity column. After loading, the column was washed with 2 mL of Washing Buffer 3 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 0.2% Tween 20, 30 mM imidazole), and finally eluted with 2 mL of Elution Buffer-3 (50 mM Tris-HCl pH 8.0, 200 mM NaCl, 10% glycerol, 0.2% Tween 20, 300 mM imidazole). The eluted sample was loaded into a cation exchange column (RESOURSE Q 1 mL) and eluted with a gradient of 15-35% Buffer B (20 mM Tris-HCl pH 8.0, 2 M NaCl, 0.1% Tween-20). The fraction collected at the 40 s conductivity value was the desired 1:6 coupled nanopore. The sample was purified by 12% SDS-PAGE to detect heptamer purification (Figure 21) and then used for later use.

[0204] Example 7. Preparation of RPN complex

[0205] PN preparation: The R2Bm mutant was mixed with the corresponding NetB-conjugated nanopore at a ratio of 2:1, incubated on ice for 1 h, and then incubated at 16 °C for 2 h. The mixture was purified by heparin column chromatography using buffer A (50 mM Tris-HCl pH 7.5, 50 mM KCl, 10% glycerol, 0.02% TWEEN 20, 2 mM TCEP). The sample was first washed with 10% buffer B (50 mM Tris-HCl pH 7.5, 1.5 M KCl, 10% glycerol, 0.02% TWEEN 20, 2 mM TCEP), and then with 40% buffer B to remove the target protein. The purified sample was examined by 12% SDS-PEGA gel electrophoresis (Figure 22) and named PN.

[0206] Sample preparation: The T2 template (SEQ ID NO:50) with a final concentration of 500 nM was mixed with the corresponding DNA primer P2 (SEQ ID NO:51) at a ratio of 1:1.5 and annealed in Buffer-Anneal (20 mM Tris, pH 8.0, 60 mM KCl) at 85 °C for 3 minutes, then the temperature was decreased by 0.3 °C every 20 seconds until it reached 25 °C. DTT was then added to bring the final concentration to 5 mM.

[0207] RPN preparation: Take 2.5 μL of annealed 500 nM RNA template and mix it with PN at a ratio of 1:1.5. Add DTT to a final concentration of 5 mM, add 1 μL of recombinant RNase inhibitor (Takara, 2313A), and then add RPN buffer (50 mM Tris, pH 7.5, 50 mM KCl, 20% glycerol, 2 mM TCEP, 1 mM EDTA, 0.01% Tween-20) to a final concentration of 20 μL. Couple at 25 °C for 120 min.

[0208] RPN purification: Take 1.8 times the original volume of RNA purification magnetic beads (restored to room temperature) into a 1.5 mL centrifuge tube, place it on a magnetic rack for 5 min, carefully aspirate and discard the supernatant, and resuspend the magnetic beads in the same volume of binding buffer (20% PEG4000, 640 mM Tris, pH 7.8, 600 mM KAc, 0.02 mg / ml BSA, 60 mM KCl, 0.2% Tween-20). Add 20 μL of the incubated sequencing complex to the resuspended magnetic beads, gently pipette to mix 10–20 times, and incubate at room temperature for 10 min. Place the centrifuge tube on a magnetic rack, and after the solution becomes clear, discard the supernatant. Add 200 μL of wash buffer (binding buffer diluted 1:1), slowly rotate the centrifuge tube horizontally 360° on a magnetic rack, then discard the supernatant. Repeat the wash once, completely aspirating and discarding the supernatant. Remove the centrifuge tube from the magnetic rack and add 20 μL of elution buffer (3 μM dN6P, 28 mM Tris, pH 7.8, 0.3 mM CaCl2, 0.02 mg / ml BSA, 30% glycerol, 0.015% Tween-20). Gently pipette to mix 10–20 times and incubate at room temperature for 5 min. After centrifugation, place the tube on the magnetic rack and allow it to clarify. Then, aspirate the supernatant and store it in a new PCR tube. The tube can then be stored at -80°C or directly used for sequencing. Samples should be protected from freeze-thaw cycles to control the impact of these processes on sequencing quality.

[0209] Example 8. Linear RNA Sequencing

[0210] As shown in Figure 23, the apparatus was ready. Various RPN complex samples prepared using T2 (specific sequence SEQ ID NO:50) as a template and P2 (specific sequence SEQ ID NO:51) as a primer were diluted to appropriate concentrations using Buffer-Seq (650mM KAc, 5% ethylene glycol, 20mM Tris, pH 7.5) and added to the apparatus. After the RPN complexes were embedded into the artificial phospholipid membrane, excess RPN complexes were washed away using Buffer-Seq. Finally, the prepared sequencing reagents were added to begin the test and recording. The test used a 500Hz AC current with an excitation voltage of (100mV, -100mV), and the detector measured the current value passing through the nanopore every 200μs. Figure 24 (R2Bm08) shows the sequencing signal, where different current blocking depths correspond to different bases. The blocking currents from shallow to deep (i.e., from top to bottom) correspond to the four bases C, G, T, and A, respectively. Figure 24 shows that the sequencing signal is clear and easy to identify, and the capture rate (i.e. the probability of current blocking in each cycle during the signal duration) is good. The figure also shows the analysis and sequence alignment results of the corresponding current signal.

[0211] Table 6. Linear RNA sequencing accuracy and DW for different R2Bm mutants

[0212] Table 6 shows the sequencing accuracy and DW (average of 4 bases) for the R2Bm mutant (NetB01+NetB02). The results show that any single point or any combination of bases on R2Bm can be found: C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E346K, E454K, E454N, E461H, E461Y, F. Mutations in 530Y, H342K, H962A, K1026A, K1029A, L270W, L513F, N659K, R1027E, R961A, V434T, V453K, V453P, and Y350R can improve sequencing accuracy by affecting the depth of wave (DW) of the sequencing signal (Table 6 and Figure 25 compare the sequencing signals of R2Bm08 and R2Bm09 with their DW).

[0213] Example 9. Circular RNA Sequencing

[0214] The sequencing apparatus and procedure were the same as in Example 8. The naturally occurring circular RNA circPVT1(homo) (T3, specific sequence SEQ ID NO:52) was used as a template, and P3 (specific sequence SEQ ID NO:53) was used as a primer. The combination of R2Bm reverse transcriptase and NetB-coupled nanopores was used (R2Bm and NetB mutant information is shown in Table 7, where the coupled nanopore protein is the NetB01+NetB mutant). Sequencing was performed for 30 minutes.

[0215] Table 7. Information on R2Bm and NetB mutants

[0216] The sequencing system disclosed herein reads circular RNA and generates clear sequencing signals. The sequencing results of R2Bm08, compared with the RNA sequence (Figures 26A-C), show a median accuracy of nearly 94%, with more than 10 rolling circles. Further consistency sequence correction further improves the accuracy to 99.9% (Figure 27). The overall sequencing results are shown in Table 8: both full-length and truncated R2Bm mutants can generate sequencing signals, with slight differences in accuracy at multiple truncation positions, but the overall difference is not significant. The sequencing accuracy of dual-coupling is significantly higher than that of single-coupling. Differences in the quality of sequencing signals generated by different coupling positions and coupling pairs on R2Bm and NetB directly affect sequencing accuracy. Meanwhile, the accuracy of multiple coupling methods is high. In summary, the sequencing system disclosed herein is not limited by the coupling method; different coupling methods can generate clear sequencing signals and high sequencing accuracy.

[0217] Table 8. Sequencing accuracy and median read length of circular RNA generated by different "R2Bm-NetB coupled nanopore" complexes

[0218] This embodiment demonstrates that the sequencing system of this disclosure can perform rolling reads of circular RNA and improve the accuracy of sequencing results through sequence consistency correction. The sequencing method of this disclosure is currently the only means capable of directly sequencing circular RNA.

[0219] Example 10. Direct RNA sequencing based on PsRT reverse transcriptase

[0220] This embodiment uses a reverse transcriptase derived from the type II intron of *Planococcus salinus* (hereinafter referred to as PsRT) (specific sequence: SEQ ID NO:48). Mutations of D15R, L63R, K233A, A234D, and E257R are added to PsRT, and HisTag and avidin are fused to its N-terminus, while SpyCatcher is fused to its C-terminus, resulting in the designation PsRT01 (specific sequence: SEQ ID NO:49). The purification process for PsRT01 is the same as in Example 1.

[0221] This embodiment uses αHL-coupled nanopores that have undergone mutation modification. The preparation process is the same as in Example 6. It is assembled from ordinary monomer αHL01 (specific sequence is SEQ ID NO:46) and αHL02 (specific sequence is SEQ ID NO:47) coupled nanopores with C-terminal SpyTag and modified with Biotin at N17C site.

[0222] The synthesis steps and mass spectrometry identification results of the four modified nucleotides used in this embodiment are shown in Figures 28-35.

[0223] RPN preparation: 2.5 μL of annealed 500 nM T1 template was mixed with 2.5 μL of 400 nM PsRT01, DTT was added to a final concentration of 5 mM, and the mixture was incubated at 25 °C for 20 min. Then, 8 μL of 10 nM αHL was added to couple the nanopores, and the mixture was coupled at 25 °C for 40 min.

[0224] The sequencing apparatus and process were the same as in Example 8, using the modified nucleotides shown in Figures 28, 30, 32, and 34. The test was performed using a 500Hz AC current with an excitation voltage of (120mV, -180mV), and the detector measured the current through the nanopore every 200μs.

[0225] The sequencing signal is shown in Figure 36. Different current blocking depths correspond to different bases, and the sequencing signal with the sequence TGGGCTGAC was obtained, in which "CTGAC" completely matches the template sequence.

[0226] This embodiment demonstrates that the RNA direct sequencing scheme disclosed herein is not limited to specific reverse transcriptases, specific nanopores, or specific modified nucleotides.

[0227] Sequence information

Claims

1. RNA sequencing methods, including sequencing-by-synthesis using reverse transcriptase or its mutants.

2. The method of claim 1, wherein the RNA sequencing is nanopore sequencing.

3. The method of claim 1 or 2, wherein the RNA is linear or circular RNA.

4. The method according to any one of claims 1-3, wherein the reverse transcriptase is R2Bm or PsRT.

5. The method of any one of claims 1-4, wherein the reverse transcriptase or its mutant is coupled to the nanoporin through one, two or more sites.

6. The method according to any one of claims 1-5, wherein the coupling site is located at the C-terminus, N-terminus, and / or middle of the amino acid sequence of the reverse transcriptase or its mutant.

7. The method according to any one of claims 1-6, wherein the nanoporin is NetB, αHL, or a mutant thereof.

8. The method according to any one of claims 1-7, wherein the reverse transcriptase R2Bm or its mutant has 1-288 amino acid residues deleted at its N-terminus.

9. The method according to any one of claims 1-8, wherein the reverse transcriptase R2Bm or its mutant comprises any one or any combination of the following mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F.

10. The method according to any one of claims 1-9, wherein the reverse transcriptase or its mutant and the nanoporin are coupled by covalent or non-covalent bonds.

11. The method of any one of claims 1-10, wherein the covalent bond comprises a SpyCatcher-SpyTag bond.

12. The method according to any one of claims 1-11, wherein the non-covalent bond comprises avidin-biotin binding, Cl7-Im7 binding, or BimBinder-BimTag binding.

13. The method of any one of claims 1-12, wherein the reverse transcriptase R2Bm or its mutant is fused with SpyCatcher / SpyTag / BimBinder / Cl7 at the N or C end, and at least one subunit of the nanoporin NetB or its mutant comprises SpyTag / SpyCatcher / BimTag / Im7.

14. The method according to any one of claims 1-13, wherein the SpyCatcher / SpyTag / BimBinder / Cl7 is fused to the N-terminus of the nanoporin NetB, or the SpyTag is linked to the nanoporin NetB or its mutant via a cysteine ​​residue contained in the nanoporin NetB or its mutant.

15. The method of any one of claims 1-14, wherein the nanoporin NetB mutant comprises an I27C, D96C, or N236C mutation.

16. The method of any one of claims 1-15, wherein the reverse transcriptase R2Bm or its mutant is fused with SpyCatcher / SpyTag / BimBinder / Cl7 / antibiotin, and at least one subunit of the nanoporin NetB or its mutant comprises SpyTag / SpyCatcher / BimTag / Im7.

17. The method according to any one of claims 1-16, wherein the Cl7 or avidin replaces, or inserts into, any of the positions 297-303, 670-676, or 376-383 of the reverse transcriptase R2Bm or its mutant.

18. The method of any one of claims 1-17, wherein the Im7 is fused to the N or C terminus of the nanoporin NetB or a mutant thereof.

19. The method of any one of claims 1-18, wherein the reverse transcriptase R2Bm or a mutant thereof is further fused with a solubilizing tag, such as TrxA or MBP.

20. The method according to any one of claims 1-19, wherein the reverse transcriptase is PsRT or a mutant thereof, and the nanoporin is αHL or a mutant thereof.

21. The method of any one of claims 1-20, wherein the PsRT mutant comprises mutants D15R, L63R, K233A, A234D and / or E257R.

22. The method of any one of claims 1-21, wherein the PsRT mutant is further fused with an avidin at the N-terminus and with a SpyCatcher at the C-terminus.

23. The method of any one of claims 1-22, wherein the nanoporin αHL is fused to its C-terminus with a SpyTag and includes an N17C mutation.

24. The method of any one of claims 1-23, wherein the nanoporin αHL is linked to biotin at the N17C mutation site.

25. The method of any one of claims 1-24, wherein the reverse transcriptase and nanoporin include a purification tag, such as HisTag, at the N-terminus or C-terminus.

26. The method of any one of claims 1-25, wherein the amino acid sequence of the reverse transcriptase R2Bm is as shown in SEQ ID NO:1, or comprises an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:

1.

27. The method of any one of claims 1-26, wherein the amino acid sequence of the reverse transcriptase R2Bm mutant is as shown in any one of SEQ ID NO:2-34, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:2-34.

28. The method of any one of claims 1-27, wherein the amino acid sequence of at least one subunit of the nanoporin NetB mutant is as shown in any one of SEQ ID NO:36-45, or includes an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:36-45.

29. The method of any one of claims 1-28, wherein the amino acid sequence of the PsRT mutant is as shown in SEQ ID NO:48, including an amino acid sequence that has at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:

48.

30. The method according to any one of claims 1-29, wherein the amino acid sequence of the nanoporin αHL is as shown in SEQ ID NO:46, including an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:

46.

31. The method according to any one of claims 1-30, wherein it is performed at room temperature.

32. The method according to any one of claims 1-31, wherein in Mn 2+ or Mg 2+ It is performed under the condition that it exists.

33. The method according to any one of claims 1-32, wherein Mn 2+ Concentrations of 0.2-1.2 mM, Mg 2+ The concentration is above 30mM.

34. The method according to any one of claims 1-33, wherein it is carried out in the presence of 350-650 mM KAc.

35. The method of any one of claims 1-34, wherein the reverse transcriptase is capable of synthesizing DNA strands using modified nucleotides as substrates.

36. The method according to any one of claims 1-35, wherein the reverse transcriptase is R2Bm or a mutant thereof, and the structure of the modified nucleotide is as follows:

37. The method according to any one of claims 1-36, wherein the reverse transcriptase is PsRT or a mutant thereof, and the structure of the modified nucleotide is as follows:

38. A reverse transcriptase, which is a modified R2Bm, said modification comprising one or more of the following: 1) Introduce any one or any combination of the following mutations: A531D, A533G, C337P, C660P, D334N, D360H, D490N, D529N, D707N, E345K, E345Y, E346K, E454K, E454N, E461H, E461Y, F530Y, H342K, H962A, K1026A, K1029A, L270W, N659K, P456H, P459H, R1027E, R961A, T349K, V434T, V453K, V453P, Y350R, L513F; 2) The N-terminus of 1-288 amino acid residues is missing; 3) Introduce one or more coupling moieties that can bind to the coupling moieties on the nanoporous protein; 4) Introduce solubilizing tags, such as TrxA or MBP; 5) Introduce purification tags, such as HisTag.

39. The reverse transcriptase of claim 38, wherein the coupling portion is capable of forming a covalent or non-covalent binding with the nanoporin.

40. The reverse transcriptase of claim 38 or 39, wherein the coupling portion is located at the C-terminus, N-terminus, and / or middle of the amino acid sequence of the reverse transcriptase or its mutant.

41. The reverse transcriptase according to any one of claims 38-40, wherein the coupling portion is SpyCatcher or SpyTag, avidin or biotin, Cl7 or Im7 or BimBinder-BimTag.

42. The reverse transcriptase according to any one of claims 38-41, wherein the SpyCatcher, SpyTag, BimBinder, or Cl7 is located at the N-terminus.

43. The reverse transcriptase according to any one of claims 38-42, wherein the Cl7 is located at the C or N terminus.

44. The reverse transcriptase according to any one of claims 38-43, wherein the SpyCatcher / SpyTag / BimBinder replaces its positions 297-303, 670-676, or 376-383, or is inserted into any position among positions 297-303, 670-676, or 376-383.

45. The reverse transcriptase according to any one of claims 38-44, wherein the amino acid sequence of R2Bm is as shown in SEQ ID NO:1, or includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:

1.

46. ​​The reverse transcriptase according to any one of claims 38-45, wherein the amino acid sequence is as shown in any one of SEQ ID NO:2-34, or includes an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in any one of SEQ ID NO:2-34.

47. A reverse transcriptase, which is a modified PsRT, said modification comprising one or more of the following: 1) Introduce mutations D15R, L63R, K233A, A234D and / or E257R; 2) Introduce one or more coupling moieties capable of binding to coupling moieties on nanoporous proteins; and 3) Introduce purification tags, such as HisTag.

48. The reverse transcriptase of claim 47, wherein the coupling portion comprises avidin and / or SpyCatcher.

49. The reverse transcriptase of claim 47 or 48, wherein the avidin is fused to the N-terminus of the reverse transcriptase and the SpyCatcher is fused to the C-terminus of the reverse transcriptase.

50. The reverse transcriptase according to any one of claims 47-49, wherein the amino acid sequence is as shown in SEQ ID NO:48, or comprises an amino acid sequence having at least 90% sequence identity with the amino acid sequence shown in SEQ ID NO:

48.

51. A sequencing conjugate comprising the reverse transcriptase and nanoporous protein conjugated thereto as described in any one of claims 38-50.

52. The sequencing conjugate of claim 51, wherein the reverse transcriptase and the nanoporin bind through two or more conjugation sites.

53. The sequencing conjugate of claim 51 or 52, wherein the nanoporin is nanoporin NetB, αHL, or a mutant thereof.

54. An RNA sequencing kit comprising the reverse transcriptase as described in any one of claims 38-50 or the sequencing conjugate as described in any one of claims 51-53.

55. The sequencing kit of claim 54, further comprising modified nucleotides.

56. The sequencing kit of claim 55, wherein the modified nucleotide is selected from one or more of the following compounds: