Variant reverse transcriptase
A truncated variant reverse transcriptase with a sequence-specific DNA binding domain enhances RNA template copying efficiency, particularly on circular templates, addressing inefficiencies in existing reverse transcriptases by achieving higher cDNA yields and improved template switching.
Patent Information
- Application Number
- PCT/US2025/011846
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-24
AI Technical Summary
Existing reverse transcriptases are inefficient in copying certain RNA templates, particularly at lower temperatures, and lack optimal performance on circular templates.
A variant reverse transcriptase with a truncated N-terminus and high sequence identity to residues 83-841 of SEQ ID NO: 1, which includes a fusion protein with a sequence-specific DNA binding domain, exhibits enhanced efficiency in copying RNA templates, especially at lower temperatures and on circular templates.
The variant reverse transcriptase achieves higher cDNA yields and improved template switching efficiency, with reduced non-templated addition, making it more effective for RNA template copying and second strand synthesis.
Smart Images

Figure US2025011846_24072025_PF_FP_ABST
Abstract
Description
[0001] VARIANT REVERSE TRANSCRIPTASE
[0002] CROSS-REFERENCING
[0003] This application claims the benefit of provisional application serial no 63 / 622,714, filed on January 19, 2024, which application is incorporated herein in its entirety.
[0004] SEQUENCE LISTING
[0005] This application is filed with a Sequence Listing in electronic form as a Sequence Listing XML, "NEB-479" created on January 15, 2025, and having a size of 2,691 bytes. The contents of the Sequence Listing XML are incorporated by reference herein in its entirety.
[0006] BACKGROUND
[0007] Reverse transcriptases are multi-functional enzymes that typically have multiple enzymatic activities, including an RNA-dependent DNA polymerization activity, a DNA- dependent DNA polymerization activity, and an RNaseH activity that catalyzes the cleavage of RNA in RNA-DNA hybrids. These enzymes, which are used to synthesize complementary DNA (cDNA) using RNA as a template, were first identified in RNA viruses. Subsequently, reverse transcriptases have been isolated and purified directly from virus particles, cells, and tissues (e.g., see Kacian et aL, 1971, Biochim. Biophys. Acta 46: 365-83; Yang et aL, 1972, Biochem. Biophys. Res. Comm. 47: 505-11; Gerard et aL, 1975, J. ViroL 15: 785-97; Liu et aL, 1977, Arch. Virol. 55 187-200; Kato et aL, 1984, J. ViroL Methods 9: 325-39; Luke et aL, 1990, Biochem. 29: 1764-69 and Le Grice et al., 1991, J. ViroL 65: 7004-07). More recently, mutants and fusion proteins have been created in the quest for improved properties such as thermostability, fidelity and activity.
[0008] SUMMARY
[0009] Provided herein is a reverse transcriptase comprising an amino acid sequence that is at least 90% identical to amino acid residues 83-841 of SEQ ID NO: 1 and has a truncated N- terminus relative to SEQ ID NO: 1. In some embodiments, the reverse transcriptase has an improvement in one or more properties. For example, the present reverse transcriptase may be more efficient at copying certain RNA templates (e.g., longer and / or circular RNA templates) than other reverse transcriptases, particularly at lower temperatures. Kits, reaction mixes and methods that include the reverse transcriptase are also provided.
[0010] These and other features of the present teachings are set forth herein.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The skilled artisan will understand that the drawings, described below, are for illustration purposes only. The drawings are not intended to limit the scope of the present teachings in any way.
[0013] Figure 1A-1B: Temperature dependence of DNA synthesis yield by Induro® (NEB), FBu, BMo, and ProtoScript® II Reverse Transcriptases (NEB). Synthesis was performed using a 1 kb circular RNA template (Figure 1A) or a 4 kb linear RNA template (Figure IB). cDNA generated by reverse transcriptases were all subjected to second strand synthesis using ProtoScript II and analyzed on a TapeStation® (Agilent) using Genomic DNA Tape and ladder.
[0014] Figure 2: Temperature dependence of rolling-circle reverse transcription. Induro, BMo and FBu were used to reverse transcribe a 1.2 kb circular RNA. Products were analyzed on a 1.2% agarose, lx TBE gel. Lanes marked M indicate 1 kb plus DNA ladder (NEB).
[0015] Figure 3: Temperature dependence of rolling-circle reverse transcription. Induro and FBu were used in rolling-circle reverse transcription and analyzed on a 1.2% agarose, lx TBE gel. Lanes marked M indicate 1 kb plus DNA ladder. Lanes marked L indicate reverse transcription product using linear template of same length and composition as circular template.
[0016] Figure 4: Mg21dependence of non-templated addition (NTA) for FBu. NTA profiles for concentrations of Mg2t1 mM and 5 mM determined by capillary electrophoresis (CE). The number of NTA is labeled on x-axis.
[0017] Figure 5: Non-templated addition (NTA) profiles of BMo and FBu for the first and second NTA as determined by LC-MS. The fraction of each addition is represented by the size of the bar.
[0018] Figure 6A-6B: Capillary electrophoresis (CE) traces reporting the non-templated nucleotide addition (NTA) of FBu. NTA was monitored by CE using a FAM labeled primer in the absence of RNA template (dashed) or in the presence of RNA template (solid). NTA is observed in Mg2+(Figure 6A) and Mn2+(Figure 6B).
[0019] Figure 7A-7B: Capillary electrophoresis (CE) traces reporting the template switching activity of FBu. Template switching was monitored by CE using a FAM labeled primer in the absence of RNA template and template switching oligonucleotide (dashed) or in the presence of RNA template and template switching oligonucleotide (solid). The template switching oligo nucleotide can be either DNA (Figure 7A) or RNA (Figure 7B).
[0020] DESCRIPTION
[0021] All patents and publications, including all sequences disclosed within such patents and publications, referred to herein are expressly incorporated by reference.
[0022] Numeric ranges are inclusive of the numbers defining the range. Unless otherwise indicated, nucleic acids are written left to right in 5' to 3' orientation; amino acid sequences are written left to right in amino to carboxy orientation, respectively.
[0023] The headings provided herein are not limitations of the various aspects or embodiments of the invention. Accordingly, the terms defined immediately below are more fully defined by reference to the specification as a whole.
[0024] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as would be commonly understood by one of ordinary skill in the art to which this invention belongs. Singleton, et al., DICTIONARY OF MICROBIOLOGY AND MOLECULAR BIOLOGY, 2D ED., John Wiley and Sons, New York (1994), and Hale & Markham, THE HARPER COLLINS DICTIONARY OF BIOLOGY, Harper Perennial, N.Y. (1991) provide one of skill with the general meaning of many of the terms used herein. Still, certain terms are defined below for the sake of clarity and ease of reference.
[0025] As used herein, the term "reverse transcriptase" refers to a polymerase that can make first-strand cDNA using an RNA as a template. Such enzymes are commonly referred to as RNA- directed DNA polymerases and have IUBMB activity EC 2.7.7.49. In some cases, a reverse transcriptase may have other activities, e.g., a DNA-directed DNA polymerase activity, an RNAseH activity and / or an endonuclease activity.
[0026] As used herein, the term "template" refers to a nucleic acid molecule that can be used as a template for a reverse transcriptase. The present reverse transcriptase can copy RNA templates (to make cDNA) and can copy DNA templates (e.g., to make second strand cDNA). A template may be complex (e.g., total RNA, polyA+ RNA, mRNA, or first strand cDNA made from the same, etc.) or not complex (e.g., an enriched RNA, an in vitro transcribed product, synthetic RNA, or first strand cDNA made from the same). The template may contain one or more modified nucleobase, which can be a natural nucleobase (e.g., 5-methylcytosine, N6- methyladenosime, N7-methylguanosine, 5-hydroxymethylcytosine, N-lmethyladenosine, pseudouridine) or engineered nucleobase (e.g., 5-bromouracil, 5-methyl-2-deoxycytidine, 5- azacytidine, 5-fluorouracil, dye-tagged nucleobases), or both.
[0027] The term "cDNA" refers to a strand of DNA copied from an RNA template. cDNA is complementary to the RNA template.
[0028] The term "second strand cDNA" refers to a strand of DNA copied from a cDNA template.
[0029] A "mutant" or "variant" protein may have one or more amino acid substitutions, deletions (including truncations) or additions (including deletions) relative to a wild type. A variant may have less than 100% sequence identity to the amino acid sequence of a naturally occurring protein but may have an amino acid sequence that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% identical to the amino acid sequence of the naturally occurring protein. A fusion protein is a type of protein composed of a plurality of polypeptide components that are unjoined in their naturally occurring state. Fusion proteins may be a combination of two, three or even four or more different proteins. The term polypeptide includes fusion proteins, including, but not limited to, a fusion of two or more heterologous amino acid sequences, a fusion of a polypeptide with: a heterologous targeting sequence, a linker, a purification tag (e.g. a tag that binds to a binding partner to facilitate purification and / or a tag that enhances solubility of the reverse transcriptase), an epitope tag, a detectable fusion partner, such as a fluorescent protein, P-galactosidase, luciferase, etc., and the like. A fusion protein may have one or more heterologous domains added to the N- terminus, C-terminus, and / or the middle portion of the protein. If two parts of a fusion protein are "heterologous," they are not part of the same protein in its natural state.
[0030] The term "non-naturally occurring" refers to a composition that does not exist in nature. Variant proteins are non-naturally occurring. In some embodiments, the term "non-naturally occurring" refers to a protein that has an amino acid sequence and / or a post-translational modification pattern that is different to the protein in its natural state. A non-naturally occurring protein may have one or more amino acid substitutions, deletions or insertions at the N- terminus, the C-terminus and / or between the N- and C-termini of the protein. A "non-naturally occurring" protein may have an amino acid sequence that is different from a naturally occurring amino acid sequence (i.e., having less than 100% sequence identity to the amino acid sequence of a naturally occurring protein) but that is at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98% or at least 99% identical to the naturally occurring amino acid sequence. In certain cases, a non-naturally occurring protein may contain an N-terminal methionine or may lack one or more post-translational modifications (e.g., glycosylation, phosphorylation, etc.) if it is produced by a different (e.g., bacterial) cell.
[0031] In the context of a nucleic acid, the term "non-naturally occurring" refers to a nucleic acid that contains: a) a sequence of nucleotides that is different to a nucleic acid in its natural state (i.e. having less than 100% sequence identity to a naturally occurring nucleic acid sequence), b) one or more non-naturally occurring nucleotide monomers (which may result in a non-natural backbone or sugar that is not G, A, T or C) and / or c) may contain one or more other modifications (e.g., an added label or other moiety) to the 5'- end, the 3' end, and / or between the 5'- and 3' -ends of the nucleic acid.
[0032] In the context of a preparation, the term "non-naturally occurring" refers to: a) a combination of components that are not combined by nature, e.g., because they are at different locations, in different cells or different cell compartments; b) a combination of components that have relative concentrations that are not found in nature; c) a combination that lacks something that is usually associated with one of the components in nature; d) a combination that is in a form that is not found in nature, e.g., dried, freeze dried, crystalline, aqueous; and / or e) a combination that contains a component that is not found in nature. For example, a preparation may contain a "non-naturally occurring" buffering agent (e.g., Tris, HEPES, TAPS, MOPS, tricine or MES), a detergent, a dye, a reaction enhancer or inhibitor, an oxidizing agent, a reducing agent, a solvent, or a preservative that is not found in nature.
[0033] The term "template-switching" refers to a reverse transcription reaction in which the reverse transcriptase switches template from an RNA molecule to a synthetic oligonucleotide (which usually contains two or three Gs at its 3' end, thereby copying the sequence of the synthetic oligonucleotide onto the end of the cDNA. Template switching is generally described in Matz et al., Nucl. Acids Res. 1999 27: 1558-1560 and Wu et al., Nat Methods. 2014 11: 41-6. In template switching, a primer hybridizes to an RNA molecule. This primer serves as a primer for a reverse transcriptase that copies the RNA molecule to form a complementary cDNA molecule. In copying the RNA molecule, the reverse transcriptase commonly travels beyond the 5' end of the mRNA to add non-template nucleotides to the 3' end of the cDNA (typically Cs). Addition of an oligonucleotide that has ribonucleotides or deoxyribonucleotides that are complementary to the non-template nucleotides added onto the cDNA (e.g., a "template switching" oligonucleotide that typically has two or three Gs at its 3' end), the reverse transcriptase will jump templates from the RNA template to the oligonucleotide template, thereby producing a cDNA molecule that has the complement of the template switching oligonucleotide at the 3' end.
[0034] The term "RNAseH activity" refers to an activity that hydrolyzes the RNA in an RNA / DNA hybrid. Many reverse transcriptases have an RNAseH activity that can be inactivated by truncation or by substitution.
[0035] The term "primer" refers to an oligonucleotide that is capable, upon forming a duplex with a polynucleotide template, of acting as a point of initiation of nucleic acid synthesis and being extended from its 3' end along the template so that an extended duplex is formed. The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. Primers are of a length compatible with their use in synthesis of primer extension products, and can be in the range of between 8 to 100 nucleotides in length, such as 10 to 75, 15 to 60, 15 to 40, 18 to 30, 20 to 40, 21 to 50, 22 to 45, 25 to 40, and so on, more typically in the range of between 18 to 40, 20 to 35, 21 to 30 nucleotides long, and any length between the stated ranges. Primers are usually single stranded. Primers have a 3' hydroxyl.
[0036] The term "primer extension" as used herein refers to both the synthesis of DNA resulting from the polymerization of individual nucleoside triphosphates using a primer as a point of initiation, and to the joining of additional oligonucleotides to the primer to extend the primer. Primers can incorporate additional features which allow for the detection or immobilization of the primer but do not alter the basic property of the primer, that of acting as a point of initiation of DNA synthesis. For example, primers may contain an additional nucleic acid sequence at the 5' end which does not hybridize to the target nucleic acid, but which facilitates cloning of the amplified product. The region of the primer which is sufficiently complementary to the template to hybridize is referred to herein as the hybridizing region. The terms "target region" and "target nucleic acid" refer to a region or subsequence of a nucleic acid which is to be reverse transcribed.
[0037] As noted above, provided herein is a reverse transcriptase, i.e., an RNA-directed DNA polymerase, comprising an amino acid sequence that is at least 90% identical to amino acid residues 83-841 of SEQ ID NO: 1 and has a truncated N-terminus relative to SEQ ID NO: 1. In any embodiment, the amino acid sequence may be at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% identical to amino acid residues 83-841 of SEQ ID NO: 1. In some embodiments, the amino acid sequence is identical to amino acid residues 83-841 of SEQ ID NO: 1. In any embodiment, the amino acid sequence lacks (i.e., does not have) up to 82 amino acids from the N-terminus of SEQ ID NO: 1. In these embodiments, the amino acid sequence lacks (i.e., does not have) any contiguous sequence of at least 10, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, or at least 80 amino acids at the N terminus of SEQ ID NO: 1. In some embodiments, the amino acid sequence lacks the first 82 amino acids of SEQ ID NO: 1. In addition to having an RNA-directed DNA polymerase activity, the reverse transcriptase may additionally have a DNA-directed DNA polymerase activity and, as such, may be able to copy a DNA template.
[0038] In any embodiment, the reverse transcriptase may be a fusion protein that comprises a reverse transcriptase domain and an exogenous domain, e.g., a non-specific DNA binding domain, a sequence-specific DNA binding domain, or a purification tag, e.g., a His or MBP tag or the like, at either terminus. In some embodiments, the polypeptide may comprise a sequencespecific DNA binding protein domain. Such domains have been shown to increase the processivity of other polymerases (see, e.g., US 2016 / 0160193). In some embodiments the sequence-specific DNA binding protein domain may be a DNA binding protein domain listed in the following table (as found in US 2016 / 0160193).
[0039] Table 1: DNA binding domains
[0040] Reverse transcriptases have been crystallized (see, e.g., Das et al., Structure. 2004 12: 819-29), the structure-functional relationships in reverse transcriptase have been studied (see, e.g., Cote et al., Virus Res. 2008 134: 186-202, Georgiadis et al., Structure. 1995 3: 879-92 and Crowther et al., Proteins 2004 57: 15-26) and many mutations in reverse transcriptases are known (see, e.g., Yasukawa et aL, J. BiotechnoL 2010 150: 299-306, Arezi et al Nucleic Acids Res. 2009 37: 473-81 and Konishi et aL, Biochem. Biophys. Res. Commun. 2014 454:269-74, among many others).
[0041] In any embodiment, the reverse transcriptase may lack endonuclease activity (e.g., may contain or omit C-terminal endonuclease domain of SEQ ID NO: 1) and / or the reverse transcriptase may lack RNAseH activity. See, e.g., Kotewicz et al. (Nucleic Acids Res. 1988 16: 265-77) and Schultz et al., (J. Virol. 1996 70: 8630-8).
[0042] A nucleic acid (e.g., an expression vector) comprising a nucleotide sequence encoding a subject reverse transcriptase is also provided. A subject nucleic acid may be produced by any method. Since the genetic code and recombinant techniques for manipulating nucleic acid are known, the design and production of nucleic acids encoding a subject fusion protein is well within the skill of an artisan. In certain embodiments, standard recombinant DNA technology (Ausubel, et al, Short Protocols in Molecular Biology, 3rd ed., Wiley & Sons, 1995; Sambrook, et al., Molecular Cloning: A Laboratory Manual, Second Edition, (1989) Cold Spring Harbor, N.Y.) methods are used.
[0043] A method for reverse transcribing a nucleic acid template is also provided. In some embodiments, the method may comprise: (a) combining the reverse transcriptase, a primer, the nucleic acid template, and dNTPs to produce a reaction mix and (b) incubating the reaction mix to produce a copy of the template. In some embodiments, the template may be RNA. In these embodiments, the copy may be cDNA (i.e., cDNA that is made by extending the primer using the RNA as a template). In these embodiments, the RNA template may be any type of RNA template, e.g., total RNA, polyA+RNA, capped RNA, enriched RNA etc. In alternative embodiments, the template may be DNA (e.g., first strand cDNA). In these embodiments, the copy may be second strand cDNA. The nucleic acid can be from any source, e.g., bacteria, mammals, an in vitro transcription reaction, etc., methods for the making of which are known. The nucleic acid template (RNA or DNA) may contain linear nucleic acid molecules that are at least 1 kb in length, e.g., at least 2kb, at least 3kb or at least 4 kb. In some embodiments, the nucleic acid molecules may be circular. In some embodiments, the circular nucleic acid molecules may be at least 1 kb in length, e.g., at least 2kb, at least 3kb or at least 4 kb. The primer in the reaction mix may be any type of primer, e.g., an oligo-dT primer, a random primer, or a gene-specific primer, for example, which primers are commonly used to make cDNA or other nucleic acid molecules. In some embodiments, the reaction mix may comprise a template switching oligonucleotide, as described above. In any embodiment, the reaction mix may comprise Mg2+and / or Mn2+at a suitable concentration.
[0044] In some embodiments, the reaction mix may be incubated at a temperature that is lower than 42°C, e.g., at a temperature in the range of 14°C to 40°C. In some embodiments, the reaction mix may be incubated at a temperature in the range of 14°C to 26°C, e.g., 14°C to 23°C or 23°C to 40°C, e.g., 26°C to 37°C. In some embodiments, the population of product molecules (e.g., cDNA molecules) produced by the method may be at least 20%, at least 40%, at least 60%, or at least 80% full length.
[0045] If an oligo-dT or a random primer is used in the method, then the primer may be used to make a cDNA library from RNA isolated from cells, that can be sequenced or used for gene expression analysis. Alternatively, if a gene-specific primer is used, then the method may be used for RT-PCR (e.g., quantitative RT-PCR) or sequencing, and other similar analyses. In particular embodiments, the method may be used to perform reverse transcription-rolling circle amplification (RT-RCA), which can be used to identify circRNA (a class of endogenous non-coding RNA found in mammalian cells) with high selectivity very rapidly (see, e.g., Liu et al Analytica Chimica Acta 1101: 8: 169-175). Rolling circle amplification products can be sequenced using long read sequencing technologies, such as SMRT sequencing (developed by Pacific Biosciences (PacBio)) and nanopore sequencing (developed by Oxford Nanopore Technologies) are described by the publication, including Logsdon et al. (Nature Reviews Genetics 2020 21: 597- 614), which is herein incorporated by reference in its entirety.
[0046] Briefly, in SMRT sequencing, fluorescently labelled deoxynucleoside triphosphates (dNTPs) are incorporated into the newly synthesized strand, a fluorescent dNTP is held in the detection volume, and a light pulse from the well excites the fluorophore. A camera detects the light emitted from the excited fluorophore, which records the wavelength and the position of the incorporated base in the nascent strand. The DNA sequence is determined by the changing fluorescent emission that is recorded within zero-mode waveguides.
[0047] In nanopore sequencing, a long DNA strand is tagged with sequencing adapters preloaded with a motor protein on one or both ends. The DNA is combined with tethering proteins and loaded onto the flow cell for sequencing. The flow cell contains protein nanopores embedded in a synthetic membrane. The tethering proteins bring the molecules to be sequenced towards the nanopores and as the motor protein unwinds the DNA, an electric current is applied, which drives the negatively charged DNA through the pore. The DNA is sequenced as it passes through the pore and causes characteristic changes in the current.
[0048] Also provided is a reaction mix. In some embodiments, the reaction mix may comprise a subject reverse transcriptase, a primer, a nucleic acid template (which may be DNA or RNA) and dNTPs. In any embodiment, the nucleic acid template may be a circular molecule or a linear molecule. In any embodiment, the template may be at least 1 kb in length, e.g., at least 2 kb, at least 3 kb or at least 4 kb in length. In some embodiments, the reaction mix may comprise a template switching oligonucleotide, as described above. In any embodiment, the reaction mix may comprise Mg2+and / or Mn2+at a suitable concentration, and may optionally comprise a detergent. The reaction mix may be used in the range of 14°C to 40°C. In some embodiments, the reaction mix may be used at a temperature in the range of 14°C to 26°C, e.g., 14°C to 23°C or 26°C to 40°C, e.g., 26°C to 37°C.
[0049] Also provided by this disclosure are kits that contain the reverse transcriptase as described above. In certain embodiments, the kit may comprise any of the components listed above, e.g., one or more of a buffer, dNTPs, a primer, and a template switching oligonucleotide, for example. The reverse transcriptase may be in a storage buffer that contains 10-50% glycerol or lacks glycerol in some embodiments. The kit may additionally contain other agents, including buffers and other components described above. The various components of the kit may be present in separate containers or certain compatible components may be pre-combined into a single container, as desired. In addition to the above-mentioned components, the subject kit may further include instructions for using the components of the kit to practice the subject method.
[0050] Embodiments include but are not limited to:
[0051] Embodiment 1. A reverse transcriptase comprising an amino acid sequence that: is at least 90% identical to amino acid residues 83-841 of SEQ ID NO: 1; and has a truncated N- terminus relative to SEQ ID NO: 1.
[0052] Embodiment 2. The reverse transcriptase of embodiment 1, comprising an amino acid sequence that is at least 95% identical to amino acid residues 83-841 of SEQ ID NO: 1.
[0053] Embodiment 3. The reverse transcriptase of embodiment 1 or 2, comprising an amino acid sequence that is identical to amino acid residues 83-841 of SEQ ID NO: 1.
[0054] Embodiment 4. The reverse transcriptase of any prior embodiment, wherein the amino acid sequence lacks up to 82 amino acids from the N-terminus of SEQ ID NO: 1.
[0055] Embodiment 5. The reverse transcriptase of any prior embodiment, wherein the reverse transcriptase lacks endonuclease activity.
[0056] Embodiment 6. The reverse transcriptase of any prior embodiment, wherein the reverse transcriptase is a fusion protein that comprises a reverse transcriptase domain and an exogenous domain.
[0057] Embodiment 7. The reverse transcriptase of embodiment 6, wherein the exogenous domain is a sequence-specific DNA binding domain.
[0058] Embodiment 8. The reverse transcriptase of embodiment 6, wherein the exogenous domain comprises a purification tag.
[0059] Embodiment 9. The reverse transcriptase of embodiment 8, wherein the exogenous domain is a purification tag. Embodiment 10. The reverse transcriptase of embodiment 8 or 9, wherein the purification tag is maltose binding protein.
[0060] Embodiment 11. A method for copying a nucleic acid template, comprising combining the reverse transcriptase of any prior embodiment with a primer, the nucleic acid template, and dNTPs to produce a reaction mix; and incubating the reaction mix to produce a copy of the template.
[0061] Embodiment 12. The method of embodiment 11, wherein the nucleic acid template is DNA.
[0062] Embodiment 13. The method of embodiment 11, wherein the nucleic acid template is RNA, and the copy is cDNA.
[0063] Embodiment 14. The method of any embodiment, wherein the nucleic acid template is a circular nucleic acid or a linear nucleic acid.
[0064] Embodiment 15. The method of any embodiment, wherein the nucleic acid template is at least 1 kb in length.
[0065] Embodiment 16. The method of any embodiment, wherein the reaction mix further comprises a template switching oligonucleotide.
[0066] Embodiment 17. The method of any of embodiment, wherein the reaction mix further comprises Mg2+ or Mn2+.
[0067] Embodiment 18. The method of any of embodiment, wherein the incubation is done at a temperature in the range of 15°C-40°C.
[0068] Embodiment 19. The method of any embodiment, wherein the incubation is done at a temperature in the range of 14°C to 26°C, 14°C to 23°C, 26°C to 40°C, or 26°C to 37°C.
[0069] Embodiment 20. A reaction mix comprising: the reverse transcriptase of any embodiments 1-10; a primer; a nucleic acid template; and dNTPs.
[0070] Embodiment 21. The reaction mix of embodiment 20, wherein the nucleic acid template is DNA.
[0071] Embodiment 22. The reaction mix of embodiment 20, wherein the nucleic acid template is RNA.
[0072] Embodiment 23. The reaction mix of any embodiment, wherein the nucleic acid template is a circular nucleic acid or a linear nucleic acid.
[0073] Embodiment 24. The reaction mix of any embodiment, wherein the nucleic acid template is at least 1 kb in length. Embodiment 25. The reaction mix of any embodiment, wherein the reaction mix further comprises a template switching oligonucleotide.
[0074] Embodiment 26. The reaction mix of any embodiment, wherein the reaction mix comprises Mg2+ or Mn2+.
[0075] Embodiment 27. A kit comprising: a reverse transcriptase of any of embodiments 1-10; and one or more of a buffer, dNTPs, and a primer.
[0076] Embodiment 28. The kit of embodiment 27, further comprising a template switching oligonucleotide.
[0077] Aspects of the present teachings can be further understood in light of the following examples, which should not be construed as limiting the scope of the present teachings in any way.
[0078] In the following examples, the N-terminally truncated reverse transcriptase from the type-2 retrotransposable element R2DM of Fasciolopsis buski, as defined by Genbank accession no. KAA0201068.1 is referred to as "FBu". An N-terminally truncated R2 reverse transcriptase from Bombyx mori as defined by Genbank accession no. AAB59214.1 is referred to as "Brno". See Upton et al (Proc Natl Acad Sci. 2021 118: e2107900118) and Pimentel (J. Biol. Chem. 2022 298:101624).
[0079] The full length amino acid sequence of the reverse transcriptase from the type-2 retrotransposable element R2DM of Fasciolopsis buski (Genbank accession no. KAA0201068.1) is set forth below as SEQ. ID NO: 1.
[0080] MSRCRNVRWTDQECSRLLSLAGRREEGTSIARLSTVLASEFPTRSREAVRLRLKALRRVHPT WGGLGSGDPITAATGASAPVVPNNEWSTRLLEVVMAELMVHSDESLGSHELLQLVRDFN EGKSTRQEAAEGLERLMATSFPHVWQAREQRPSTGRTLGLTRKKVRRVQYAKVQSLYKR RTKDAADMVLSGDWSSSHLSERRQPEGQSSFWRHLFEQKSVRDERPVQGLVNHWQVLE PITANEVKQAAKAIGDSAAGMDKVNAGALLTKDLTVVAQLFNVMLLLEMPTQQLSKARV TLIPKAQTPSGPSDYRPIAVSSVVLRILHKIMAQRMMKCIKLERLQVAFQKRDGCMEAAET LNACLREAQEKSKNLAAAFVDVSKAFDTVSHDSILRAAQRQGFPPPLLNYLRRLYDGSTVQ LCGEEVRCRRGVRQGDPLSPILFIAVIDDVLSSLPCFGFPLGSERIVDVLAYADDLVLFTENEA ALQSKLNGLADALALVGMTVNASKSRALTITANKHNKTVVLRPVSYTIGNSKIRGMSVSDR VQYLGLSFGWKGKIPVKHTGVLEGWIKNVTQAPLKPYQRMSILRNHIVPRLLHGMVLGAA HRNTLKAVDIMLRQAVKAWLRLPKDTTSAFLHGPVNAGGLGIPCLSVMVPLAQRKRLENL ARQTDPTIGVVREKECFRKFVRQANLPIQVGSKIVLSKEEARKAWADALVESKDGRALEGG EVDSASHSWVREPSKLPAKVFIRGVKLRGGLLPTKVRAARGRAVPQNEVICKGRCGQPESI DHILQSCPITHDVRCERHDRVVRKLGKELSKTMERVWLEPIVPSGSTFCKPRDAG The fusion protein used in the following experiments has an N-terminal truncation of 82 amino acids. As such, the reverse transcriptase used in the following experiments comprises amino acid residues 83-841 of SEQ ID NO: 1 and lacks the N-terminal 82 amino acids of SEQ ID NO: 1 (underlined above).
[0081] Example 1: Construct design, protein expression and purification
[0082] A coding sequence for a fusion protein comprising an N-terminally truncated reverse transcriptase from the type-2 retrotransposable element R2DM of Fasciolopsis buski (composed of amino acid residues 83-841 SEQ ID NO: 1) was cloned into pMalc5E. The fusion protein (which contains an N-terminal maltose binding portion (MPB) domain was expressed and purified using the following protocol:
[0083] 1) Using plasmids described above, transform into C3013 cells. Plate on LB-Amp for overnight growth at 37°C.
[0084] 2) Inoculate 8.5 mL of LB-Amp liquid media with a single colony and shake at 30°C for ~8 hours. OD should reach ~1.5.
[0085] 3) Starting at ~5:00 PM, Add 1 mL of starter culture above to 1 L of LB-Amp equilibrated to 20°C. Shake over night at 20°C (~5 PM - 7 AM). OD595should be between 0.4 and 0.6.
[0086] 4) Induce with 0.3 mM IPTG. Reduce temperature to 16°C and shake for 9 hours.
[0087] 5) Harvest cells by centrifugation. Pellets can be frozen at this point in -80°C at this point or moved directly to lysis and purification phase.
[0088] 6) Take pellet up in 20 mL of Lysis Buffer: 50 mM Tris pH 7.5, IM NaCI, 2 mM MgCI2, 0.05% P-Mercaptoethanol, 10% glycerol, lx protease inhibitors (PMSF). Perform this step on ice. Pipette to resuspend.
[0089] 7) Sonicate for 3 minutes IOS / IOS (on / off) 60% power, on ice water.
[0090] 8) Centrifuge at 18k RPM for 20 minutes at 4°C. If desired, filter with 0.22 pm filter.
[0091] 9) Load sample onto 5 mL MBP-Trap column using AKTA PURE. Make sure to keep lysate on ice the entire time. Wash column with 5 column volumes of lx lysis buffer after loading.
[0092] 10) Wash column with 10 column volumes of 0.5x Lysis buffer.
[0093] 11) Elute with 50mL of 0.5x Lysis Buffer supplemented with 20 mM Maltose. Collect in flask at 4°C.
[0094] 12) Dilute sample 5-fold with Heparin A. Load sample onto 5mL heparin column equilibrated with Heparin A Buffer: 5 mM Tris pH 7.5, 400 mM NaCI, 1 mM MgCI2, 2% Glycerol, 0.2 mM DTT. Wash column with 60 mL of Heparin A. 13) Perform elution with 5% incremental gradient with Heparin B Buffer: 20 mM Tris pH 7.5, 2M NaCI, 2 mM MgCI2, 10% glycerol, 1 mM DTT. (Protein eluted around 50% Heparin B. Elution was tuned to elute over finer gradient). Analyze fractions on SDS polyacrylamide gel and pool fractions.
[0095] 14) Concentrate pooled fraction in 100 kD Amicon unit. Exchange into 50 mM Tris pH 7.5, 600 mM NaCI, 4 mM MgCI2, 5% Glycerol, 2mM DTT. After concentration, dilute l / 2x with 100% glycerol.
[0096] The isolated fusion protein was >95% pure.
[0097] Example 2: Biochemical characterization
[0098] The FBu fusion protein was tested for its ability to produce cDNA using 1 kb, 2 kb, and 4 kb linear templates, as well as a 1.5 kb circular template.
[0099] The following reaction conditions were used for the linear templates: 50 mM Tris pH 8.0, 10 mM NaCI, 5 mM MgCb, 2 mM DTT, 0.5 mM dNTP.
[0100] The following reaction conditions were used for the circular template: 50 mM Tris pH 8.0, 250 mM NaCI, 4.5 mM MgCI2, 0.02% Tween-20.
[0101] Temperature dependence
[0102] First strand cDNA was generated by several reverse transcriptases, including the Fasciolopsis buski reverse transcriptase (referred to as "FBu"), a group II intron-encoded reverse transcriptase sold commercially by New England Biolabs and referred to as "Induro," the BMo reverse transcriptase (see above), and ProtoScript II. Second strand cDNA was made using ProtoScript II and products were analyzed on the TapeStation.
[0103] The temperature dependence of cDNA synthesis of the reverse transcriptases was analyzed. Synthesis yields for the 1 kilobase and 4 kilobase linear templates at various temperatures are shown in Figures 1A and IB. These results show that the FBu reverse transcriptase produces approximately the same or more cDNA as the best reverse transcriptase tested (Induro) at lower temperatures and produces more cDNA (i.e., has a greater yield) than the BMo and ProtoScript II reverse transcriptases.
[0104] Data showing the activity of these enzymes on circular template at temperatures in the range of 35°C to 50°C is shown in Figure 2. This data shows that unlike the BMo reverse transcriptase, the FBu reverse transcriptase is active on circular templates in this temperature range. Data showing the activity of the Induro and FBu reverse transcriptases on the circular template in lower temperatures (16°C to 26°C) is shown in Figure 3. This data shows that FBu is more active than Induro reverse transcriptase on circular templates at lower temperatures.
[0105] The following table quantifies the yield of cDNA produced using the circular template by the Induro and FBu reverse transcriptases.
[0106] Table 2: cDNA product yield at various temperatures
[0107] Other experiments revealed that FBu has only a modest difference in error rate between cDNA and second strand synthesis, which is not the case for other R2 RTs. Table 3. Error rates (x IO-6) for cDNA and second-strand synthesis
[0108] Non-templated addition (NTA) is correlated with template switching (TS). Reverse transcriptases that have a high propensity to template switch also have greater amount of non- templated addition. FBu readily performs non-templated addition (Figure 4). While +1 and +2 non-templated addition have been shown to promote template switching, > +2 non-templated addition is thought to inhibit template switching. Non-templated addition for both BMo and FBu is Mg2+dependent. Increasing concentration of Mg2+promotes greater number of non- templated nucleotide additions. Low Mg2+concentration conditions (1 mM Mg2+) improve template switching relative to higher concentrations for R2-RTs generally.
[0109] While the number of non-templated addition impacts template switching efficiency, the composition of non-templated addition is also relevant. Having balanced proportions of each nucleotide in the NTA enables capture of more diverse transcripts. Relative to ProtoScript II, R2- RT has greater diversity of NTA at both +1 and +2 positions (Figure 5). This is thought to be what enables BMo to obtain low-bias template switching. The abundance of T in the RT profiles is advantageous. Not only do reverse transcriptases usually select against T addition, non- templated addition with T can help capture the majority of the transcriptome that encodes 3' A, such as poly and CCA tails found in tRNAs and snoRNAs. Therefore, FBu is expected to outperform many RTs in tRNA sequencing.
[0110] Capillary electrophoresis (CE) traces showing the non-templated nucleotide addition (NTA) of FBu are shown in Figures 6A and 6B. This data shows that non-templated addition is observed in the presence of Mg2+and Mn2+. These data show that FBu is highly active in NTA, a hall mark of reverse transcriptases that readily perform template switching and tailing.
[0111] Capillary electrophoresis (CE) traces showing the template switching activity of FBu are shown in Figures 7A and 7B. This data shows that the template switching oligonucleotide can be either DNA or RNA.
[0112] FBu is useful for performing second strand synthesis. Second strand synthesis is performed by obtaining single stranded cDNA, treated to remove RNA by either RNase of heat denaturation under basic conditions, performing second strand DNA synthesis on this starting material. 30 pL of single stranded cDNA at 0.5 pM is combined with 2 pL of a primer complementary to the 3' end of the cDNA. The length of this primer can range from 6 to 30 nucleotides with variable sequence. The combined sample is heated to a temperature above the melting temperature of the annealed primer and template and allowed to cool to room temperature. A concentrated solution containing a buffer, monovalent and divalent salt is added to the solution so that the components will all be at a suitable concentration at 50 plvolume. dNTPs are added to the solution so that their concertation will be between 0.1 - 1 mM at a final volume of 50 pL. Finally, water and enzyme are added to the solution and the reaction is allowed to proceed at a given temperature.
[0113] While the present invention has been described with reference to the specific embodiments thereof, it should be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departing from the true spirit and scope of the invention. In addition, many modifications may be made to adapt a particular situation, material, composition of matter, process, process step or steps, to the objective, spirit, and scope of the present invention. All such modifications are intended to be within the scope of the claims appended hereto.
Claims
CLAIMSWhat is claimed is:
1. A reverse transcriptase comprising an amino acid sequence that: is at least 90% identical to amino acid residues 83-841 of SEQ ID NO: 1; and has a truncated N-terminus relative to SEQ ID NO: 1.
2. The reverse transcriptase of claim 1, comprising an amino acid sequence that is at least 95% identical to amino acid residues 83-841 of SEQ ID NO: 1.
3. The reverse transcriptase of claim 1 or 1, comprising an amino acid sequence that is identical to amino acid residues 83-841 of SEQ ID NO: 1.
4. The reverse transcriptase of any prior claim, wherein the amino acid sequence lacks up to 82 amino acids from the N-terminus of SEQ ID NO: 1.
5. The reverse transcriptase of any prior claim, wherein the reverse transcriptase lacks endonuclease activity.
6. The reverse transcriptase of any prior claim, wherein the reverse transcriptase is a fusion protein that comprises a reverse transcriptase domain and an exogenous domain.
7. The reverse transcriptase of claim 6, wherein the exogenous domain is a sequencespecific DNA binding domain.
8. The reverse transcriptase of claim 6, wherein the exogenous domain comprises a purification tag.
9. The reverse transcriptase of claim 8, wherein the exogenous domain is a purification tag.
10. The reverse transcriptase of claim 8 or 9, wherein the purification tag is maltose binding protein.
11. A method for copying a nucleic acid template, comprising:(a) combining the reverse transcriptase of any prior claim with a primer, the nucleic acid template, and dNTPs to produce a reaction mix; and(b) incubating the reaction mix to produce a copy of the template.
12. The method of claim 11, wherein the nucleic acid template is DNA.
13. The method of claim 11, wherein the nucleic acid template is RNA, and the copy is cDNA.
14. The method of any claim, wherein the nucleic acid template is a circular nucleic acid or a linear nucleic acid.
15. The method of any claim, wherein the nucleic acid template is at least 1 kb in length.
16. The method of any claim, wherein the reaction mix further comprises a template switching oligonucleotide.
17. The method of any of claim, wherein the reaction mix further comprises Mg2+or Mn2+.
18. The method of any of claim, wherein the incubation is done at a temperature in the range of 15°C-40°C.
19. The method of any claim, wherein the incubation is done at a temperature in the range of 14°C to 26°C, 14°C to 23°C, 26°C to 40°C, or 26°C to 37°C.
20. A reaction mix comprising: the reverse transcriptase of any claims 1-10; a primer; a nucleic acid template; and dNTPs.
21. The reaction mix of claim 20, wherein the nucleic acid template is DNA.
22. The reaction mix of claim 20, wherein the nucleic acid template is RNA.
23. The reaction mix of any claim, wherein the nucleic acid template is a circular nucleic acid or a linear nucleic acid.
24. The reaction mix of any claim, wherein the nucleic acid template is at least 1 kb in length.
25. The reaction mix of any claim, wherein the reaction mix further comprises a template switching oligonucleotide.
26. The reaction mix of any claim, wherein the reaction mix comprises Mg2+or Mn2+.
27. A kit comprising: a reverse transcriptase of any of claims 1-10; and one or more of a buffer, dNTPs, and a primer.
28. The kit of claim 1 , further comprising a template switching oligonucleotide.
Citation Information
Patent Citations
Fusion Polymerase and Method for Using the Same
US20160160193A1