Systems, compositions, and methods involving retrotransposons and functional fragments thereof
Patent Information
- Application Number
- JP2024513336
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-08
- Filing Date
- 2022-09-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies have not fully exploited the potential of transposable elements for DNA manipulation and gene editing applications, despite their prevalence in eukaryotic genomes.
Development of an engineered retrotransposase system comprising a retrotransposase configured to transpose cargo nucleotide sequences to target nucleic acid loci, utilizing RNA sequences and double-stranded DNA structures, with specific recognition sequences and high sequence identity, enabling precise genetic manipulation.
Facilitates efficient and targeted genetic modification by allowing precise transposition of cargo nucleotide sequences into desired genomic locations, enhancing the utility of transposable elements in DNA manipulation and gene editing.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of U.S. Provisional Application No. 63 / 241,943, filed September 8, 2021, entitled “SYSTEMS AND METHODS FOR TRANSPOSING CARGO NUCLEOTIDE SEQUENCES,” which is incorporated herein by reference in its entirety. [Background technology]
[0002] Transposable elements are mobile DNA sequences that play important roles in gene function and evolution. Transposable elements are found in almost all types of life forms, but their prevalence varies among organisms, and the majority of eukaryotic genomes encode transposable elements (at least 45% in humans).
[0003] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in XML format, and is incorporated by reference in its entirety. The XML copy created on September 7, 2022 is named 55921-734_601_SL.xml and is 1,677,029 bytes in size. Summary of the Invention
[0004] Although fundamental research on transposable elements was carried out in the 1940s, their potential utility in DNA engineering and gene editing applications has only recently been recognized.
[0005] In some aspects, the disclosure provides a method for the preparation of a retrotransposase comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase has a nucleic acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 109%, at least about 110, at least about 111, at least about 112, at least about 113, at least about 114, at least about 115, at least about 116, at least about 117, at least about 118, at least about 119, at least about 120, at least about 122, at least about 124, at least about 126, at least about 128, at least about 129, at least about 130, at least about 131, at least about 132, at least about 133, at least about 134, at least about 135, at least about 136, at least about 137, at least about 138, at least about 139, at least about 140, at least about 141, at least about 142, at least about 143, at least about 144, at least about 145 and a retrotransposase comprising an RT domain, an endonuclease domain, and a sequence having 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of any one of SEQ ID NOs: 1-29 or 393-401, or a variant thereof. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29 or 393-401, or a variant thereof. In some embodiments, the retrotransposase further comprises a catalytic D, QG, [Y / F]XDD, or LG motif conserved to any of the sequences in Figure 2A. In some embodiments, the retrotransposase further comprises a CX [2-3]In some embodiments, the retrotransposase further comprises a C Zn finger motif. In some embodiments, the retrotransposase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 6, 7, 8, 14, or 402, or a variant thereof. In some embodiments, the system further comprises (c) a double-stranded DNA sequence comprising a target nucleic acid locus. In some embodiments, the double-stranded DNA sequence comprises a 5' recognition sequence and a 3' recognition sequence configured to interact with the retrotransposase, wherein the 5' recognition sequence comprises a GG nucleotide sequence and the 3' recognition sequence comprises a TGAC nucleotide sequence. In some embodiments, the RNA is an in vitro transcribed RNA. In some embodiments, the RNA comprises a sequence 5' to a cargo sequence or a sequence 3' to a cargo sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RNA cognate of any one of SEQ ID NOs: 761-798, its complement, or its reverse complement. In some embodiments, the RNA comprises a sequence encoding a retrotransposase. In some embodiments, the heterologous engineered cargo nucleotide sequence comprises an expression cassette.
[0006] In some embodiments, the disclosure provides a method for the preparation of a 5′ sequence capable of encoding an RNA sequence configured to interact with a retrotransposase, (a) a 5′ sequence capable of encoding an RNA sequence configured to interact with a retrotransposase, (b) a heterologous cargo sequence, and (c) a sequence encoding a retrotransposase configured to interact with an RNA cognate of the 5′ sequence, wherein the retrotransposase has a homology ratio of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, or at least about 86% to a reverse transcriptase (RT) or endonuclease domain of any one of SEQ ID NOs: 1-29 or 393-401. and (d) a 3' sequence capable of encoding an RNA sequence configured to interact with a retrotransposase. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of any one of SEQ ID NOs: 1-29 or 393-401, or a variant thereof. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29 or 393-401, or a variant thereof. In some embodiments, the retrotransposase further comprises a catalytic D, QG, [Y / F]XDD, or LG motif conserved to any of the sequences in Figure 2A. In some embodiments, the retrotransposase further comprises a CX [2-3]In some embodiments, the retrotransposase further comprises a C Zn finger motif. In some embodiments, the retrotransposase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 6, 7, 8, 14, or 402, or a variant thereof. In some embodiments, the 5' or 3' sequence comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RNA cognate of any one of SEQ ID NOs: 761-798, its complement, or its reverse complement.
[0007] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) providing a primer oligonucleotide that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, or at least about 86% relative to a reverse transcriptase domain of any one of SEQ ID NOs: 1-29, 393-401, or 427-439. and synthesizing a primer oligonucleotide-primed cDNA from a template using a reverse transcriptase comprising a sequence having at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 799-894 or 427-439, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, synthesizing the cDNA comprises incubating the template RNA molecule, the primer oligonucleotide, and the reverse transcriptase in a reaction mixture under conditions suitable for elongation of a DNA sequence from the RNA template. In some embodiments, the reaction mixture contains dNTPs, a reaction buffer, divalent metal ions, Mg 2+ , or Mn 2+ Further includes:
[0008] In some aspects, the disclosure provides a protein comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs: 1-29, 393-401, or 427-439, or a variant thereof, wherein the sequence is fused at the N-terminus or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 799-894, 427-439, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
[0009] In some aspects, the disclosure provides nucleic acids encoding any of the proteins described herein.
[0010] In some aspects, the disclosure provides a nucleic acid encoding an open reading frame, the open reading frame having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 129, at least about 130, at least about 131, at least about 132, at least about 133, at least about 134, at least about 135, at least about 136, at least about 137, at least about 138, at least about 139, at least about 140, at least about 141, at least about 14 or a variant thereof, having about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, and (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT or endonuclease domain of any one of SEQ ID NOs: 1-29, 393-401, or 427-439, or a variant thereof.
[0011] In some embodiments, the disclosure provides a method for the synthesis of a nucleic acid sequence comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase has a nucleic acid sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at and a retrotransposase comprising an RT domain or an endonuclease domain comprising a sequence having about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any of the sequences, or variants thereof. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or variants thereof. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a conserved CX [2-3] In some embodiments, the system further comprises (c) a double-stranded DNA sequence comprising a target locus. In some embodiments, the RNA is in vitro transcribed RNA. In some embodiments, the RNA comprises a sequence encoding a retrotransposase.
[0012] In some aspects, the disclosure provides a method for the preparation of a retrotransposase comprising: (a) a 5′ sequence capable of encoding an RNA sequence configured to interact with a retrotransposase; (b) a heterologous cargo sequence; and (c) a sequence encoding a retrotransposase configured to interact with an RNA cognate of the 5′ sequence, wherein the retrotransposase has a homology ratio of at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 110%, at least about 111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132, at least about 133, at least about 1 The present invention provides an engineered DNA sequence comprising a sequence comprising an RT domain, an endonuclease domain, and (d) a 3' sequence capable of encoding an RNA sequence configured to interact with a retrotransposase, the RT domain comprising a sequence having at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the retrotransposase, or a variant thereof. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a conserved CX [2-3] C further contains a Zn finger motif.
[0013] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), the method comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing cDNA initiated by the primer oligonucleotide from the template using a reverse transcriptase that comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, divalent metal ions, Mg 2+ , or Mn 2+ Further includes.
[0014] In some aspects, the disclosure provides a protein comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of SEQ ID NO: 402 or 895, or a variant thereof, wherein the sequence is fused N- or C-terminally to a non-retrotransposase domain or affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
[0015] In some aspects, the disclosure provides nucleic acids encoding an open reading frame, the open reading frame encoding an RT or endonuclease domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT or endonuclease domain of SEQ ID NO: 402 or 895, or a variant thereof, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof.
[0016] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing the primer oligonucleotide-initiated cDNA from the template using a reverse transcriptase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to a reverse transcriptase domain of any one of SEQ ID NOs: 555-728, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, 607, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the primer oligonucleotide comprises at least one phosphorothioate linkage.In some embodiments, synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, a divalent metal ion, Mg. 2+ , or Mn 2+ Further includes.
[0017] In some embodiments, the present disclosure provides a protein comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs: 555-728, or a variant thereof, wherein the sequence is fused at the N-terminus or C-terminus to a non-retrotransposase domain or affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, 607, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 30-32, 40-50, 740-756, 757-760, or a variant thereof.In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-558, 561-567, 569, 570, 575, or a variant thereof.
[0018] In some aspects, the disclosure provides a nucleic acid encoding an open reading frame, the open reading frame encoding an RT or endonuclease domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RT or endonuclease domain of any one of SEQ ID NOs:555-728, or a variant thereof, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least 80% sequence identity to the RT or endonuclease domain of any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, 607, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, or a variant thereof.
[0019] In some aspects, the disclosure provides a nucleic acid comprising a sequence comprising an open reading frame (ORF) comprising a sequence encoding a reverse transcriptase domain or maturase domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain or maturase domain of any one of SEQ ID NOs:729-733, or a variant thereof, wherein (a) the open reading frame is optimized for expression in an organism, the organism being different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the ORF encodes a protein having at least 80% sequence identity to any one of SEQ ID NOs: 729-733, or a variant thereof. In some embodiments, the ORF is optimized for expression in a bacterial organism, or the organism is E. coli. In some embodiments, the ORF is optimized for expression in a mammalian organism, or the organism is a primate organism. In some embodiments, the primate organism is Homo sapiens. In some embodiments, the ORF comprises an affinity tag operably linked to a sequence encoding a reverse transcriptase domain or a maturase domain, the ORF having at least 80% sequence identity to any one of SEQ ID NOs: 298-302. In some embodiments, the ORF comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 303-307. In some embodiments, the reverse transcriptase domain or maturase domain comprises a conserved Y[I / L]DD active site motif of any one of SEQ ID NOs: 729-733.
[0020] In some embodiments, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing the primer oligonucleotide-initiated cDNA from the template using a reverse transcriptase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to a reverse transcriptase domain of any one of SEQ ID NOs: 440-554, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 526, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongation of a DNA sequence from an RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, a divalent metal ion, Mg 2+ , or Mn 2+ Further includes.
[0021] In some embodiments, the present disclosure provides a protein comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs: 440-554, or a variant thereof, wherein the sequence is fused at the N-terminus or C-terminus to a non-retrotransposase domain or affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 526, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the sequence is fused N- or C-terminally to an affinity tag.
[0022] In some aspects, the disclosure provides a nucleic acid encoding an open reading frame, the open reading frame encoding an RT domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT domain of any one of SEQ ID NOs: 440-554, or a variant thereof, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes an RT having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 526, or a variant thereof. In some embodiments, the open reading frame comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 356-373.
[0023] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) providing a primer oligonucleotide that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, or at least about 84% to a reverse transcriptase domain of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, 627-673. and synthesizing a primer oligonucleotide-primed cDNA from a template using a reverse transcriptase comprising a sequence having about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the reverse transcriptase domain comprises a conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, or 627-673. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, 633, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the primer oligonucleotide comprises at least six consecutive nucleotides having at least 80% sequence identity to any one of SEQ ID NOs: 340-341, 342-344, 345-346, 347-351, 352, or 353-355.In some embodiments, synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, a divalent metal ion, Mg. 2+ , or Mn 2+ Further includes.
[0024] In some embodiments, the present disclosure provides a protein comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, 627-673, or a variant thereof, wherein the sequence is fused at the N-terminus or C-terminus to a non-retrotransposase domain or affinity tag. In some embodiments, the reverse transcriptase domain comprises a conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, or 627-673. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, 633, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the sequence is fused N-terminally or C-terminally to an affinity tag.
[0025] In some aspects, the present disclosure provides a nucleic acid encoding an open reading frame (ORF) optimized for expression in an organism, the open reading frame having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109 ... %, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RT domain, or a variant thereof, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the reverse transcriptase domain comprises a conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs:609-610, 611-615, 616-617, 618-622, 623, 624-626, or 627-673. In some embodiments, the nucleic acid further encodes an RT having at least 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, 633, or a variant thereof. In some embodiments, the ORF comprises a sequence encoding an affinity tag. In some embodiments, the open reading frame comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 308-309, 310-312, 313-314, 315-319, 320, 321-323, or 174-180. In some embodiments, the organism is different from the source of the RT domain.In some embodiments, the ORF comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 324-325, 326-328, 329-330, 331-335, 336, 327-329, or 181-187.
[0026] In some aspects, the present disclosure provides a synthetic oligonucleotide comprising at least 6 consecutive nucleotides having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 340-341, 342-344, 345-346, 347-351, 352, or 353-355. In some embodiments, the synthetic oligonucleotide comprises DNA nucleotides. In some embodiments, the oligonucleotide further comprises at least one phosphorothioate linkage.
[0027] In some embodiments, the present disclosure provides a vector comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 340-341, 342-344, 345-346, 347-351, 352, or 353-355.
[0028] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein.
[0029] In some aspects, the disclosure provides a host cell comprising any of the nucleic acids described herein. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the nucleic acid comprises an open reading frame (ORF) encoding a retrotransposase, a fragment thereof, or a reverse transcriptase domain, the open reading frame being selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a raP ... BAD promoter, the strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to a sequence encoding a retrotransposase, a fragment thereof, or a reverse transcriptase domain.
[0030] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.
[0031] In some aspects, the disclosure provides a method of producing a retrotransposase, a fragment thereof, or a reverse transcriptase domain, comprising culturing any of the host cells described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the retrotransposase, a fragment thereof, or a reverse transcriptase domain by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or the increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to affinity or ion affinity chromatography specific for the affinity tag.
[0032] In some aspects, the disclosure provides an in vitro transcribed mRNA comprising an RNA homolog of any of the nucleic acids described herein.
[0033] In some aspects, the disclosure provides an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, the cargo nucleotide sequence configured to interact with a retrotransposase; and (b) a retrotransposase, (i) the retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase derived from an uncultured microorganism. In some embodiments, the cargo nucleotide sequence is engineered. In some embodiments, the cargo nucleotide sequence is heterologous. In some embodiments, the cargo nucleotide sequence does not have the sequence of a wild-type genomic sequence present in the organism. In some embodiments, the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to a documented retrotransposase. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. In some embodiments, the retrotransposase comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the retrotransposase. In some embodiments, the NLS comprises a sequence that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 896-911. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm.In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0034] In some aspects, the disclosure provides an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, the cargo nucleotide sequence configured to interact with a retrotransposase; and (b) a retrotransposase, (i) the retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase comprising a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29. In some embodiments, the retrotransposase is derived from an uncultured microorganism. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to a documented retrotransposase. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. In some embodiments, the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW with parameters of the Smith-Waterman homology search algorithm. In some embodiments, the sequence identity is determined by the BLASTP homology search algorithm using parameters of word length (W) of 3, expectation (E) of 10, and gap cost at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0035] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide encoding the engineered retrotransposase system according to any one of the aspects or embodiments described herein.
[0036] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a retrotransposase, the retrotransposase being derived from an uncultured microorganism, and the organism is not an uncultured microorganism. In some embodiments, the retrotransposase comprises a variant having at least 75% sequence identity to any one of SEQ ID NOs: 1-29. In some embodiments, the retrotransposase comprises a sequence encoding one or more nuclear localization sequences (NLSs) proximal to the N-terminus or C-terminus of the retrotransposase. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 896-911. In some embodiments, the NLS comprises SEQ ID NO: 897. In some embodiments, the NLS is proximal to the N-terminus of the retrotransposase. In some embodiments, the NLS comprises SEQ ID NO: 896. In some embodiments, the NLS is proximal to the C-terminus of the retrotransposase. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0037] In some aspects, the present disclosure provides a vector comprising a nucleic acid according to any one of the aspects or embodiments described herein. In some embodiments, the vector further comprises a nucleic acid encoding a cargo nucleotide sequence configured to form a complex with a retrotransposase. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0038] In some aspects, the disclosure provides a cell comprising the vector of any one of the aspects or embodiments described herein.
[0039] In some aspects, the disclosure provides a method of producing a retrotransposase comprising culturing a cell according to any of the aspects or embodiments described herein.
[0040] In some aspects, the disclosure provides a method for binding, nicking, cleaving, marking, modifying, or transposing a double-stranded deoxyribonucleic acid polynucleotide, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a transposase configured to transpose a cargo nucleotide sequence to a target nucleic acid locus, wherein the transposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29. In some embodiments, the retrotransposase is derived from an uncultured microorganism. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to a documented retrotransposase. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is translocated via a ribonucleic acid polynucleotide intermediate. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0041] In some aspects, the disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering to a target nucleic acid locus an engineered retrotransposase system according to any one of the aspects or embodiments described herein, wherein the retrotransposase is configured to transpose a cargo nucleotide sequence to the target nucleic acid locus, and the complex is configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, marking, modifying, or transposing the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA). In some embodiments, the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC).
[0042] In some aspects, the disclosure provides a method according to any one of the aspects or embodiments described herein, wherein delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid according to any one of the aspects or embodiments described herein, or a vector according to any one of the aspects or embodiments described herein. In some embodiments, delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding a retrotransposase. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding a retrotransposase is operably linked. In some embodiments, delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding a retrotransposase. In some embodiments, delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, the retrotransposase does not induce cleavage at or proximal to the target nucleic acid locus.
[0043] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding a heterologous transposase having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, or a variant thereof. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a araP promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a raP ...BAD In some embodiments, the open reading frame is operably linked to a nucleotide sequence encoding an affinity tag linked in-frame to the sequence encoding the retrotransposase. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a Tobacco Etch Virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of the host cell.
[0044] In some aspects, the disclosure provides a culture comprising a host cell according to any one of the aspects or embodiments described herein in a suitable liquid medium.
[0045] In some aspects, the disclosure provides a method of producing a retrotransposase, comprising culturing a host cell according to any one of the aspects or embodiments described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the retrotransposase by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or the increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cell after culturing and lysing the host cell to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC or ion affinity chromatography. In some embodiments, the open reading frame comprises a sequence encoding an IMAC affinity tag linked in frame to the sequence encoding the retrotransposase. In some embodiments, the IMAC affinity tag is linked in frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the IMAC affinity tag by contacting the retrotransposase with a protease corresponding to the protease cleavage site. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the retrotransposase.
[0046] In some aspects, the disclosure provides a method of disrupting a locus in a cell, comprising contacting the cell with a composition comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, the cargo nucleotide sequence configured to interact with a retrotransposase; and (b) a retrotransposase, (i) the retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, (ii) the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, and (iii) the retrotransposase has at least equivalent transposition activity to a demonstrated retrotransposase in the cell. In some embodiments, the transposition activity is measured in vitro by introducing the retrotransposase into a cell comprising the target nucleic acid locus and detecting transposition of the target nucleic acid locus in the cell. In some embodiments, the composition comprises 20 pmoles or less of retrotransposase. In some embodiments, the composition comprises 1 pmol or less of retrotransposase.
[0047] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding any of the proteins described herein. In some embodiments, the host cell is an E. coli cell or a mammalian cell. In some embodiments, the host cell is an E. coli cell, the E. coli cell is a lambda DE3 lysogen, or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a raP ... BADIn some embodiments, the open reading frame is operably linked to a nucleotide sequence encoding an affinity tag linked in-frame to the protein-coding sequence. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a strep tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the protein-coding sequence via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of the host cell.
[0048] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.
[0049] In some aspects, the disclosure provides a method of producing any of the proteins described herein, comprising culturing any of the host cells described herein encoding any of the proteins described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the protein. In some embodiments, the induction of expression of the nuclease is by the addition of an additional chemical agent or an increased amount of a nutrient, or by an increased or decreased temperature. In some embodiments, the additional chemical agent or increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract comprising the protein. In some embodiments, the method further comprises isolating the protein. In some embodiments, the isolating comprises subjecting the protein extract to IMAC, ion exchange chromatography, anion exchange chromatography, or cation exchange chromatography. In some embodiments, the host cell comprises a nucleic acid comprising an open reading frame comprising a sequence encoding an affinity tag linked in frame to a sequence encoding the protein. In some embodiments, the affinity tag is linked in frame to the protein-encoding sequence via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the affinity tag by contacting the protein with a protease corresponding to the protease cleavage site. In some embodiments, the affinity tag is an IMAC affinity tag. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the protein.
[0050] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other and different embodiments, and its several details can be modified in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
[0051] Incorporation by Reference All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. [Brief description of the drawings]
[0052] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings.
[0053] [Figure 1] Genomic context of bacterial retrotransposons is shown. MG140-1 is a predicted retrotransposase (arrow) that encodes a Zn finger DNA binding domain and a reverse transcriptase domain. Regions flanking the retrotransposase display secondary structures that may represent binding sites for the retrotransposase (secondary structure box and zoomed image). Regions of similarity with other homologs indicate putative target sites where the retrotransposon has integrated. [Figure 2A] Figure 2 shows a multiple sequence alignment (MSA) of MG retrotransposase protein sequences of the MG140 family. Figure 2A shows the MSA of the reverse transcriptase domain. The conserved catalytic residues D, QG, [Y / F]ADD, and LG are highlighted on the consensus sequence. [Figure 2B] Figure 2A shows a multiple sequence alignment (MSA) of MG retrotransposase protein sequences of the MG140 family. Figure 2B shows the MSA of the Zn finger domain and the endonuclease domain. The Zn finger motif (CX[2-3]C) and nuclease catalytic residues that are part of the endonuclease domain are highlighted on the consensus sequence. [Diagram 3] Phylogenetic gene tree of MG genes and reference retrotransposase genes. Figure 3A shows that microbial MG retrotransposases (black branch on clade 4) are more closely related to eukaryotes than viral retrotransposases (gray branch on clade 6). Clade 1: telomerase reverse transcriptase, class 2: group II intron reverse transcriptase, class 3: eukaryotic R1-type retrotransposase, clade 4: microbial and eukaryotic R2 retrotransposase, clade 5: eukaryotic retrovirus-associated reverse transcriptase, and class 6: viral reverse transcriptase. Figure 3B shows clades 3 and 4 from the phylogenetic gene tree from Figure 3A. Some microbial MG retrotransposases contain multiple Zn finger motifs (vertical rectangles), a conserved RVT_1 reverse transcriptase domain, and APE / RLE or other endonuclease domains) (top and bottom panels. Some microbial MG retrotransposases lack the endonuclease domain (middle panel). [Figure 4] Phylogenetic tree inferred from multiple sequence alignment of reverse transcriptase domains from diverse enzymes. RT sequences are derived from DNA as well as RNA ensembles. A reference RT was included in the tree for classification purposes. [Figure 5A] Phylogenetic tree inferred from multiple sequence alignments of RT domains identified from a novel family of non-LTR retrotransposases (MG140, MG146, and MG147) and a related RT (MG148). [Figure 5B]We present data demonstrating that while non-LTR retrotransposases (MG140, MG146, and MG147) contain an RT domain, an endonuclease domain (Endo), and multiple zinc-binding ribbon motifs, the MG148 RT family lacks the endonuclease domain. [Figure 6A] We present data demonstrating that the MG140 R2 retrotransposase contains RT and endonuclease (EN) domains, as well as multiple zinc fingers, and shares 24%-26% average amino acid identity (AAI) with the reference Danio rerio R2 retrotransposase (R2Dr). [Figure 6B] Figure 1 shows data demonstrating that the MG140-47 R2 retrotransposon is integrated into the 28S rRNA gene. Alignment of the MG140-47 contig to the reference (GQ398061) ribosomal RNA operon shows a large gap in the reference 28S rDNA gene due to the integration of the R2 element (dotted box) into the MG140-47 28S rDNA gene. [Figure 7A] The genomic context of the MG145-45 retrotransposon is shown. The enzyme contains an RT domain and a zinc finger domain. A partial 18S rDNA gene at the 5' end and a polyA tail at the 3' end likely delineate the boundaries of the transposon. [Figure 7B] An alignment of the MG140-3, MG140-8, and MG140-45 genomic sequences is shown, indicating conservation of the 18S rRNA gene at position 200 of the alignment, and integration of the R2 element into the 18S rDNA gene (arrow). [Figure 8A] 1 shows a contig encoding the MG146-1 retrotransposase having the RT domain and the endonuclease domain. [Figure 8B]Shown is the MG140-17-R2 retrotransposon, which encodes three genes predicted to be involved in mobilization: an RNA recognition motif gene (RRM), an endonuclease enzyme, and a reverse transcriptase with RT and RNAse H domains. [Figure 9A] The genomic context of two members of the MG148 family of RTs is shown, with predicted genes not associated with RTs shown as white arrows. [Figure 9B] Nucleotide sequence alignment of five members of the MG148 family is shown (annotated arrow above the consensus sequence) showing the conserved region upstream of the RT (box below the sequence). [Figure 10] Screening of the in vitro activity of the RTns family of enzymes by qPCR (MG140). Activity was detected by qPCR using primers amplifying the full-length cDNA product derived from a primer extension reaction containing the respective RT. Samples are from RT reactions containing 100 nM of substrate. Negative control: water control without template in PURExpress reactions, Positive control 1: R2Tg (Taeniopygia guttata), Positive control 2: R2Bm (Bombyx mori). The two positive controls are verified R2 retrotransposons. Active candidates, defined as at least 10-fold higher signal than the negative controls, are marked in dark grey, candidates inactive in these conditions are in light grey. [Figure 11] Screening of the in vitro activity of the RTns family of enzymes by qPCR (MG146, MG147, MG148). Activity was detected by qPCR using primers amplifying full-length cDNA products derived from primer extension reactions containing the respective RT. Samples are from RT reactions containing 100 nM substrate. Negative control: water control without template in PURExpress reactions, Positive control 1: R2Tg (Taeniopygia guttata), a verified R2 retrotransposon. Active candidates, defined as at least 10-fold higher signal than the negative control, are marked in dark grey, candidates inactive in these conditions are in light grey. [Figure 12] Assay for assessing the fidelity of R2 and R2-like candidates by next generation sequencing. cDNA products obtained from primer extension reactions were PCR amplified and library prepared for NGS. Trimmed reads were aligned to the reference sequence and misincorporation frequency was calculated. Background: water control without template in PURExpress reactions, Positive control 1: R2Tg (Taeniopygia guttata). [Figure 13A] Phylogenetic tree inferred from multiple sequence alignments of full-length group II intron RTs identified from novel families from diverse classes. [Figure 13B] A summary table of the MG families of group II introns is shown. AAI: average pairwise amino acid identity of the MG family to the reference group II intron sequence. [Figure 14A] Screening of GII intron class C candidates MG153-1 through MG153-21 and MG153-25 through MG153-27 for in vitro activity by primer extension assay. For Figures 14A through 14C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - MarathonRT control RT. Bold numbering corresponds to gel lanes with active novel candidates. Results are representative of two independent experiments. Lane numbers 5 through 14 in Figure 14A correspond to novel candidates MG153-1 through MG153-10. Arrows in Figures 14A through 14C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 14B]Screening of GII intron class C candidates MG153-1 through MG153-21 and MG153-25 through MG153-27 for in vitro activity by primer extension assay. For Figures 14A through 14C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - MarathonRT control RT. Bold numbering corresponds to gel lanes with active novel candidates. Results are representative of two independent experiments. Lane numbers 5 through 14 in Figure 14B correspond to novel candidates MG153-11 through MG153-20. Arrows in Figures 14A through 14C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 14C] Screening of GII intron class C candidates MG153-1 to MG153-21 and MG153-25 to MG153-27 for in vitro activity by primer extension assay. For Figures 14A-C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - MarathonRT control RT. Bold numbering corresponds to gel lanes with active novel candidates. Results are representative of two independent experiments. Lane numbers 5-8 in Figure 14C correspond to novel candidates MG153-21, MG153-25, MG153-26, and MG153-27, respectively. Arrows in Figures 14A-C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 14D] Figure 14D shows the screening of in vitro activity of GII intron class C candidates MG153-1 to MG153-21 and MG153-25 to MG153-27 by primer extension assay. Figure 14D shows detection of full-length cDNA production by qPCR. Dark grey bars correspond to RTs generating products at least 10-fold above background. Results were determined from two technical replicates. [Figure 15A]Screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 for in vitro activity by primer extension assay. For Figures 15A to 15C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4 to 13 in Figure 15A correspond to novel candidates MG153-28 to MG153-37. Arrows in Figures 15A to 15C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 15B] Screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 for in vitro activity by primer extension assay. For Figures 15A to 15C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4 to 13 in Figure 15B correspond to novel candidates MG153-39 to MG153-48. Arrows in Figures 15A to 15C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 15C] Screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 for in vitro activity by primer extension assay. For Figures 15A to 15C, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4 to 13 in Figure 15C correspond to novel candidates MG153-49 to MG153-57. Arrows in Figures 15A to 15C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 15D]Figure 15D shows the screening of in vitro activity of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 by primer extension assay. Figure 15D shows detection of full-length cDNA production by qPCR. Dark grey bars correspond to RTs generating products at least 10-fold above background. Results were determined from two technical replicates. [Figure 16A] Screening of the GII intron class D MG165 family of reverse transcriptases for in vitro activity by primer extension assay. For Figure 16A, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4-12 - novel candidates MG165-1-9. Bold numbering corresponds to gel lanes with active novel candidates. Arrows in Figure 16A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 16B] Figure 16B shows the in vitro activity screening of the GII intron class D MG165 family of reverse transcriptases by primer extension assay. Figure 16B shows the quantification of full-length cDNA production by qPCR. Dark grey bars correspond to RTs that generate products at least 10-fold higher than background. Results were determined from two technical replicates. [Figure 17A] Screening of the in vitro activity of the GII intron class F MG167 family of reverse transcriptases by primer extension assay. For Figure 17A, lane numbers correspond to the following: 1 - PURExpress no template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - novel candidate MG167-1 to 8. Bold numbering corresponds to gel lanes with active novel candidates. Arrows in Figure 17A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 17B]Figure 17B shows the in vitro activity screening of the GII intron class F MG167 family of reverse transcriptases by primer extension assay. Figure 17B shows the quantification of full-length cDNA production by qPCR. Dark grey bars correspond to RTs that generate products at least 10-fold higher than background. Results were determined from two technical replicates. [Figure 18] Assay for assessing the fidelity of GII intron class C RT candidates from the MG153 family by next-generation sequencing. cDNA products obtained from primer extension reactions were PCR amplified and library-prepared for NGS. Trimmed reads were aligned to the reference sequence and the frequency of misincorporation was calculated. Results were determined from two independent experiments. [Figure 19A] Screening to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells is shown. Figure 19A shows detection of 542 bp (top) and 100 bp (bottom) PCR products by agarose gel analysis. Lanes not relevant to the experiments described in Figures 19A and 19B are covered with black boxes. [Figure 19B] Screening to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells is shown. Figure 19B shows detection of 542 bp (top) and 100 bp (bottom) PCR products by D1000 TapeStation. Lanes not relevant to the experiments described in Figures 19A and 19B are covered with black boxes. [Figure 19C] Screening to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells is shown. Figure 19C shows detection of a 542 bp PCR product by D1000 TapeStation for additional candidates. [Figure 20A] Phylogenetic tree of full-length G2L4-like RTs. The reference G2L4 sequence and the MG172 candidate (dots) are highlighted. [Figure 20B]Refer to columns 277-280 and show data demonstrating that MG172 RT represents catalytic residues involved in reverse transcriptase function. [Figure 21A] Phylogenetic tree of full-length LTR RTs. Reference LTR RT sequences and MG151 candidates (dots) are highlighted. [Figure 21B] The genomic context of MG151-82 RT (labeled ORF7) is shown, with predicted domains shown as dark boxes and long terminal repeats (LTRs) indicated as arrows flanking the LTR transposon. [Figure 21C] 3D structure prediction of MG151-82 showing the protease domain, RT domain, RNAse H domain, and integrase domain. [Figure 22] A multiple sequence alignment of the full-length pol protein sequence is shown to highlight the protease, RT-RNAse H, and integrase domains. The catalytic residues of the RT, RNAse H, and integrase domains of MMLV RT are indicated by bars under each domain. The protease domain of the MMLV reference sequence is not shown in the alignment. [Figure 23A] Screening of viral candidates MG151-80 to MG151-97 for in vitro activity by primer extension assay. For Figure 23A, lane numbers correspond to: 1 - RNA template annealed to primer, 2 - MMLV control RT, 3 - Ty3 control RT, 4-9 - novel candidates MG151-80 to 85, 10 - RT control. Arrows in Figures 23A-C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 23B]Screening of viral candidates MG151-80 to MG151-97 for in vitro activity by primer extension assay. For Figure 23B, lane numbers correspond to: 1 - RNA template annealed to primer, 2 to 12 - novel candidates MG151-87 to 97, 13 - MMLV control RT. Arrows in Figures 23A to 23C indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 23C] Figure 23C shows the screening of the in vitro activity of virus candidates MG151-80 to MG151-97 by primer extension assay.Figure 23D shows the testing of the in vitro activity of Ty3 control RT in different buffer conditions. Lane numbers correspond to: 1 - PURExpress no template control, 2 - Buffer A (40 mM Tris-HCl pH 7.5, 0.2 M NaCl, 10 mM MgCl2, 1 mM TCEP), 3 - Buffer B (20 mM Tris pH 7.5, 150 mM KCl, 5 mM MgCl2, 1 mM TCEP, 2% PEG-8000), 4 - Buffer C (10 mm Tris-HCl pH 7.5, 80 mm NaCl, 9 mm MgCl2, 1 mM TCEP, 0.01% (v / v) Triton X-100), 5 - Buffer D (10 mM Tris pH 7.5, 130 mM NaCl, 9 mM MgCl2, 1 mM TCEP, 10% glycerol). Arrows in Figures 23A-C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 24A]Testing of in vitro RT processivity and priming parameters of candidates MG151-89, MG151-92, and MG151-97 on structured RNA templates is shown. For Figures 24A and 24B, lane 1: 6, 10, and 16 nucleotide oligomarkers (arrows), lane 2: 8, 13, and 20 nucleotide oligomarkers, lane 3: 43 and 55 nucleotide oligomarkers, lane 4 and 10: 6 nucleotide primer, lane 5 and 11: 8 nucleotide primer, lane 6 and 12: 10 nucleotide primer, lane 7 and 13: 13 nucleotide primer, lane 8 and 14: 16 nucleotide primer, lane 9 and 15: 20 nucleotide primer. Lanes 4-9 in Figure 24A correspond to reverse transcription reactions containing MMLV with various primer lengths. MMLV reverse transcribes through structured RNA hairpins. Lanes 10-15 correspond to reverse transcription reactions containing MG151-89 with various primer lengths. MG151-89 prefers primer lengths of 16 and 20 nucleotides and appears to terminate reverse transcription at structured RNA hairpins. [Figure 24B] Testing of in vitro RT processivity and priming parameters of candidates MG151-89, MG151-92, and MG151-97 on structured RNA templates is shown. For Figures 24A and 24B, lane 1: 6, 10, and 16 nucleotide oligomarkers (arrows), lane 2: 8, 13, and 20 nucleotide oligomarkers, lane 3: 43 and 55 nucleotide oligomarkers, lane 4 and 10: 6 nucleotide primer, lane 5 and 11: 8 nucleotide primer, lane 6 and 12: 10 nucleotide primer, lane 7 and 13: 13 nucleotide primer, lane 8 and 14: 16 nucleotide primer, lane 9 and 15: 20 nucleotide primer. Lanes 4-9 in Figure 24B correspond to reverse transcription reactions containing MG151-92 with various primer lengths. Lanes 10-15 correspond to reverse transcription reactions containing MG151-97 with various primer lengths. Neither MG151-92 nor MG151-97 appear to be active under these experimental conditions. [Diagram 25]Phylogenetic analysis of 2407 retron RTs is shown, and initial candidates selected for downstream characterization in vitro are highlighted. Nine of the 16 experimentally validated retrons in the literature were added and highlighted in the tree. Grey stars represent candidate MG154-MG159 and MG173 family members. [Figure 26] Protein alignment of several retron-RT candidates selected for downstream characterization in vitro. The retron-specific motifs and the catalytic XXDD core common to all demonstrated reverse transcriptases are indicated on the figure. [Figure 27] Figure 27A shows the genomic context of the MG157-1 retron (RT labeled with an arrow on a thick black line). The retron non-coding RNA (ncRNA) is highlighted with a dotted box. Figure 27B shows an inset showing the MG157-1 retron ncRNA with its adjacent inverted repeats. Figure 27C shows the predicted structure of the MG157-1 retron ncRNA. [Figure 28A] The genomic context of the MG160-3 retron-like single domain RT is shown. The region upstream from the RT (dotted box) is conserved across MG160 members. [Figure 28B] 3D structure prediction of MG160-3 showing the RT domain aligned to the group II intron cryo-EM structure. [Figure 28C] The predicted structures of the 5'UTRs of five MG160 members are shown. [Figure 29A] Screening of retron-like candidates MG160-1 through MG160-6 and MG160-8 for in vitro activity by primer extension assay. Lane numbers in FIG. 29A correspond to the following samples: 1-PURExpress no template control, 2-MMLV control RT, 3-TGIRT-III control RT, 4-10-novel candidates MG160-1 through MG160-6 and MG160-8. Bold numbering corresponds to gel lanes with active novel candidates. Arrows in FIG. 29A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 29B] Figure 29B shows the screening of retron-like candidates MG160-1 to MG160-6 and MG160-8 for in vitro activity by primer extension assay. Figure 29B shows the quantification of full-length cDNA production by qPCR. Dark grey bars correspond to RTs generating products at least 10-fold above background. Results were determined from two technical replicates. [Figure 30A] Figure 30A shows the cell-free expression of retron RT candidates and the generation of retron ncRNA by in vitro transcription. Figure 30A shows the confirmation of retro RT protein production in cell-free expression system. Lanes correspond to: 1: ladder, 2: no template control, 3: MG156-1 (39 kDa), 4: MG156-2 (40 kDa), 5: MG157-1 (38 kDa). [Figure 30B] Figure 30B shows the cell-free expression of retron RT candidates and the generation of retron ncRNA by in vitro transcription. Figure 30B shows the confirmation of retro RT protein production in the cell-free expression system. Lanes correspond to: 1: ladder, 2: no template control, 3: MG157-2 (37 kDa), 4: MG157-5 (43 kDa), 5: MG159-1 (53 kDa), 6: Ec86 (38 kDa, positive control retron RT). [Figure 30C] Figure 30C shows the cell-free expression of retron RT candidates and the generation of retron ncRNA by in vitro transcription. Figure 30C shows the generation of retron ncRNA templates by in vitro transcription. Lanes correspond to the following ncRNAs corresponding to the following retrons-1: MG154-1, 2: MG154-2, 3: MG155-1, 4: MG155-2, 5: MG155-3, 6: MG156-1, 7: MG156-2, 8: MG157-1, 9: MG157-2, 10: MG157-5, 11: MG158-1, 12: MG159-1, 13: Ec86, 14: MG155-4, 15: MG173-1, 16: MG155-5. [Diagram 31]MG140-1 shows the domain architecture demonstrating that the R2 retrotransposon integrates into the 28S rRNA gene. The R2 retrotransposase (light grey arrow) contains multiple Zn fingers, as well as the RT and endonuclease domains. MG140-1 is flanked by the 5'UTR and 3'UTR that define the transposon boundaries. MG140-1 integrates precisely between the G and T nucleotides in the target site motif GGTAGC. [Diagram 32] RT activity by primer extension using DNA oligos containing phosphorothioate linkage modifications is shown. Lane numbers correspond to the following: 1: PURExpress no template control with PS modified primer 1, 2: PURExpress no template control with PS modified primer 2, 3: PURExpress no template control with PS modified primer 3, 4: MMLV RT with unmodified primer, 5: MMLV RT with PS modified primer 1, 6: MMLV RT with PS modified primer 2, 7: MMLV RT with PS modified primer 3, 8: TGIRT-III with unmodified primer, 9: TGIRT-III with PS modified primer 1, 10: TGIRT-III with PS modified primer 2, 11: TGIRT-III with PS modified primer 3, 12: MG153-9 with unmodified primer, 13: MG153-9 with PS modified primer 1, 14: MG153-9 with PS modified primer 2, 15 MG153-9 with PS modified primer 3. MMLV RT and TGIRT-III are control RTs. [Diagram 33]1 shows the screening of retron RT activity on RNA template by primer extension assay. Lane numbers correspond to the following: 1: PURExpress no template control, 2: MMLV control RT, 3: MG154-1, 4: MG155-1, 5: MG155-2, 6: MG155-3, 7: MG156-2, 8: MG157-1, 9: MG157-2, 10: MG157-5, 11: MG158-1, 12: MG159-1, 13: Ec86 control retron RT, 14: Sa163 control retron RT, 15: St85 control retron RT. Bold lanes correspond to novel retron RTs that show primer extension activity on the tested substrates. [Diagram 34] Screening of the ability of MG153 GII-derived RT to synthesize cDNA in mammalian cells. Detection of a 542 bp cDNA synthesis PCR product was assayed by Taqman qPCR. cDNA activity was normalized to an active TGIRT control, where TGIRT represents a value of 1. The Y-axis is shown in log10 scale. [Figure 35A] Protein expression of MG153 GII-derived RT by immunoblot. Figures 35A and 35B: Cells were transfected with plasmids containing candidate RTs, and protein expression was assessed by immunoblot to detect HA peptide fused to the N-terminus of the RT. All lanes were normalized to total protein concentration. The white arrow indicates a band at 2X the expected molecular size of the protein, indicating a protein dimer. Lanes not relevant to the experiment described in Figures 35A and 35B are covered with black boxes. [Figure 35B] Protein expression of MG153 GII-derived RT by immunoblot. Figures 35A and 35B: Cells were transfected with plasmids containing candidate RTs, and protein expression was assessed by immunoblot to detect HA peptide fused to the N-terminus of the RT. All lanes were normalized to total protein concentration. The white arrow indicates a band at 2X the expected molecular size of the protein, indicating a protein dimer. Lanes not relevant to the experiment described in Figures 35A and 35B are covered with black boxes. [Figure 35C]Figure 35C: Multiple sequence alignment of GII-derived RT. The region shown corresponds to positions 196 to 201 of the alignment. The dimerization motif CAQQ is highlighted. [Diagram 36] Relative activity of GII-derived RT normalized to protein expression is shown. cDNA synthesis was detected by Taqman qPCR and protein expression was detected by immunoblot. Activity against TGIRT was normalized to total protein concentration. Y-axis is shown on a linear scale.
[0054] Brief Description of the Sequence Listing The Sequence Listing submitted herewith provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0055] MG140 SEQ ID NOs: 1 to 29 and 393 to 401 show the full-length peptide sequences of the MG140 translocation protein.
[0056] SEQ ID NOs: 374 to 386 show the nucleotide sequences of genes encoding HA-His-tagged MG140 reverse transcriptase proteins.
[0057] SEQ ID NOs: 761 to 798 show the nucleotide sequences of MG140 UTRs.
[0058] SEQ ID NOs: 799 to 894 show the full-length peptide sequences of the MG140 reverse transcriptase protein.
[0059] MG146 SEQ ID NOs: 402 and 895 show the full-length peptide sequence of the MG140 translocation protein.
[0060] SEQ ID NO: 387 shows the nucleotide sequence of the gene encoding the HA-His tagged MG146 reverse transcriptase protein.
[0061] MG147 SEQ ID NO: 388 shows the nucleotide sequence of the gene encoding HA-His tagged MG147 reverse transcriptase protein.
[0062] MG148 SEQ ID NOs: 403 to 426 show the full-length peptide sequences of MG148 reverse transcriptase protein.
[0063] SEQ ID NOs: 389 to 392 show the nucleotide sequences of genes encoding HA-His-tagged MG148 reverse transcriptase proteins.
[0064] MG149 SEQ ID NOs: 427 to 439 show the full-length peptide sequences of the MG149 reverse transcriptase protein.
[0065] MG151 SEQ ID NOs: 440 to 554 show the full-length peptide sequences of the MG151 reverse transcriptase protein.
[0066] SEQ ID NOs: 356 to 362 show the nucleotide sequences of genes encoding TwinStrep-tagged MG151 reverse transcriptase proteins.
[0067] SEQ ID NOs: 363 to 373 show the nucleotide sequences of genes encoding strep-tagged MG151 reverse transcriptase proteins.
[0068] MG153 SEQ ID NOs: 555 to 608 show the full-length peptide sequences of the MG153 reverse transcriptase protein.
[0069] SEQ ID NOs: 30 to 32 and 40 to 50 show the nucleotide sequences of fusion proteins containing the MG153 reverse transcriptase protein and the MS2 coat protein (MCP).
[0070] SEQ ID NOs: 66 to 119 show the nucleotide sequences of genes encoding strep-tagged MG153 reverse transcriptase proteins.
[0071] SEQ ID NOs: 120 to 173 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG153 reverse transcriptase protein.
[0072] SEQ ID NOs: 740 to 756 show the nucleotide sequences of genes encoding MCP-tagged MG153 reverse transcriptase proteins.
[0073] MG154 SEQ ID NOs: 609-610 show the full-length peptide sequences of the MG154 reverse transcriptase protein.
[0074] SEQ ID NOs: 308 to 309 show the nucleotide sequences of genes encoding strep-tagged MG154 reverse transcriptase proteins.
[0075] SEQ ID NOs: 324-325 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG154 reverse transcriptase protein.
[0076] SEQ ID NOs: 340 to 341 show the nucleotide sequence of an ncRNA compatible with MG154 nuclease.
[0077] MG155 SEQ ID NOs: 611 to 615 show the full-length peptide sequences of the MG155 reverse transcriptase protein.
[0078] SEQ ID NOs: 310 to 312 show the nucleotide sequences of genes encoding strep-tagged MG155 reverse transcriptase proteins.
[0079] SEQ ID NOs: 326-328 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG155 reverse transcriptase protein.
[0080] SEQ ID NOs: 342 to 344 show the nucleotide sequences of ncRNAs compatible with MG155 nuclease.
[0081] MG156 SEQ ID NOs: 616-617 show the full-length peptide sequences of MG156 reverse transcriptase protein.
[0082] SEQ ID NOs: 313 to 314 show the nucleotide sequences of genes encoding strep-tagged MG156 reverse transcriptase proteins.
[0083] SEQ ID NOs: 329-330 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG156 reverse transcriptase protein.
[0084] SEQ ID NOs: 345 to 346 show the nucleotide sequence of an ncRNA compatible with MG156 nuclease.
[0085] MG157 SEQ ID NOs: 618 to 622 show the full-length peptide sequences of MG157 reverse transcriptase protein.
[0086] SEQ ID NOs: 315 to 319 show the nucleotide sequences of genes encoding strep-tagged MG157 reverse transcriptase proteins.
[0087] SEQ ID NOs: 331 to 335 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG157 reverse transcriptase protein.
[0088] SEQ ID NOs: 347 to 351 show the nucleotide sequences of ncRNAs compatible with MG157 nuclease.
[0089] MG158 SEQ ID NO: 623 shows the full-length peptide sequence of the MG158 reverse transcriptase protein.
[0090] SEQ ID NO: 320 shows the nucleotide sequence of the gene encoding a strep-tagged MG158 reverse transcriptase protein.
[0091] SEQ ID NO: 336 shows the nucleotide sequence of the E. coli codon-optimized gene encoding the MG158 reverse transcriptase protein.
[0092] SEQ ID NO: 352 shows the nucleotide sequence of an ncRNA compatible with MG158 nuclease.
[0093] MG159 SEQ ID NOs: 624 to 626 show the full-length peptide sequences of the MG159 reverse transcriptase protein.
[0094] SEQ ID NOs: 321 to 323 show the nucleotide sequences of genes encoding strep-tagged MG159 reverse transcriptase proteins.
[0095] SEQ ID NOs: 337-339 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG159 reverse transcriptase protein.
[0096] SEQ ID NOs: 353 to 355 show the nucleotide sequence of an ncRNA compatible with MG159 nuclease.
[0097] MG160 SEQ ID NOs: 627 to 673 show the full-length peptide sequences of the MG160 reverse transcriptase protein.
[0098] SEQ ID NOs: 174 to 180 show the nucleotide sequences of genes encoding strep-tagged MG160 reverse transcriptase proteins.
[0099] SEQ ID NOs: 181-187 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG160 reverse transcriptase protein.
[0100] MG163 SEQ ID NOs: 674 to 678 show the full-length peptide sequences of the MG163 reverse transcriptase protein.
[0101] SEQ ID NOs: 188 to 192 show the nucleotide sequences of genes encoding strep-tagged MG163 reverse transcriptase proteins.
[0102] SEQ ID NOs: 193-197 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG163 reverse transcriptase protein.
[0103] MG164 SEQ ID NOs: 679 to 683 show the full-length peptide sequences of the MG164 reverse transcriptase protein.
[0104] SEQ ID NOs: 198 to 202 show the nucleotide sequences of genes encoding strep-tagged MG164 reverse transcriptase proteins.
[0105] SEQ ID NOs: 203-207 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG164 reverse transcriptase protein.
[0106] MG165 SEQ ID NOs: 684 to 692 show the full-length peptide sequences of MG165 reverse transcriptase protein.
[0107] SEQ ID NOs: 208 to 216 show the nucleotide sequences of genes encoding strep-tagged MG165 reverse transcriptase proteins.
[0108] SEQ ID NOs: 217-225 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG165 reverse transcriptase protein.
[0109] SEQ ID NOs: 757 to 759 show the nucleotide sequences of genes encoding MCP-tagged MG165 reverse transcriptase proteins.
[0110] MG166 SEQ ID NOs: 693 to 697 show the full-length peptide sequences of MG166 reverse transcriptase protein.
[0111] SEQ ID NOs: 226 to 230 show the nucleotide sequences of genes encoding strep-tagged MG166 reverse transcriptase proteins.
[0112] SEQ ID NOs: 231-235 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG166 reverse transcriptase protein.
[0113] MG167 SEQ ID NOs: 698 to 702 show the full-length peptide sequences of the MG167 reverse transcriptase protein.
[0114] SEQ ID NOs: 236 to 240 show the nucleotide sequences of genes encoding strep-tagged MG167 reverse transcriptase proteins.
[0115] SEQ ID NOs: 241-245 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG167 reverse transcriptase protein.
[0116] SEQ ID NOs: 759 to 760 show the nucleotide sequences of genes encoding MCP-tagged MG167 reverse transcriptase proteins.
[0117] MG168 SEQ ID NOs: 703 to 707 show the full-length peptide sequences of MG168 reverse transcriptase protein.
[0118] SEQ ID NOs: 246 to 250 show the nucleotide sequences of genes encoding strep-tagged MG168 reverse transcriptase proteins.
[0119] SEQ ID NOs: 251-255 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG168 reverse transcriptase protein.
[0120] MG169 SEQ ID NOs: 708 to 718 show the full-length peptide sequences of the MG169 reverse transcriptase protein.
[0121] SEQ ID NOs: 256 to 266 show the nucleotide sequences of genes encoding strep-tagged MG169 reverse transcriptase proteins.
[0122] SEQ ID NOs: 267-277 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG169 reverse transcriptase protein.
[0123] MG170 SEQ ID NOs: 719 to 728 show the full-length peptide sequences of the MG170 reverse transcriptase protein.
[0124] SEQ ID NOs: 278 to 287 show the nucleotide sequences of genes encoding strep-tagged MG170 reverse transcriptase proteins.
[0125] SEQ ID NOs: 288-297 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG170 reverse transcriptase protein.
[0126] MG172 SEQ ID NOs: 729 to 733 show the full-length peptide sequences of the MG172 reverse transcriptase protein.
[0127] SEQ ID NOs: 298 to 302 show the nucleotide sequences of genes encoding strep-tagged MG172 reverse transcriptase proteins.
[0128] SEQ ID NOs: 303-307 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG172 reverse transcriptase protein.
[0129] MG173 SEQ ID NOs: 734-735 show the full-length peptide sequences of the MG173 reverse transcriptase protein.
[0130] Other Arrays SEQ ID NOs: 736 to 738 show the nucleotide sequences of phosphorothioate-modified primers.
[0131] SEQ ID NO: 739 shows the nucleotide sequence of a Taqman probe for qPCR. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0132] While various embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be used.
[0133] The practice of some of the methods disclosed herein may involve immunological, biochemical, chemical, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques, unless otherwise indicated. See, for example, Sambrook and Green, Molecular Cloning: A Laboratory Manual, 4th Edition (2012); the series Current Protocols in Molecular Biology (FMA Usubel, et al. eds.); the series Methods In Enzymology (Academic Press, Inc.), PCR2: A Practical Approach (MJ MacPherson, BD Hames and GR Taylor eds. (1995)), Harlow and Lane, eds. (1988) Antibodies, A Laboratory Manual, and Culture of Animal Cells: A Manual of Basic Technique and Specialized Applications, 6th Edition (RI Freshney, ed. (2010)) (incorporated herein in its entirety by reference).
[0134] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. Furthermore, to the extent the terms "comprising," "including," "having," "having," "having," or variations thereof are used in either the detailed description and / or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0135] The term "about" or "approximately" means within an acceptable error range for a particular value as determined by one of ordinary skill in the art, which depends in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, "about" can mean within one or more standard deviations, as is customary in the art. Alternatively, "about" can mean within a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0136] As used herein, "cell" generally refers to a biological cell. A cell may be the basic structural, functional, or biological unit of a living organism. A cell may originate from any organism having one or more cells. Some non-limiting examples include prokaryotic cells, eukaryotic cells, bacterial cells, archaeal cells, single-cell eukaryotic cells, protozoan cells, cells from plants (e.g., plant crops, fruits, vegetables, grains, soybeans, corn, maize, wheat, seeds, tomatoes, rice, cassava, sugarcane, pumpkins, hay, potatoes, cotton, cannabis, tobacco, flowering plants, conifers, gymnosperms, ferns, club mosses, hornworts, bryophytes, mosses), algae cells (e.g., Botryococcus braunii, Chlamydomonas reinhardtii, Nannochloropsis gaditana, Chlorella pyrenoidosa, Sargassum patens, etc.), and cells from other organisms (e.g., cereals, vegetables, fruits, and vegetables). C. Agardh, etc.), seaweed (e.g., kelp), fungal cells (e.g., yeast cells, cells from mushrooms), animal cells, cells from vertebrates (e.g., fruit flies, cnidarians, echinoderms, nematodes, etc.), cells from vertebrates (e.g., fish, amphibians, reptiles, birds, mammals), cells from mammals (e.g., pigs, cows, goats, sheep, rodents, rats, mice, non-human primates, humans, etc.), etc. In some cases, the cells are not derived from a naturally occurring organism (e.g., the cells may be synthetically produced and sometimes referred to as artificial cells).
[0137] As used herein, the term "nucleotide" generally refers to a base-sugar-phosphate combination. A nucleotide may include synthetic nucleotides. A nucleotide may include synthetic nucleotide analogs. A nucleotide may be a monomeric unit of a nucleic acid sequence (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide may include ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP) and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives may include, for example, [αS]dATP, 7-deaza-dGTP and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used herein, the term nucleotide may refer to dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Examples of dideoxyribonucleoside triphosphates include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP, and ddTTP. Nucleotides may be unlabeled or detectably labeled, such as by using a moiety that includes an optically detectable moiety (e.g., a fluorophore). Labeling may also be performed using quantum dots. Detectable labels may include, for example, radioisotopes, fluorescent labels, chemiluminescent labels, bioluminescent labels, and enzyme labels. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2′7′-dimethoxy-4′5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N′,N′-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4′dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2′-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer, Foster City, Calif.; fluoro-conjugated deoxynucleotides, fluoro-conjugated Cy3-dCTP, fluoro-conjugated Cy5-dCTP, fluoro-conjugated fluoroX-dCTP, fluoro-conjugated Cy3-dUTP, and fluoro-conjugated Cy5-dUTP available from Amersham, Arlington Heights, Ill.; Fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2′-dATP available from Mannheim, Indianapolis, Ind.; and Molecular Examples of chromosomal labeling nucleotides available from Probes, Eugene, Oreg. include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. Nucleotides may also be labeled or marked by chemical modification. The chemically modified single nucleotide may be a biotin-dNTP.Some non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0138] The terms "polynucleotide", "oligonucleotide", and "nucleic acid" are generally used interchangeably to refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-stranded, double-stranded, or multiple-stranded form. A polynucleotide may be exogenous or endogenous to a cell. A polynucleotide may be present in a cell-free environment. A polynucleotide may be a gene or a fragment thereof. A polynucleotide may be DNA. A polynucleotide may be RNA. A polynucleotide may have any three-dimensional structure and may perform any function. A polynucleotide may contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). If present, modifications to the nucleotide structure may be imparted before or after assembly of the polymer. Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, heterologous nucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. Non-limiting examples of polynucleotides include coding or non-coding regions of genes or gene fragments, loci (locuses) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, isolated RNA of any sequence, cell-free polynucleotides including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. The sequence of nucleotides may be interrupted by non-nucleotide components.
[0139] The term "transfection" or "transfected" generally refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a genetic sequence encoding a complete protein or a functional portion thereof. See, e.g., Sambrook et al., 1989, Molecular Cloning: A Laboratory Manual, 18.1-18.88, which is incorporated herein by reference in its entirety.
[0140] The terms "peptide", "polypeptide" and "protein" are used interchangeably herein and generally refer to a polymer of at least two amino acid residues linked by peptide bonds. The term does not refer to a particular length of the polymer, and is not intended to imply or distinguish whether the peptide is produced using recombinant technology, chemical or enzymatic synthesis, or naturally occurring. The term applies to naturally occurring amino acid polymers as well as amino acid polymers that include at least one modified amino acid. In some embodiments, the polymer may be interrupted by non-amino acids. The term includes amino acid chains of any length, including full-length proteins and proteins with or without secondary or tertiary structure (e.g., domains). The term also encompasses amino acid polymers that have been modified by any other manipulation, such as, for example, disulfide bond formation, glycosylation, lipid formation, acetylation, phosphorylation, oxidation, and conjugation with a labeling component. As used herein, the terms "amino acid" and "amino acids" generally refer to natural and unnatural amino acids, including, but not limited to, modified amino acids and amino acid analogs. Modified amino acids may include natural amino acids and unnatural amino acids, which are chemically modified to include a non-naturally occurring group or chemical moiety on the amino acid. An amino acid analog may refer to an amino acid derivative. The term "amino acid" includes both D- and L-amino acids.
[0141] As used herein, "non-natural" may generally refer to a nucleic acid or polypeptide sequence that is not present in a naturally occurring nucleic acid or protein. Non-natural may refer to an affinity tag. Non-natural may refer to a fusion. Non-natural may refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, or deletion. The non-natural sequence may exhibit or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, an ubiquitination activity, etc.) that may also be exhibited by the nucleic acid or polypeptide sequence to which the non-natural sequence is fused. The non-natural nucleic acid or polypeptide sequence may be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence that encodes the chimeric nucleic acid or polypeptide.
[0142] As used herein, the term "promoter" generally refers to a regulatory DNA region that controls the transcription or expression of a gene and may be located adjacent to or overlapping the nucleotide or region of nucleotides at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often called transcription factors, which promote the binding of RNA polymerase to DNA, thereby resulting in gene transcription. A "basal promoter", also called a "core promoter", may generally refer to a promoter that contains all the basic elements to promote the transcriptional expression of an operably linked polynucleotide. Eukaryotic basal promoters may contain a TATA-box or CAAT box.
[0143] As used herein, the term "expression" generally refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) or by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression includes splicing of the mRNA in eukaryotic cells.
[0144] As used herein, "operably linked," "operably linked," "operably linked," or grammatical equivalents thereof generally refer to the juxtaposition of genetic elements, such as promoters, enhancers, polyadenylation sequences, and the like, where the elements are in a relationship that allows them to operate in an expected manner. For example, a regulatory element, which may include a promoter sequence or an enhancer sequence, is operably linked to a coding region if the regulatory element helps to initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region so long as this functional relationship is maintained.
[0145] As used herein, a "vector" generally refers to a polymer or an association of polymers that contains or associates with a polynucleotide and can be used to mediate delivery of the polynucleotide to a cell. Examples of vectors include plasmids, viral vectors, liposomes, and other gene delivery vehicles. A vector generally includes genetic elements, such as control elements, operably linked to a gene to facilitate expression of the gene in a target.
[0146] As used herein, "expression cassette" and "nucleic acid cassette" are generally used interchangeably to refer to a combination of nucleic acid sequences or elements that are operably linked for expression. In some embodiments, an expression cassette refers to a combination of regulatory elements and a gene or genes to which they are operably linked for expression.
[0147] A "functional fragment" of a DNA or protein sequence generally refers to a fragment that retains a biological activity (either functional or structural) substantially similar to the biological activity of the full-length DNA or protein sequence. The biological activity of a DNA sequence may be the ability to affect expression in a manner attributable to the full-length sequence.
[0148] As used herein, an "engineered" subject generally refers to a subject that has been modified by human intervention. By way of non-limiting examples, a nucleic acid may be modified by altering its sequence to a sequence that does not occur in nature, a nucleic acid may be modified by ligating to a nucleic acid with which it is not naturally associated such that the ligated product has a function not present in the original nucleic acid, an engineered nucleic acid may be synthesized in vitro with a sequence that does not occur in nature, a protein may be modified by changing its amino acid sequence to a sequence that does not occur in nature, and an engineered protein may acquire a new function or property. An "engineered" system includes at least one engineered component.
[0149] As used herein, "synthetic" and "artificial" may generally be used interchangeably to refer to proteins or domains thereof that have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains.
[0150] As used herein, the term "transposable element" refers to a DNA sequence that can move from one location to another in a genome (e.g., they can be "transposed"). Transposable elements can be broadly divided into two classes. Class I transposable elements, or "retrotransposons," are transposed via transcription and translation of an RNA intermediate, and then reintegrate into the genome at their new location via reverse transcription (a process mediated by reverse transcriptase). Class II transposable elements, or "DNA transposons," are transposed via a complex of single- or double-stranded DNA flanked on both sides by a transposase. Further characteristics of this family of enzymes can be found, for example, in Nature Education 2008,1(1),204, and Genome Biology 2018,19(199),1-12, each of which is incorporated herein by reference.
[0151] As used herein, the term "retrotransposon" refers to a class I transposable element that functions according to a bipartite "copy and paste" mechanism involving an RNA intermediate. "Retrotransposase" refers to an enzyme involved in the transposition of retrotransposons. In some embodiments, a retrotransposase comprises a reverse transcriptase domain. In some embodiments, a retrotransposase further comprises one or more zinc finger domains. In some embodiments, a retrotransposase further comprises an endonuclease domain.
[0152] The terms "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences generally refer to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a certain percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix setting parameters of a word length (W) of 3, an expectation (E) of 10, and gap costs at 11, an extension of 1, and using a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of a word length (W) of 2, an expectation (E) of 1,000,000, and PAM30 scoring setting gap costs at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (default parameters for BLASTP are available in BLAST at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW using Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE using default parameters; MAFFT using parameters of 2 retrees and 1000 maximum repeats; Novafold using default parameters; HMMER hmmalign using default parameters.
[0153] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" generally refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, e.g., as determined by the alignment producing the highest or "optimized" percent identity score.
[0154] The term "open reading frame" or "ORF" generally refers to a nucleotide sequence that can code for a protein or a portion of a protein. An open reading frame can begin with a start codon (e.g., represented in standard codes as AUG for RNA molecules and ATG for DNA molecules) and can be read with a codon triplet until the frame ends with a stop codon (e.g., represented in standard codes as UAA, UGA, or UAG for RNA molecules and TAA, TGA, or TAG for DNA molecules).
[0155] The present disclosure includes any variant of the enzymes described herein that have one or more conservative amino acid substitutions. Such conservative substitutions can be made in the amino acid sequence of a polypeptide without destroying the three-dimensional structure or function of the polypeptide. Conservative substitutions can be achieved by replacing amino acids with similar hydrophobicity, polarity, and R chain length. Additionally or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that are mutated between species without changing the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of the retrotransposase protein sequences described herein (e.g., an MG140 family retrotransposase described herein, or any other family retrotransposase described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions that do not destroy the activity of one or more important active site residues of retrotransposase. In some embodiments, a functional variant of any of the proteins described herein lacks at least one substitution of the conserved or functional residues called out in FIG. 2.In some embodiments, a functional variant of any of the proteins described herein lacks all of the substitutions of conserved or functional residues called out in FIG.
[0156] The disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein having substitutions of one or more catalytic residues to reduce or eliminate activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include disruptive substitutions of at least one, at least two, or all three catalytic residues called out in FIG. 2.
[0157] Conservative substitution tables providing functionally similar amino acids are available in a variety of references (see, for example, Creighton, Proteins: Structures and Molecular Properties (WH Freeman & Co.; 2nd edition (December 1993)). Each of the following eight groups contains amino acids that are conservative substitutions for one another: 1) Alanine (A), Glycine (G), 2) Aspartic acid (D), glutamic acid (E), 3) Asparagine (N), Glutamine (Q), 4) Arginine (R), Lysine (K), 5) Isoleucine (I), Leucine (L), Methionine (M), Valine (V), 6) Phenylalanine (F), Tyrosine (Y), Tryptophan (W), 7) Serine (S), Threonine (T), and 8) Cysteine (C), Methionine (M).
[0158] Variants of any of the nucleic acid sequences described herein having one or more substitutions, deletions, or insertions are also included in the present disclosure. In some embodiments, such variants have a sequence that has at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of the nucleic acid sequences described herein.
[0159] Some of the protein sequences described herein involve the determination of specific domains (e.g., reverse transcriptase or RT domains) from the sequence of a selected larger protein (e.g., a retrotransposase). In such cases, multiple sequence alignment (MSA) with a reference larger protein (e.g., a retrotransposase) whose domains have been verified (e.g., in a 3D structure) is used to identify domain boundaries by aligning the selected protein with the larger protein that has the verified domain. When the sequences are highly divergent, so that the MSA is inconclusive, the 3D structure of the larger protein is determined and the structural domains are compared to known domains to define the boundaries. These boundaries can be further verified by ensuring the presence of critical catalytic residues for the domain within the domain boundaries.
[0160] As used herein, the term "LINE retrotransposase" generally refers to a class of autonomous non-LTR retrotransposons (Long INterspersed Elements). As used herein, the terms "R2 retrotransposase" or "R4 retrotransposase" generally refer to a subclass of LINE retrotransposases that share a similar domain architecture but differ in that R2 retrotransposases can be site-specific (e.g., integrated at specific sites in rRNA genes), while R4 retrotransposons can integrate both in rRNA genes as well as other non-specific sites that contain repeats.
[0161] overview The discovery of new transposable elements with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technology, offering the potential to improve speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of transposable elements in microorganisms and indeed in a wide variety of microbial species, there are relatively few functionally characterized transposable elements in the literature. This is in part because the vast number of microbial species are not easily cultured under laboratory conditions. Metagenomic sequencing from natural environmental niches containing a large number of microbial species could dramatically increase the number of demonstrated new transposable elements, offering the potential to expedite the discovery of new oligonucleotide editing functions.
[0162] Transposable elements are deoxyribonucleic acid sequences that can change position in the genome, often resulting in the generation or improvement of mutations. In eukaryotes, a large proportion of the genome and a large proportion of the mass of cellular DNA are attributable to transposable elements. Although transposable elements are "selfish genes" that propagate themselves at the expense of other genes, they perform a variety of important functions and have been found to be important in genome evolution. Based on their mechanism, transposable elements are classified as either class I "retrotransposons" or class II "DNA transposons".
[0163] Class I transposable elements, also called retrotransposons, function according to a two-part "copy-and-paste" mechanism involving an RNA intermediate. First, the retrotransposon is transcribed. The resulting RNA is then converted back into DNA by reverse transcriptase (generally encoded by the retrotransposon itself), and the reverse transcribed retrotransposon is integrated into its new location in the genome by an integrase. Retrotransposons are further classified into three lineages. Retrotransposons with long terminal repeats ("LTRs") encode reverse transcriptase and are flanked by long stretches of repetitive DNA. Retrotransposons with long interspersed repeats ("LINEs") encode reverse transcriptase, lack LTRs, and are transcribed by RNA polymerase II. Retrotransposons with short interspersed repeats ("SINEs") are transcribed by RNA polymerase III, but lack reverse transcriptase and instead rely on the reverse transcription machinery of other transposable elements (e.g., LINEs).
[0164] Class II transposable elements, also called DNA transposons, function according to a mechanism that does not involve an RNA intermediate. Many DNA transposons exhibit a "cut and paste" mechanism in which a transposase binds to terminal inverted repeats ("TIRs") flanking the transposon, cleaves the transposon from the donor region, and inserts it into the target region of the genome. Others, called "helitrons," involve a single-stranded DNA intermediate and exhibit a "rolling circle" mechanism mediated by an unsubstantiated protein understood to have HUH endonuclease function and 5' to 3' helicase activity. First, a circular strand of DNA is nicked to create two single DNA strands. The protein remains attached to the 5' phosphate of the nicked strand, leaving the 3' hydroxyl end of the complementary strand exposed, thus allowing the polymerase to replicate the unnicked strand. Once replication is complete, the new strand dissociates and replicates itself along with the original template strand. Yet another DNA transposon, the "pollinton," is theorized to undergo a "self-synthesizing" mechanism. Transposition is initiated by integrase excision of single-stranded extrachromosomal Pollinator elements that form racket-like structures. Pollinator undergoes replication by DNA polymerase B, and double-stranded Pollinator is inserted into the genome by integrase. In addition, some DNA transposons, such as those of the IS200 / IS605 family, proceed via a "peel and paste" mechanism in which TnpA excises a piece of single-stranded DNA (as a circular "transposon junction") from the lagging strand template of a donor gene and reinserts it into the replication fork of the target gene.
[0165] Although transposable elements have found some use as biological tools, the transposable elements demonstrated do not encompass the full range of possible biodiversity and targeting possibilities, and may not represent all possible activities. Here, we have drawn thousands of genome fragments from multiple metagenomes for transposable elements. The diversity of transposable elements demonstrated may be expanded, and novel systems may be developed into highly targetable, compact, and precise gene editing agents.
[0166] MG enzyme In some aspects, the present disclosure provides novel retrotransposases. These candidates may represent one or more novel subtypes, and several subfamilies may be identified. These retrotransposases are less than about 1,400 amino acids in length. These retrotransposases may simplify delivery and expand therapeutic applications.
[0167] In some aspects, the present disclosure provides novel retrotransposases, which may be MG140 as described herein (see Figures 1 and 2).
[0168] In one aspect, the present disclosure provides engineered retrotransposase systems discovered through metagenomic sequencing. In some embodiments, metagenomic sequencing is performed on samples. In some embodiments, samples can be collected from various environments. Such environments can be human microbiomes, animal microbiomes, high temperature environments, low temperature environments. Such environments can include sediments.
[0169] In one aspect, the present disclosure provides an engineered retrotransposase system comprising a retrotransposase. In some embodiments, the retrotransposase is derived from an uncultured microorganism. The retrotransposase may be configured to bind to a 3' untranslated region (UTR). The retrotransposase may bind to a 5' untranslated region (UTR).
[0170] In one aspect, the disclosure provides an engineered retrotransposase system comprising a retrotransposase. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0171] In some embodiments, the retrotransposase comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase may be substantially identical to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0172] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain.
[0173] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to a documented retrotransposase.
[0174] In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0175] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a single-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0176] In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a human genomic polynucleotide sequence.
[0177] In some embodiments, the retrotransposase may include a variant having one or more nuclear localization sequences (NLS). The NLS may be proximal to the N-terminus or C-terminus of the retrotransposase. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 896-911 or a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 896. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 897.
[0178] [Table 1]
[0179] In some embodiments, sequences may be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithms, or the CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. Sequence identity may be determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment.
[0180] In one aspect, the present disclosure provides a deoxyribonucleic acid polynucleotide encoding the engineered retrotransposase system described herein.
[0181] In one aspect, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence. In some embodiments, the engineered nucleic acid sequence is optimized for expression in an organism. In some embodiments, the retrotransposase is from an uncultured microorganism. In some embodiments, the organism is not an uncultured organism.
[0182] In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0183] In some embodiments, the retrotransposase comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase may be substantially identical to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0184] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain.
[0185] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to a documented retrotransposase.
[0186] In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0187] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a single-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0188] In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a human genomic polynucleotide sequence.
[0189] In some embodiments, the retrotransposase may include a variant having one or more nuclear localization sequences (NLS). The NLS may be proximal to the N-terminus or C-terminus of the retrotransposase. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 896-911 or a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 896. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 897.
[0190] In some embodiments, the organism is a prokaryote. In some embodiments, the organism is a bacterium. In some embodiments, the organism is a eukaryote. In some embodiments, the organism is a fungus. In some embodiments, the organism is a plant. In some embodiments, the organism is a mammal. In some embodiments, the organism is a rodent. In some embodiments, the organism is a human.
[0191] In one aspect, the disclosure provides an engineered vector. In some embodiments, the engineered vector comprises a nucleic acid sequence encoding a retrotransposase. In some embodiments, the retrotransposase is derived from an uncultured microorganism.
[0192] In some embodiments, the engineered vector comprises a nucleic acid described herein. In some embodiments, the nucleic acid described herein is a deoxyribonucleic acid polynucleotide described herein. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus.
[0193] In one aspect, the disclosure provides a cell comprising the vector described herein.
[0194] In one aspect, the disclosure provides a method of producing a retrotransposase. In some embodiments, the method comprises culturing a cell.
[0195] In one aspect, the present disclosure provides a method of binding, nicking, cleaving, marking, modifying, or translocating a double-stranded deoxyribonucleic acid polynucleotide. The method may include contacting the double-stranded deoxyribonucleic acid polynucleotide with a retrotransposase. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0196] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain.
[0197] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to a documented retrotransposase.
[0198] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a single-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0199] In some embodiments, the retrotransposase is from an uncultured microorganism. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0200] In one aspect, the present disclosure provides a method for modifying a target nucleic acid locus. The method may include delivering an engineered retrotransposase system as described herein to a target nucleic acid locus. In some embodiments, the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus.
[0201] In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, marking, modifying, or rearranging the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA) or ribonucleic acid (RNA). In some embodiments, the target nucleic acid comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is in a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, or a human cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC).
[0202] In some embodiments, delivery of the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid described herein or a vector described herein. In some embodiments, delivery of the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding a retrotransposase. In some embodiments, the nucleic acid comprises a promoter. In some embodiments, the open reading frame encoding the retrotransposase is operably linked to a promoter.
[0203] In some embodiments, delivery of the engineered retrotransposase system to the target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the retrotransposase. In some embodiments, delivery of the engineered retrotransposase system to the target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, delivery of the engineered retrotransposase system to the target nucleic acid locus comprises delivering a deoxyribonucleic acid (DNA) encoding an engineered guide RNA operably linked to a ribonucleic acid (RNA) pol III promoter.
[0204] In some embodiments, the retrotransposase does not induce a cut at or proximal to the target nucleic acid locus.
[0205] In one aspect, the disclosure provides a host cell comprising an open reading frame encoding a heterologous retrotransposase. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0206] In some embodiments, the retrotransposase comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase may be substantially identical to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0207] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain.
[0208] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to a documented retrotransposase.
[0209] In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0210] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0211] In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype.
[0212] In some embodiments, the open reading frame is selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, BAD promoter, the strong leftward promoter from phage lambda (pL promoter), or any combination thereof.
[0213] In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to the sequence encoding the retrotransposase. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof.
[0214] In some embodiments, the open reading frame is codon optimized for expression in the host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of the host cell.
[0215] In one aspect, the disclosure provides a culture comprising a host cell described herein in a compatible liquid medium.
[0216] In one aspect, the disclosure provides a method of producing a retrotransposase, comprising culturing a host cell described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the retrotransposase by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or the increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC, or ion affinity chromatography. In some embodiments, the open reading frame comprises a sequence encoding an IMAC affinity tag linked in-frame to the sequence encoding the retrotransposase. In some embodiments, the IMAC affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the IMAC affinity tag by contacting the retrotransposase with a protease corresponding to the protease cleavage site. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the retrotransposase.
[0217] In one aspect, the disclosure provides a method of disrupting a gene locus in a cell. In some embodiments, the method includes contacting the cell with a composition comprising a retrotransposase. In some embodiments, the retrotransposase has at least equivalent transposition activity to a demonstrated retrotransposase in the cell. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0218] In some embodiments, the retrotransposase comprises a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. In some embodiments, the retrotransposase may be substantially identical to any one of SEQ ID NOs: 1-29, 393-735, or 799-895.
[0219] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain.
[0220] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to a documented retrotransposase.
[0221] In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0222] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a single-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0223] In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a eukaryotic genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a fungal genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a plant genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a mammalian genomic polynucleotide sequence. In some embodiments, the retrotransposase comprises a sequence that is complementary to a human genomic polynucleotide sequence.
[0224] In some embodiments, the retrotransposase may include a variant having one or more nuclear localization sequences (NLS). The NLS may be proximal to the N-terminus or C-terminus of the retrotransposase. The NLS may be added to the N-terminus or C-terminus of any one of SEQ ID NOs: 896-911 or a variant having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to any one of SEQ ID NOs: 896-911. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 896. In some embodiments, the NLS may comprise a sequence substantially identical to SEQ ID NO: 897.
[0225] In some embodiments, the transposition activity is measured in vitro by introducing a retrotransposase into a cell containing a target nucleic acid locus and detecting transposition of the target nucleic acid locus in the cell. In some embodiments, the composition comprises 20 pmoles or less of retrotransposase. In some embodiments, the composition comprises 1 pmol or less of retrotransposase.
[0226] The disclosed system may be used for a variety of applications, such as, for example, binding to nucleic acid molecules (e.g., sequence-specific binding), nucleic acid editing (e.g., gene editing), etc. Such systems may be used, for example, to address (e.g., remove or replace) genetically inherited mutations that may cause disease in a subject, to inactivate genes to confirm their function in cells, as diagnostic tools to detect disease-causing genetic elements (e.g., via cleavage of reverse-transcribed viral RNA or amplified DNA sequences that code for disease-causing mutations), as inactivated enzymes combined with probes to target specific nucleotide sequences (e.g., sequences that code for antibiotic resistance in bacteria), to inactivate viruses by targeting viral genomes or to prevent them from infecting host cells, to add genes or modify metabolic pathways to engineer organisms to produce valuable small molecules, macromolecules, or secondary metabolites, to establish gene drive elements for evolutionary selection, and as biosensors to detect cellular perturbations by exogenous small molecules and nucleotides. EXAMPLES
[0227] In accordance with IUPAC convention, the following abbreviations are used throughout the examples: A=Adenine C=Cytosine G=guanine T=Thymine R = adenine or guanine Y = cytosine or thymine S = guanine or cytosine W = adenine or thymine K = guanine or thymine M = adenine or cytosine B=C, G, or T D=A, G, or T H=A, C, or T V=A, C, or G
[0228] Example 1 - Methods for metagenomic analysis of new proteins Metagenomic samples were collected from sediments, soils, and animals. Deoxyribonucleic acid (DNA) was extracted using Zymobiomics DNA miniprep kit and sequenced on an Illumina HiSeq® 2500. Samples were collected with the consent of the owners. Additional raw sequence data from public sources included animal microbiome, sediment, soil, hot springs, hydrothermal vents, ocean, peat, permafrost, and sewage sequences. The metagenomic sequence data was searched using a hidden Markov model generated based on validated retrotransposase protein sequences to identify new retrotransposases. Novel retrotransposase proteins identified by the search were aligned against validated proteins to identify potential active sites. This metagenomic workflow resulted in the delineation of the MG140 family described herein.
[0229] Example 2 - Discovery of the MG140 family of retrotransposases Analysis of the data from the metagenomic analysis of Example 1 revealed a new cluster of putative retrotransposase systems, including one undescribed family (MG140). Protein sequences corresponding to these new enzymes and their exemplary subdomains are set forth as SEQ ID NOs: 1-29, 393-401, and 799-894.
[0230] Example 3 - Incorporation of reverse transcription DNA in vitro activity (predictive) Integrase activity can be performed via expression in an E. coli lysate-based expression system (e.g., myTXTL, Arbor Biosciences). The components used for in vitro testing are three plasmids: an expression plasmid with a retrotransposon gene under a T7 promoter, a target plasmid, and a donor plasmid containing 5' and 3' UTR sequences recognized by the retrotransposase around a selection marker gene (e.g., Tet resistance gene). The lysate-based expression products, target DNA, and donor plasmid are incubated to allow transposition to occur. Transposition is detected via PCR. In addition, transposition products are tagmented with T5 and sequenced via NGS to determine the insertion site on a population of transposition events. Alternatively, in vitro transposition products can be transformed into E. coli under antibiotic (e.g., Tet) selection, and growth occurs when the selection marker is stably inserted into the plasmid. Either single colonies or populations of E. coli can be sequenced to determine the insertion site.
[0231] Incorporation efficiency can be measured via ddPCR or qPCR of the experimental output of target DNA with incorporated cargo, normalized to the amount of unmodified target DNA, also measured via ddPCR.
[0232] This assay may also be performed on purified protein components rather than from lysate-based expression. In this case, proteins are expressed in E. coli protease-deficient B strain under a T7 inducible promoter, cells are lysed using sonication, and His-tagged proteins of interest are purified using HisTrap FF (GE Lifescience) Ni-NTA affinity chromatography on an AKTA Avant FPLC (GE Lifescience). Purity is determined using SDS-PAGE and densitometry in ImageLab software (Bio-Rad) of protein bands resolved on InstantBlue Ultrafast (Sigma-Aldrich) oomassie-stained acrylamide gels (Bio-Rad). Proteins are desalted in a storage buffer composed of 50 mM Tris-HCl, 300 mM NaCl, 1 mM TCEP, 5% glycerol, pH 7.5 (or other buffer determined for maximum stability) and stored at -80°C. After purification, the transposon gene is added to the target DNA and donor plasmid described above in a reaction buffer, e.g., 26 mM HEPES pH 7.5, 4.2 mM TRIS pH 8, 50 μg / mL BSA, 2 mM ATP, 2.1 mM DTT, 0.05 mM EDTA, 0.2 mM MgCl2, 30-200 mM NaCl, 21 mM KCl, 1.35% glycerol (measured to be pH 7.5) supplemented with 15 mM MgOAc2.
[0233] Example 4 - Verification of retrotransposon ends via gel shift (predictive) Retrotransposon ends are tested for retrotransposase binding via electrophoretic mobility shift assay (EMSA). In this case, target DNA fragments (100-500 bp) are end-labeled with FAM via PCR with FAM-labeled primers. 3'UTR RNA and 5'UTR RNA are generated and purified in vitro using T7 RNA polymerase. Retrotransposase protein is synthesized in an in vitro transcription / translation system (e.g., PURExpress). After synthesis, 1 μL of protein is added to 50 nM of labeled DNA and 100 ng of 3' or 5' UTR RNA in a 10 μL reaction in binding buffer (e.g., 20 mM HEPES pH 7.5, 2.5 mM Tris pH 7.5, 10 mM NaCl, 0.0625 mM EDTA, 5 mM TCEP, 0.005% BSA, 1 μg / mL poly(dI-dC), and 5% glycerol). Binding is incubated at 30° for 40 minutes, and then 2 μL of 6X loading buffer (60 mM KCl, 10 mM Tris pH 7.6, 50% glycerol) is added. Binding reactions are resolved and visualized on a 5% TBE gel. A shift of the 3' or 5' UTR in the presence of retrotransposase protein and target DNA can be attributed to successful binding and is indicative of retrotransposase activity. This assay can also be performed with truncations or mutations of the retrotransposase, as well as using E. coli extracts or purified protein.
[0234] Example 5 - Verification of target DNA cleavage (predictive) To confirm that retrotransposase is responsible for cleavage of target DNA, short (approximately 140 bp) DNA fragments are labeled at both ends with FAM via PCR using FAM-labeled primers. In vitro transcription / translation retrotransposase products are preincubated with 1 μg of Rnase A (negative control), or 3'UTR, 5'UTR, or non-specific RNA fragments (controls), followed by incubation with labeled target DNA at 37° C. The DNA is then analyzed on a denaturing gel. Cleavage of one or both strands of DNA can result in labeled fragments of different sizes that migrate at different speeds on the gel.
[0235] Example 6 - Integrase activity in E. coli (predicted) The engineered E. coli strain is transformed with a plasmid expressing the retrotransposon gene and a plasmid containing a temperature-sensitive origin of replication with a selectable marker flanked by the 5' and 3' UTRs of the retrotransposon involved in integration. Transformants induced to express these genes are then screened for transfer of the marker to the genomic target by selection at the restrictive temperature for plasmid replication, and marker integration within the genome is confirmed by PCR.
[0236] Integration is screened using an unbiased approach. Briefly, purified gDNA is tagged with Tn5, and then the DNA of interest is PCR amplified using primers specific for Tn5 tagmentation and selectable marker. The amplicons are then prepared for NGS sequencing. Analysis of the resulting sequences is trimmed from transposon sequences, and flanking sequences are mapped to the genome to determine the insertion position and the insertion ratio.
[0237] Example 7 - Integration of reverse transcribed DNA into mammalian genomes (predictive) To demonstrate targeting and cleavage activity in mammalian cells, the integrase protein is purified in E. coli or sf9 cells with two NLS peptides either at the N-terminus, C-terminus, or both termini of the protein sequence. In this procedure, a plasmid is synthesized containing a selectable neomycin resistance marker (NeoR) or a fluorescent marker flanked by the 5' and 3' UTR regions involved in transposition and under the control of the CMV promoter. Cells are transfected with the plasmid, allowed to recover for 4-6 hours for RNA transcription, and then electroporated with the purified integrase protein. Antibiotic resistance integration into the genome is quantified by G418-resistant colony counts (selection to begin 7 days after transfection) and positive transposition by the fluorescent marker is assayed by fluorescence-activated cell cytometry. 7-10 days after the second transfection, genomic DNA is extracted and used for preparation of NGS libraries. Off-target frequency is assayed by fragmenting the genome and preparing amplicons of the transposon marker and flanking DNA for NGS library preparation. At least 40 different target sites are selected to test the activity of each targeting system.
[0238] Integration into mammalian cells can also be assessed via RNA delivery. An RNA encoding a retrotransposase with two NLSs is designed and capped and polyA-tailed. A second RNA is designed to contain a selectable neomycin resistance marker (NeoR) or a fluorescent marker flanked by the 5'UTR and 3'UTR regions. The RNA construct is introduced into mammalian cells via Lipofectamine™ RNAiMAX or TransIT®-mRNA transfection reagent. Ten days after transfection, genomic DNA is extracted to measure transposition efficiency using ddPCR and NGS.
[0239] Example 8 - Bioinformatic discovery of RT An extensive assembly-driven metagenomic database of microbial, viral, and eukaryotic genomes was mined to obtain predicted proteins with reverse transcriptase function. Over 4.5 million RT proteins were predicted based on having hits to the Pfam domains PF00078 and PF07727, of which 3.4 million had significant e-values (1 × 10 -5 After filtering for complete ORFs with 70% or more RT (reverse transcriptase) domain coverage and predicted catalytic residues ([F / Y]XDD), approximately 500,000 proteins were retained for further analysis. RT domains were extracted from this set of proteins as well as from reference sequences obtained from public databases. Domain sequences were clustered at 50% identity over 80% coverage using Mmseqs2 easy-cluster (see Bioinformatics 2016 May 1;32(9):1323-30, incorporated herein by reference in its entirety), representative sequences (26,824 in total) were aligned with MAFFT using the parameter globalpair-large (see Bioinformatics 2016;32:3246-3251, incorporated herein by reference in its entirety), and the domain alignments were used to infer phylogenetic trees with FastTree2 (see Plos One 2010;5:e9490, incorporated herein by reference in its entirety). Phylogenetic analysis of RT domains suggests that many different classes of RTs with high sequence diversity were recovered (Figure 4).
[0240] Example 9 - Exemplary non-LTR retrotransposons (MG140, MG146, MG147, MG148, and MG149 families) Retrotransposon bioinformatic analysis Non-long terminal repeat (non-LTR) retrotransposases can integrate large cargos into target sites through reverse transcription of RNA templates. Non-LTR retrotransposases were identified within the R2 / R4 and LINE clades from the phylogenetic tree in Figure 4. Full-length proteins containing the RT domain classified as R2, R4, and LINE were clustered at 99% sequence identity, and representative sequences were aligned with MAFFT using the parameter globalpair-large. A phylogenetic tree was inferred from this alignment to depict the R2 / R4 retrotransposase family, as well as other RT-related families (Figure 5A).
[0241] R2 is a non-LTR retrotransposon that integrates cargo via target-primed reverse transcription (TPRT). Many R2 enzymes of the MG140 family contain an RT domain, as well as an endonuclease domain and multiple Zn-binding ribbon motifs that depict Zn fingers (Figures 5B and 6A). Some R2 retrotransposons integrate into the 28S rDNA, as shown by the border of the MG140-47 (SEQ ID NO: 395) R2 retrotransposon adjacent to a fragment of the 28S rDNA gene (Figure 6B). Other retrotransposons integrate into the 18S rRNA gene and contain polyA or polyT tails that define the 3' end of the transposon (Figure 7). The precise target binding site, as well as the 5'-UTR, 3'-UTR, and poly-T may be involved in accurate and specific integration.
[0242] Retrotransposon MG146-1 (sequence number 402), derived from an archaeal genome, contains an RT domain, a Zn-binding ribbon motif, and an endonuclease domain, and the domain architecture within the enzyme differs from that of other single ORF non-LTR retrotransposons (Figure 8A).
[0243] The MG147 family member MG140-17-R2 (SEQ ID NO: 18) retrotransposon is organized into three ORFs flanked by 5'UTR and 3'UTR (Figure 8B). The RNA recognition motif (RRM) genes are likely involved in recognition of the RNA template, while the endonuclease genes are likely involved in target site recognition and nicking. ORF3 is the enzyme involved in reverse transcription of the template and contains the RT domain, Zn-binding ribbon motif, and RNAse-H domain.
[0244] The MG148 family contains highly divergent RT homologs that are predicted to be active by the presence of all of the predicted catalytic residues. Nucleotide-level alignment of several family members did not reveal conserved regions within the 5'UTR that might be involved in RT function, activity, or recruitment (Figure 9B).
[0245] Testing the in vitro activity of retrotransposon RT (reverse transcriptase) by qPCR The in vitro activity of retrotransposon RT was evaluated by primer extension reaction containing RT enzyme from a cell-free expression system (PURExpress, NEB) and 100 nM RNA template (200 nt) annealed to a DNA primer in a reaction buffer containing 40 mM Tris-HCl (pH 7.5), 0.2 M NaCl, 10 mM MgCl2, 1 mM TCEP, and 0.5 mM dNTPs. The resulting full-length cDNA products were quantified by qPCR by extrapolating values from a standard curve generated with a specific concentration of DNA template.
[0246] MG140-3 (SEQ ID NO:3), MG140-6 (SEQ ID NO:6), MG140-7 (SEQ ID NO:7), MG140-8 (SEQ ID NO:8), MG140-13 (SEQ ID NO:14), and MG146-1 (SEQ ID NO:402) are active through primer extension (Figures 10 and 11). A preliminary assessment of fidelity was performed for MG140-3 and MG146-1, which resulted in 1.5-fold and 1.35-fold higher relative error rates than MMLV, respectively (Figure 12). For fidelity measurements, the resulting full-length cDNA products generated in the primer extension assay described above were PCR amplified, library prepared, and subjected to next-generation sequencing. Trimmed reads were aligned to the reference sequence, and the frequency of misincorporation was calculated.
[0247] Integration site Some non-LTR retrotransposons (e.g., MG140 family, such as MG140-1) integrate into 28S rDNA genes by targeting a specific GGTGAC motif, with the insertion site predicted to be between the second (G) and third (T) positions. The N-terminus of such retrotransposon proteins contains three zinc (Zn) fingers (two of the CCHH type and one of the CCHC type) followed by a reverse transcriptase (RT) domain with a YADD active site. The C-terminus of such retrotransposon proteins contains an endonuclease domain with an additional CCHC Zn finger. The proteins are flanked by 5'UTR and 3'UTR, which are 289 bp and 478 bp in length, respectively (Figure 31).
[0248] Example 10 - Group II intron RTs (MG153, MG163, MG164, MG165, MG166, MG167, MG168, MG169, and MG170 families) Group II bioinformatic analysis Group II introns can incorporate large cargos into target sites via reverse transcription of an RNA template. RT domains from group II introns were identified and depicted in a phylogenetic tree in Figure 4. Over 10,000 unique full-length group II intron proteins containing RT domains from contigs with >2 kb of sequence flanking the RT enzyme were aligned with MAFFT using the parameter globalpair-large. A phylogenetic tree was inferred from this alignment to further identify group II intron families (Figure 13). Group II intron enzymes can be classified into classes A-G, ML, and CL, and their domain architecture includes an RT domain predicted to be active, as well as a maturase domain involved in intron recruitment. Some group II intron proteins contain an additional endonuclease domain that may be involved in target recognition and cleavage. A number of candidates from all identified families were nominated for laboratory characterization.
[0249] Testing the in vitro activity of group II intron RTs, classes C, D, and F The in vitro activity of GII intron class C (MG153), class D (MG165), and class F (MG167) RTs was assessed by primer extension reactions containing RT enzymes derived from a cell-free expression system (PURExpress, NEB). The expression constructs were codon-optimized for E. coli and contained an N-terminal single Strep tag. Expression of the RTs was confirmed by SDS-PAGE analysis. The substrate for the reaction was 100 nM RNA template (200 nt) annealed to a 5'-FAM-labeled primer. The reaction buffer contained the following components: 50 mM Tris-HCl (pH 8.0), 75 mM KCl, 3 mM MgCl2, 10 mM DTT, and 0.5 mM dNTPs. After 1 hour incubation at 37°C, the reaction was quenched via incubation with RnaseH (NEB) followed by the addition of 2X RNA loading dye (NEB). The resulting cDNA products were separated on a 10% denaturing polyacrylamide gel and visualized using ChemiDoc with Gel Green settings. RT activity was also assessed by qPCR using primers that amplify full-length cDNA products. Products from the primer extension assay were diluted to ensure that cDNA concentrations were within the linear range of detection. The amount of cDNA was quantified by extrapolating values from a standard curve generated using specific concentrations of DNA template.
[0250] By detecting cDNA products on denaturing gels, the following GII intron class C candidates were active under these experimental conditions: MG153-1 to MG153-6 (SEQ ID NO: 555 to 560), MG153-9 (SEQ ID NO: 563), MG153-10 (SEQ ID NO: 564), MG153-12 (SEQ ID NO: 566), MG153-13 (SEQ ID NO: 567), MG153-15 (SEQ ID NO: 569), MG153-18 (SEQ ID NO: 572), MG153- 20 (SEQ ID NO: 574), MG153-29 to MG153-31 (SEQ ID NO: 580 to 582), MG153-33 to MG153-37 (SEQ ID NO: 584 to 588), MG153-41 (SEQ ID NO: 592), MG153-42 (SEQ ID NO: 593), MG153-45 (SEQ ID NO: 596), MG153-51 (SEQ ID NO: 602), MG153-53 (SEQ ID NO: 604), MG153-54 (SEQ ID NO: 605), and MG153-57 (SEQ ID NO: 608). (Figures 14 and 15). The active novel candidates show various degrees of apparent processivity compared to the highly processive control GII class C RTs GsI-IIC and MarathonRT, as indicated by the presence of smaller cDNA drop-off products. By qPCR, the following additional candidates were also active under these experimental conditions (cDNA was detected 10-fold above background): MG153-7 (SEQ ID NO: 561), MG153-8 (SEQ ID NO: 562), MG153-10 (SEQ ID NO: 564), MG153-11 (SEQ ID NO: 565), MG153-14 (SEQ ID NO: 568), MG153-17 (SEQ ID NO: 571), MG153-19 (SEQ ID NO: 573), MG 153-25 to MG153-28 (sequence numbers 576 to 579), MG153-32 (sequence number 583), MG153-39 (sequence number 590), MG153-40 (sequence number 591), MG153-43 (sequence number 594), MG153-47 (sequence number 598), MG153-50 (sequence number 601), MG153-55 (sequence number 606), and MG153-56 (sequence number 607) (Figures 14D and 15D).
[0251] By detecting cDNA products on denaturing gels, GII intron class D candidates MG165-1 (SEQ ID NO: 684) and MG165-5 (SEQ ID NO: 688) are active under these experimental conditions (Figure 16A). By qPCR, additional candidates MG165-4 (SEQ ID NO: 687), MG165-6 (SEQ ID NO: 689), and MG165-8 (SEQ ID NO: 691) are also active under these experimental conditions (cDNA was detected at 10-fold above background) (Figure 16B).
[0252] By detecting cDNA products on denaturing gels, GII intron class F candidates MG167-1 (SEQ ID NO: 698) and MG167-4 (SEQ ID NO: 701) are active under these experimental conditions (Figure 17A). By qPCR, additional candidates MG167-3 (SEQ ID NO: 700) and MG167-5 (SEQ ID NO: 702) are also active under these experimental conditions (cDNA was detected at 10-fold above background) (Figure 17B).
[0253] Assessment of the relative fidelity of GII intron RTs To evaluate the relative fidelity of the GII class C MG153 candidates, the resulting full-length cDNA products generated in the primer extension assay described above were PCR amplified, library prepared, and subjected to next-generation sequencing. Paired reads were merged using bbmerge.sh, which requires perfect overlap, and trimmed all non-overlapping parts (Plos One 2017;12:e0185056). The merged reads were then aligned to the reference template using BWA-MEM (Li H.2013), and the number of mismatches at each position relative to the reference was calculated using pysamstats (https: / / github.com / alimanfoo / pysamstats). Of the GII class C candidates tested, MG153-6 (SEQ ID NO: 560) and MG153-12 (SEQ ID NO: 566) have reproducibly higher error rates compared to the MMLV control RT and other GII intron class C RTs (Figure 18).
[0254] Human cell cDNA synthesis results The ability of these enzymes to produce cDNA in a mammalian environment was tested by expressing them in mammalian cells and detecting cDNA synthesis by PCR followed by agarose electrophoresis and D1000 TapeStation. Reverse transcriptase was cloned in a plasmid for mammalian expression under a CMV promoter as a fusion protein with MS2 coat protein (MCP) at the N-terminus in addition to a flag-HA tag (FH). MCP is a protein derived from MS2 bacteriophage that recognizes 20-nucleotide RNA stem loops with high affinity (sub-nanomolar Kd). Fusing RT to MCP and having the MS2 loop in the RNA template ensures that when RT is translated, it will find the RNA template and initiate cDNA synthesis from a DNA primer hybridized to the RNA template.
[0255] The plasmid containing MCP fused to the RT candidate under the CMV promoter was cloned and isolated for transfection in HEK293T cells. Transfection was performed using Lipofectamine 2000. mRNA encoding nanoluciferase (SEQ ID NO: 33) was generated using mMESSAGE mMACHINE (Thermo Fisher) according to the manufacturer's instructions. To degrade any DNA template left in the mRNA preparation, the reaction was treated with Turbo Dnase (Thermo Fisher) for 1 hour, and the mRNA was cleaned using MEGAclear Transcription Clean-Up Kit (Thermo Fisher). The mRNA was hybridized to the complementary DNA primer (SEQ ID NO: 34) in 10 mM Tris pH 7.5, 50 mM NaCl at 95°C for 2 minutes and cooled to 4°C at a rate of 0.1°C / s. The mRNA / DNA hybrids were transfected into HEK293T cells using Lipofectamine Messenger Max 6 hours after transfection with the plasmid containing the MCP-RT fusion. 18 hours after mRNA / DNA transfection, the cells were lysed using QuickExtra DNA Extraction Solution (Lucigen) and 100 μL of quick extract was added per 24 wells in a 24-well plate. Nanoluciferase is approximately 500 bp in length, and primers were designed to amplify 100 bp and 542 bp products from the newly synthesized cDNA (SEQ ID NOs: 38 and 39). The cDNA was amplified using the above-mentioned set of primers and the PCR products were detected by agarose gel electrophoresis (Figure 19A) or DNA Tape Station (Figure 19B).
[0256] Activity for the control GII intron RTs Marathon, Marathon PE2, and TGIRT was detected as indicated by the presence of 100 bp and 500 bp DNA products (Figures 19A and 19B). In addition, activity for the novel GII intron-derived RTs MG153-1 to MG153-4 (SEQ ID NOs: 555 to 558), MG153-7 to MG153-13 (SEQ ID NOs: 561 to 567), MG153-15 (SEQ ID NO: 569), MG153-16 (SEQ ID NO: 570), and MG153-21 (SEQ ID NO: 575) was also shown (Figures 19A, 19B, and 19C). The PCR product signals for the novel RTs were similar to those for Marathon and TGIRT. Taken together, this indicates that these newly discovered RTs are expressed, properly folded, and active in living mammalian cells, broadening their options for biotechnological applications.
[0257] Group II intron RTs can synthesize cDNA using modified primers The in vitro activity of RT was assessed by primer extension reactions containing RT enzyme from a cell-free expression system (PURExpress, NEB). The expression construct was codon-optimized for E. coli and contained an N-terminal single Strep tag. The substrate for the reaction was 100 nM of RNA template (202 nt) annealed to a 5'-FAM labeled DNA primer containing phosphorothioate (PS) bond modifications at various positions within the primer. Primer 1 (sequence / 56-FAM / A*G*A*C*G*GTCACAGCTTGTCTG, SEQ ID NO: 736) contains five PS bonds at the 5' end of the oligo. Primer 2 (sequence / 56-FAM / A*G*A*C*G*GTCACAGCTT*G*T*C*T*G, SEQ ID NO: 737 (where * indicates phosphorothioate bond)) contains five PS bonds at both the 5' and 3' ends of the oligo. Primer 3 (SEQ ID NO: 738, containing the sequence / 56-FAM / A*G*A*C*G*GTCACAGCTT*G*T*C*TG, where * indicates a phosphorothioate bond) differs from primer 2 in that a standard bond is replaced between the two most 3'-terminal nucleotides. The reaction buffer contained the following components: 50 mM Tris-HCl (pH 8.0), 75 mM KCl, 3 mM MgCl2, 10 mM DTT, and 0.5 mM dNTPs. After incubation at 37°C for 1 hour, the reaction was quenched via incubation with RnaseH (NEB), followed by the addition of 2X RNA loading dye (NEB). The resulting cDNA products were separated on a 10% denaturing polyacrylamide gel and visualized using ChemiDoc with Gel Green settings. Based on these results, both the control RTs MMLV (virus) and TGIRT-III (GII intron) are able to perform primer extension with all modified primers (Figure 32). The GII intron RT MG153-9 is also able to extend from all tested PS-modified DNA primers (Figure 33).
[0258] Human cell RT expression and cDNA synthesis results The ability of the new GII RT to synthesize cDNA in a mammalian cell environment was tested as described above with no modifications. cDNA synthesis was detected using PCR and analyzed by agarose gel electrophoresis or TapeStation. In order to have a quantitative readout, a Taqman qPCR assay was developed using Taqman qPCR primers previously verified with the Taqman probe listed as SEQ ID NO: 739. All tested candidates of the MG153 family were active to various degrees, with activity ranging by as much as four orders of magnitude (Figure 34). The RTs of the family tested include MG153-1 to MG153-13, MG153-15, MG153-16, MG153-18, MG153-20, MG153-21, MG153-29 to MG153-31, MG153-33 to MG153-37, MG153-45, MG153-51, MG153-53, MG153-54, MG153-57, MG165-1, MG165-5, MG167-1, and MG167-4. Some RTs (MG153-15, MG153-53, MG153-4, MG153-18, MG153-20, MG153-7, and MG153-5) were superior to the TGIRT control (Figure 34).
[0259] Immunoblots were performed to understand the protein expression and stability of GII RT in mammalian cells. Briefly, transfected cells were lysed with RIPA lysis buffer (Thermo Fisher) supplemented with protease inhibitors (80 μL per well in 24-well format). Lysates were centrifuged at 14,000 g for 10 min at 4 °C to remove insoluble aggregates. Protein was quantified using BCA. 3 or 10 μg of total protein was loaded per lane on a 4-12% polyacrylamide SDS gel (Thermo Fisher). All lanes were normalized to the same amount of protein. Protein was transferred to a PVDF membrane using the iBlot gel transfer system (Invitrogen). Protein was detected by using a rabbit HA antibody (Cell Signaling) using an HRP-based detection method. The results suggest various levels of protein expression or stability, as indicated by the intensity of the bands (Figure 35). We quantified the expression of each protein and normalized cDNA synthesis activity to total protein expression: 7 MG153RT outperformed the TGIRT control (Figure 36). Notably, MG153-15 exhibits 10-fold higher cDNA synthesis activity than TGIRT under these conditions.
[0260] Some GII-derived RTs, including MarathonRT, one of the positive controls, and MG153-1 to MG153-4 and MG153-9, form very stable dimers (Figure 35). The "CAQQ" motif was demonstrated to be involved in stable dimer formation in Marathon RT (Nat Struct Mol Biol. 2016 Jun;23(6):558-565). RTs that showed stable dimer formation on immunoblot (MG153-1 to MG153-4) also contain the CAQQ dimer formation amino acid motif (Figure 35C). Dimer formation may be an undesirable feature due to additional complexity, so RTs that do not form dimers may be optimal for certain biotechnological applications.
[0261] [Table 2]
[0262] Example 11 - G2L4 (MG172 family) G2L4 is an RT-containing sequence distantly related to group II introns (group II intron-like RT), identified in Figure 4. More than 600 novel full-length G2L4 enzymes were aligned with MAFFT using the parameter globalpair-large, and a phylogenetic tree was inferred from this alignment (Figure 20). MG172 family members contain RT and maturase domains and are predicted to have a conserved Y[I / L]DD active site motif. The motif YIDD was recently reported to show increased efficiency with shorter DNA primers in one G2L4 reference (BioRxiv10.1101 / 2022.03.14.484287). MG172 enzymes have an average length of 425 aa and share 32% AAI, highlighting the novelty of these systems.
[0263] Example 12 - LTR retrotransposons (MG151 family) LTR retrotransposon bioinformatic analysis Long terminal repeat (LTR) retrotransposons integrate into their target sites via reverse transcription of an RNA template. The MG151 family of LTR retrotransposons, including retroviral and non-viral transposons, was identified in the phylogenetic tree in Figure 4. Full-length proteins containing the LTR RT domain were aligned with MAFFT using the parameter globalpair-large. A phylogenetic tree was inferred from this alignment (Figure 21A). Over 100 non-viral and retroviral RT enzymes in the MG151 family contain RT and RnaseH domains and are predicted to be active based on the presence of catalytic residues. The LTR RT polyprotein also encodes a protease domain and an integrase domain in a similar architecture found in HIV and MMLV LTR RT (Figures 21A, 21B, 21C, and 22). RT, and other genes such as gag or envelope, are flanked by long imperfect long terminal repeats (Figure 21B). MG151 family members are diverse and novel, sharing 30% amino acid identity (Figure 22).
[0264] Polyproteins of LTR retrotransposons are naturally processed into protease, RT and Rnase H, and integrase functional units. Therefore, the MG151 RT-RNAse H functional unit boundaries were determined by a combination of sequence and structural alignment. The 3D structure of the MG151 polyprotein was predicted using Alphafold2 (Nature 2021;596:583-589, and Nucleic Acids Res 2022;50:D439-D444) and visualized with PyMOL (https: / / github.com / schrodinger / pymol-open-source). For example, for MG151-82 (sequence number 457), the predicted 3D structure identified distinct protease, RT, RNAseH, and integrase domains separated by an unstructured linker region (Figure 21C). Therefore, the RT-RNAse H functional unit was determined as two related structural domains adjacent to an unstructured loop. A trimmed variant containing the RT and RNAse H domains was designated for synthesis and laboratory characterization.
[0265] Testing the in vitro activity of LTR retrotransposon RT The in vitro activity of LTR retrotransposon RT (MG151) was assessed by primer extension reactions containing the RT enzyme from the cell-free expression system and an RNA template annealed to a 5'-FAM-labeled primer as described above in a reaction buffer containing 50 mM Tris-HCl pH 8, 75 mM KCl, 3 mM MgCl2, 1 mM TCEP, and 0.5 mM dNTPs. The resulting cDNA products were separated on a denaturing polyacrylamide gel and visualized using ChemiDoc with Gel Green settings. Based on these results, MG151-80 to MG151-84 (Figure 23A), as well as MG151-87 to MG151-90 (SEQ ID NOs: 524 to 527), and MG151-92 to MG151-95 (SEQ ID NOs: 529 to 532) (Figure 23B) are capable of synthesizing cDNA in vitro.
[0266] To determine assay conditions under which in vitro activity was observed for the control LTR retrotransposon RT, Ty3, four reaction buffers were tested: Buffer A (40 mM Tris-HCl pH 7.5, 0.2 M NaCl, 10 mM MgCl2, 1 mM TCEP), Buffer B (20 mM Tris pH 7.5, 150 mM KCl, 5 mM MgCl2, 1 mM TCEP, 2% PEG-8000), Buffer C (10 mM Tris-HCl pH 7.5, 80 mM NaCl, 9 mM MgCl2, 1 mM TCEP, 0.01% (v / v) Triton X-100), and Buffer D (10 mM Tris pH 7.5, 130 mM NaCl, 9 mM MgCl2, 1 mM TCEP, 10% glycerol). In vitro activity was observed for buffers A and B (Figure 23C).
[0267] Testing priming parameters and processivity on structural RNA templates To determine the reverse transcriptase activity of these LTR RTs on structured RNA templates, different primers of 6, 8, 10, 13, 16, and 20 nt in length were annealed onto the structured RNA scaffold. These annealed RNA / DNA hybrids were used in a cDNA generation assay equivalent to the assay used for overall activity. As shown in Figure 24, MMLV is active on structured RNA with a primer binding site of 10-20 nt, extending the template completely to the 5' end and unfolding all structures in the template. MG151-89 (SEQ ID NO: 526) is active with a primer length of 13-20 and can extend approximately 18 nt, the length of the pegRNA, until it reaches the sgRNA scaffold hairpin. MG151-92 (SEQ ID NO: 529) and MG151-97 (SEQ ID NO: 534) were not active on this template at our level of detection.
[0268] Example 13 - Retron RT (MG154, MG155, MG156, MG157, MG158, MG159, and MG160 families) Retron bioinformatic analysis Bacterial retrons are DNA elements approximately 2000 bp in length that code for a continuous non-coding RNA containing the RT-encoding gene (ret) and the inverted sequences msr and msd. Retrons use a unique mechanism for RT-DNA synthesis, where the ncRNA template folds into a conserved secondary structure insulated between two inverted repeats (a1 / a2). The retron RT recognizes the folded ncRNA, and reverse transcription is initiated from a conserved guanosine 2'OH adjacent to the inverted repeat, forming a 2'-5' linkage between the template RNA and the nascent cDNA strand. In some retrons, this 2'-5' linkage persists into the mature form of the processed RT-DNA, while in other embodiments, an exonuclease cleaves the DNA product, resulting in a free 5' end. Furthermore, RT targets msr-msd derived from the same retron as its RNA template, providing specificity that may avoid off-target reverse transcription.
[0269] More than 4031 RT domain sequences were identified as retron RTs in the phylogenetic tree in Figure 4. A subset of 2407 full-length retron protein sequences was selected for further analysis based on the presence of catalytic residues (xxDD) and conserved motifs (NaxxH and VTG) demonstrated in retron RTs (Figures 25 and 26). Retrons of the MG154-MG159 and MG173 families include members ranging from 300 to 650 aa in length, and their 5'UTRs contain predicted ncRNAs (msr-msd) flanked by inverted repeats and trimmed (Figure 27).
[0270] In addition, a divergent group of "retron-like" single domain RT sequences was identified within the retron clade in Figure 4. The single domain RTs of the MG160 family span 250-300 aa and are predicted to be active based on the presence of predicted RT catalytic residues [F / Y]XDD. Although there are no retron RT crystal or cryo-EM structures in the public database, the 3D structure prediction of MG160-3 (SEQ ID NO: 629) shows a conserved RT domain that aligns with the group II intron RT domain (Figures 28A and 28B). The 5'UTR of the MG160 family folds into a conserved secondary structure that is conserved among family members and is likely important for element activity or recruitment (Figure 28C).
[0271] In vitro activities of the MG154, MG155, MG156, MG157, MG158, and MG159 families of retron-like RTs The in vitro activity of retron RT on a general RNA template was assessed by primer extension reactions containing RT enzyme derived from a cell-free expression system (PURExpress, NEB). The expression construct was codon-optimized for E. coli and contained an N-terminal single Strep tag. The substrate for the reaction was 100 nM of RNA template (202 nt) annealed to a 5'-FAM-labeled primer. The reaction buffer contained the following components: 50 mM Tris-HCl (pH 8.0), 75 mM KCl, 3 mM MgCl2, 10 mM DTT, and 0.5 mM dNTPs. After 1 h incubation at 37°C, the reaction was quenched via incubation with RnaseH (NEB) followed by the addition of 2X RNA loading dye (NEB). The resulting cDNA products were separated on a 10% denaturing polyacrylamide gel and visualized using ChemiDoc with Gel Green settings. Based on these results, the following retron RTs are capable of performing primer extension on general RNA templates that are not their own ncRNA: MG155-2 (SEQ ID NO: 612), MG155-3 (SEQ ID NO: 613), MG156-2 (SEQ ID NO: 617), MG157-5 (SEQ ID NO: 622), and MG159-1 (SEQ ID NO: 624).
[0272] In vitro activity of the MG160 family of retron-like RTs The in vitro activity of retron-like RT (MG160 family) was assessed by primer extension reactions containing RT enzyme from a cell-free expression system (PURExpress, NEB). The expression construct was codon-optimized for E. coli and contained an N-terminal single Strep tag. The substrate for the reaction was 100 nM of RNA template (200 nt) annealed to a 5'-FAM-labeled primer. The reaction buffer contained the following components: 50 mM Tris-HCl (pH 8.0), 75 mM KCl, 3 mM MgCl2, 10 mM DTT, and 0.5 mM dNTPs. After 1 h incubation at 37°C, the reaction was quenched via incubation with RnaseH (NEB) followed by the addition of 2X RNA loading dye (NEB). The resulting cDNA products were separated on a 10% denaturing polyacrylamide gel and visualized using ChemiDoc with Gel Green settings. RT activity was also assessed by qPCR using primers that amplify full-length cDNA products. Products from primer extension assays were diluted to ensure that cDNA concentrations were within the linear range of detection. The amount of cDNA was quantified by extrapolating values from a standard curve generated using DNA templates of demonstrated concentrations.
[0273] By gel analysis, MG160-1 to MG160-4 (SEQ ID NO: 627 to 630) and MG160-6 (SEQ ID NO: 633) were active and had reduced processivity compared to the control GII intron class C RT, GsI-IIC (Figure 29). The processivity appears to be more similar to that of the retroviral control RT, MMLV, which results in a similar drop-off pattern of cDNA products (Figure 29A). By qPCR, MG160-1 to MG160-4 (SEQ ID NO: 627 to 630) were able to produce full-length cDNA, while MG160-6 (SEQ ID NO: 633) produced a product that was shorter than full-length (Figure 29B).
[0274] Cell-free expression of retron RTs (MG154, MG155, MG156, MG157, MG158, MG159, and MG173 families) and in vitro transcription of retron ncRNAs Retron RTs were produced in a cell-free expression system (PURExpress) by incubating 10 ng / μL of DNA template encoding the E. coli optimized gene with an N-terminal single Strep tag bearing PURExpress components for 2 hours at 37° C. All tested retron RTs (MG156-1 (SEQ ID NO: 616), MG156-2 (SEQ ID NO: 617), MG157-1 (SEQ ID NO: 618), MG157-2 (SEQ ID NO: 619), MG157-5 (SEQ ID NO: 622), MG159-1 (SEQ ID NO: 624)) were produced as shown by SDS-PAGE analysis (FIGS. 30A and 30B).
[0275] Retron ncRNAs were generated using HiScribe T7 in vitro transcription kit (NEB) and DNA templates encoding the respective ncRNA genes behind the T7 promoter. The reaction was then incubated with Dnase-I to remove the DNA template, and then purified by RNA cleanup kit (Monarch). The amount of ncRNA was determined by Nanodrop, and the purity was assessed by Tape Station RNA analysis (Figure 30C).
[0276] Example 14 - Testing the in vitro activity of Retron RT (prognostic) The retron RT enzyme is produced in a cell-free expression system using a construct containing an E. coli codon-optimized gene with an N-terminal single Strep tag as described above. The expression of the enzyme is confirmed by SDS-PAGE analysis. The retron RT activity on a general template is determined by the primer extension assay described above, containing a 200 nt RNA annealed to a 5'-FAM-labeled DNA primer. The resulting cDNA product is detected on a denaturing polyacrylamide gel or by qPCR using a primer specific for the full-length cDNA product.
[0277] The in vitro activity of retron RT on its own ncRNA is assessed in reactions containing buffer, dNTPs, retron RT produced from a cell-free expression system, and refolded ncRNA. RT activity is compared before and after purification of the RT from a cell-free expression system via an N-terminal single Strep tag. After incubation, half of the reactions are treated with Rnase A / T1. Products before and after Rnase A / T1 treatment are evaluated on denaturing polyacrylamide gels and visualized by SYBR gold staining. In this procedure, Rnase A / T1 is understood to digest the RNA template, resulting in a mass shift to smaller products containing ssDNA. Since Rnase H is expected to improve the homogeneity of the 5' and 3' ssDNA boundaries, the effect of Rnase H on product distribution is also assessed by gel analysis. The covalent bond between the ncRNA template and the ssDNA is confirmed by incubating the RT product with 5' to 3' ssDNA exonuclease (RecJ) before or after treatment with debranching enzyme (DBR1). RecJ is expected to be able to degrade the ssDNA after DBR1 removes the 2'-5' phosphodiester linkage between the RNA and the ssDNA.
[0278] Example 15 - Determination of retron msr-msd boundaries by NGS (predictive) The msr-msd boundaries are determined by unbiased ligation of adapter sequences to the 5' and 3' ends of the msDNA products after removal of the 2'-5' phosphodiester linkages by DBR1. The resulting ligation products are PCR amplified, library prepared, and subjected to next generation sequencing. The sequencing reads are aligned to a reference sequence to determine the 5' and 3' boundaries of the msd. The effect of the presence of Rnase H in the RT reaction on the homogeneity of the 5' and 3' msd boundaries is also evaluated.
[0279] Example 16 - Systematic evaluation of insertion sequences into msd on RT activity (predictive) Sequences of distinct length, predicted secondary structure, and GC content are inserted into the msd at selected insertion sites informed by msd boundaries determined by NGS and ncRNA secondary structure prediction, and the effect of these insertion sequences on RT activity is assessed by gel analysis or NGS as described above.
[0280] Example 17 - Testing the in vitro activity of novel RTs (prospective) RT activity is evaluated using primer extension assay, which contains RT from cell-free expression system and RNA template annealed to DNA primer described above.The obtained cDNA product is detected by denaturing polyacrylamide gel and qPCR as described above.Detection of cDNA drop-off product on denaturing gel provides a relative evaluation of the processivity of new candidates.
[0281] Example 18 - Evaluation of novel RT priming parameters (predictive) The optimal primer length is determined by testing the activity of RT on RNA templates annealed to 5'-FAM-labeled DNA primers of either 6, 8, 10, 13, 16, or 20 nucleotides in length. RT is derived from a cell-free expression system as described above. After incubation of the reaction, the reaction is quenched via the addition of Rnase H. The size distribution of the cDNA product is analyzed on a denaturing polyacrylamide gel as described above. The optimal primer length is determined as the length that allows RT to convert the maximum primer into cDNA product. The experimentally determined optimal primer length is then used in subsequent experiments, such as fidelity and processivity assays, to further characterize RT in vitro.
[0282] Example 19 - Assessment of RT Fidelity (Predictive) RT fidelity is assessed by primer extension assay as described above, except that a 14-nt unique molecular identifier (UMI) barcode is included in the primer for the reverse transcription reaction to account for errors introduced during PCR and sequencing. The resulting full-length cDNA products are PCR amplified, library prepared, and subjected to next-generation sequencing. Barcodes with more than 5 reads are analyzed. Mutations, insertions, and deletions are counted if errors are present in all sequence reads with the same barcode after alignment to the reference sequence. Errors present in one but not all sequencing reads are considered to be introduced during PCR or sequencing. In addition to identifying mutation hotspots in the RNA template, further analysis of substitution, insertion, and deletion profiles is performed. Fidelity measurements are also performed using modified bases in the template, such as pseudouridine.
[0283] Example 20 - Determination of the Progressivity Factor of RT (Predictive) RT processivity is assessed using primer extension assays containing RT enzyme derived from a cell-free expression system as described above and RNA templates ranging from 1.6 kb to 6.6 kb in length annealed to either a 5'-FAM-labeled primer (for gel analysis) or an unlabeled primer (for sequencing analysis).
[0284] The reverse transcription reaction is carried out under single cycle conditions that disfavor the rebinding of the RT enzyme that dropped off from the RNA template during cDNA synthesis. The optimal capture molecule and concentration to achieve the single cycle conditions are experimentally determined. The selected conditions are designed to provide sufficient inhibition of cDNA synthesis when incubated before the start of the reaction, but are otherwise designed not to affect the rate of the reaction. The optimal capture molecules to test include an unrelated RNA template and an unrelated RNA template annealed to DNA primers of various lengths.
[0285] Once the single cycle reaction conditions are optimized, processivity is assessed by pre-equilibrating the RT with the RNA template annealed to the DNA primer in the reaction buffer, followed by initiating the reaction with the addition of dNTPs and the capture molecule of choice. After incubating the reaction, the reaction is quenched by the addition of RnaseH. The size distribution of the cDNA products is analyzed on a denaturing polyacrylamide gel as described above, or subjected to PCR and libraries prepared for long-read sequencing. From these experiments, the processivity coefficient is quantified as the length of the template, which results in 50% of the full-length cDNA product. The median length of the cDNA products from the single cycle primer extension reaction is used to estimate the probability that the RT will dissociate on the tested template. From this, the probability that the RT will dissociate at each nucleotide position is calculated, assuming that each dissociation is an independent event and that the probability of dissociation is equal at all nucleotide positions. The processivity coefficient, which represents the length of the template at 50% of the dissociated RT, is then calculated as 1 / (2*P d ) where P d is the probability of dissociation at each nucleotide.
[0286] Example 21 - Systematic analysis of challenge structures on primer extension (predictive) To assess the effect of a challenging template on RT activity, primer extension reactions are performed as described above with modifications. The RNA template contains one of the following challenging motifs at a fixed distance (100-300 nt) downstream of the primer binding site: a homopolymer stretch, a thermodynamically stable GC-rich stem-loop, a pseudoknot, a tRNA, a GII intron, and an RNA template containing base or backbone modifications (e.g., pseudouridine, phosphothioate linkages). After quenching the reaction, the size distribution of the cDNA products is analyzed by denaturing polyacrylamide gel. Also, adapter sequences are unbiasedly ligated to the 3' ends of the cDNA products using T4 ligase. The ligated products are then PCR amplified and library-prepared for next-generation sequencing to identify, at single-base resolution, both sites of RT misincorporation / insertion / deletion and sites of RT drop-off. The extent of RT drop-off at a given position is quantified by comparing the number of sequencing reads corresponding to the drop-off product with the number of sequencing reads corresponding to the full-length product.
[0287] Example 22 - Evaluation of non-templated base addition (predictive) The non-templated addition of bases to the 5'-end of the cDNA products is evaluated by next-generation sequencing. Primer extension reactions containing RT and RNA templates derived from a cell-free expression system are performed as described above. A systematic analysis of different RNA template lengths and sequence motifs at the 5'-end is examined. Adapter sequences are ligated to the 3'-ends of the resulting cDNA products by T4 ligase without bias, resulting in the capture of all cDNA products despite the potential heterogeneous nature of their 3'-ends. The ligated products are then PCR amplified and library-prepared for next-generation sequencing. Comparison of the predicted full-length cDNA reference sequence with experimentally produced cDNA sequences that are longer than full-length allows the identification of both the type and number of base additions to the 5'-end that were not templated by RNA.
[0288] Example 23 - Determination of 5' and 3' UTR parameters for activity and processivity of R2, non-LTR, and similar systems (predictive) Proteins of interest are purified via Twin-strep tags after IPTG-induced overexpression in E. coli. Purified proteins are tested against 1 kb and 4 kb cargo flanked by the 3'UTR identified from their native context and the 5'UTR an additional 400 bp beyond the start codon. The effect of the 5' and 3' flanking sequences on activity is assayed via qPCR on sections near the end of the template to determine whether cargo with these native features produces superior results.
[0289] Example 24 - RT cDNA synthesis activity can be utilized for multiple purposes (predictive) RNA-dependent processes are important in biology, such as expression, processing, modification, and half-life. Quality control procedures in biotechnology performed on RNA utilize the conversion of RNA to cDNA. Thus, multiple RTs have been used for the generation of cDNA libraries over the years. Commercially available RTs used for these purposes include MMLV RT, AMV RT, and GsI-IIC RT (TGIRT). The first two represent retroviral RTs, while the latter is a GII intron-derived RT. GII intron-derived RTs, as well as non-LTR-derived RTs, exhibit several advantages over their retroviral counterparts. For example, they are more processive and read through structural and modified RNAs. Structural or modified RNAs may not be optimal substrates for retroviral RTs, as they generate premature termination products that can be misinterpreted as RNA fragments. In addition, the switch-templated ability of some RTs can be exploited for early adapter addition, making the adapter ligation procedure less critical during library preparation. Thus, highly processive RTs are suitable for the generation of libraries with complex RNAs. Furthermore, some highly processive RTs are generally smaller than currently used retroviral RTs, making their production and associated downstream processes easier. Some of the novel RTs described herein are superior to commercially available TGIRT enzymes, some with over 10-fold their cDNA synthesis activity. Thus, many of these novel RTs are very promising for their commercial use in cDNA synthesis kits.
[0290] While preferred embodiments of the present invention have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided for illustrative purposes only. The present invention is not intended to be limited by the specific examples provided herein. Although the present invention has been described with reference to the foregoing description, the description and explanation of the embodiments herein are not intended to be construed in a limiting sense. Numerous variations, changes, and substitutions will occur to those skilled in the art without departing from the present invention. Furthermore, it will be understood that all aspects of the present invention are not limited to the specific depictions, configurations, or relative proportions described herein, which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present invention described herein may be used in the practice of the present invention. It is therefore contemplated that the present invention will encompass any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the present invention and its methods and structures within the scope of these claims and their equivalents encompassed thereby.
[0291] [Table 3-1]
[0292] [Table 3-2]
[0293] [Table 3-3]
[0294] [Table 3-4]
[0295] [Table 3-5]
[0296]
Table 3-6
[0297]
Table 3-7
[0298]
Table 3-8
[0299]
Table 3-9
[0300]
Table 3-10
[0301]
Table 3-11
[0302]
Table 3-12
[0303]
Table 3-13
[0304]
Table 3-14
[0305]
Table 3-15
[0306]
Table 3-16
[0307]
Table 3-17
[0308]
Table 3-18
[0309]
Table 3-19
[0310]
Table 3-20
[0311]
Table 3-21
[0312]
Table 3-22
[0313]
Table 3-23
[0314]
Table 3-24
[0315]
Table 3-25
[0316]
Table 3-26
[0317]
Table 3-27
[0318]
Table 3-28
[0319]
Table 3-29
[0320]
Table 3-30
[0321]
Table 3-31
[0322]
Table 3-32
[0323]
Table 3-33
[0324]
Table 3-34
[0325]
Table 3-35
[0326]
Table 3-36
[0327]
Table 3-37
[0328]
Table 3-38
[0329]
Table 3-39
[0330]
Table 3-40
[0331]
Table 3-41
[0332]
Table 3-42
[0333]
Table 3-43
[0334]
Table 3-44
[0335]
Table 3-45
[0336]
Table 3-46
[0337] [Table 3-47]
[0338] [Table 3-48]
[0339] [Table 3-49]
[0340] Embodiment The following embodiments are not intended to be limiting in any way. Embodiment 1. An engineered retrotransposase system comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, the cargo nucleotide sequence being configured to interact with a retrotransposase; (b) a retrotransposase, An engineered retrotransposase system, wherein the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and wherein the retrotransposase is derived from an uncultured microorganism. Embodiment 2. The engineered retrotransposase system of embodiment embodiment 1, wherein the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. Embodiment 3. The engineered retrotransposase system of embodiment embodiment 1 or embodiment embodiment 2, wherein said retrotransposase comprises a reverse transcriptase domain. Embodiment 4. The engineered retrotransposase system according to any one of the embodiments of embodiment 1 to embodiment 3, wherein the retrotransposase further comprises one or more zinc finger domains. Embodiment 5. The engineered retrotransposase system according to any one of the embodiments of embodiment 1 to embodiment 4, wherein the retrotransposase further comprises an endonuclease domain. Embodiment 6. The engineered retrotransposase system according to any one of the embodiments of embodiment 1 to embodiment 5, wherein the retrotransposase has less than 80% sequence identity to a documented retrotransposase. Embodiment 7. The engineered retrotransposase system according to any one of the embodiments of embodiment 1 to embodiment 6, wherein the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). Embodiment 8. The engineered retrotransposase system of any one of the embodiments of embodiment 1 to embodiment 7, wherein the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. Embodiment 9. The engineered retrotransposase system of any one of the embodiments of embodiment 1 to embodiment 8, wherein the retrotransposase comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the retrotransposase. Embodiment 10. The engineered retrotransposase system according to any one of the embodiments of embodiment 1 to embodiment 9, wherein the NLS comprises a sequence that is at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 896 to 911. Embodiment 11. The engineered retrotransposase system according to any one of the embodiments 1 to 10, wherein the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Embodiment 12. The engineered retrotransposase system of embodiment embodiment 11, wherein said sequence identity is determined by said BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment. Embodiment 13. An engineered retrotransposase system comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, the cargo nucleotide sequence being configured to interact with a retrotransposase; (b) a retrotransposase, An engineered retrotransposase system, wherein the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. Embodiment 14 The engineered retrotransposase system of embodiment embodiment 13, wherein the retrotransposase is derived from an uncultured microorganism. Embodiment 15. The engineered retrotransposase system of embodiment embodiment 13 or embodiment embodiment 14, wherein the retrotransposase comprises a reverse transcriptase domain. Embodiment 16. The engineered retrotransposase system according to any one of embodiments 13 to 15, wherein the retrotransposase further comprises one or more zinc finger domains. Embodiment 17. The engineered retrotransposase system according to any one of embodiments 13 to 16, wherein the retrotransposase further comprises an endonuclease domain. Embodiment 18. An engineered retrotransposase system according to any one of the embodiments of embodiments 13 to 17, wherein the retrotransposase has less than 80% sequence identity to a documented retrotransposase. Embodiment 19. The engineered retrotransposase system according to any one of the embodiments of embodiment 13 to embodiment 18, wherein the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). Embodiment 20. An engineered retrotransposase system according to any one of the embodiments of embodiment 13 to embodiment 19, wherein the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. Embodiment 21. The engineered retrotransposase system according to any one of the embodiments 13 to 20, wherein the sequence identity is determined by BLASTP, CLUSTALW, MUSCLE, MAFFT, or CLUSTALW using the parameters of the Smith-Waterman homology search algorithm. Embodiment 22. The engineered retrotransposase system of embodiment embodiment 21, wherein said sequence identity is determined by said BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, with a conditional composition score matrix adjustment. Embodiment 23. A deoxyribonucleic acid polynucleotide encoding the engineered retrotransposase system according to any one of the embodiments of embodiments 1 to 22. Embodiment 24. A nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, the nucleic acid encoding a retrotransposase, the retrotransposase being derived from an uncultured microorganism, and the organism is not the uncultured microorganism. Embodiment 25. The nucleic acid of embodiment 24, wherein the retrotransposase comprises a variant having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895. Embodiment 26. The nucleic acid of embodiment embodiment 24 or embodiment embodiment 25, wherein the retrotransposase comprises a sequence encoding one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the retrotransposase. Embodiment 27. The nucleic acid of embodiment 26, wherein the NLS comprises a sequence selected from SEQ ID NOs: 896-911. Embodiment 28. The nucleic acid described in embodiment 26 or embodiment 27, wherein the NLS comprises SEQ ID NO: 897. Embodiment 29. The nucleic acid of embodiment embodiment 28, wherein the NLS is proximal to the N-terminus of the retrotransposase. Embodiment 30. The nucleic acid described in embodiment 26 or embodiment 27, wherein the NLS comprises SEQ ID NO: 896. Embodiment 31. The nucleic acid of embodiment embodiment 30, wherein the NLS is proximal to the C-terminus of the retrotransposase. Embodiment 32. The nucleic acid according to any one of embodiments 24 to 31, wherein the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human. Embodiment 33. A vector comprising a nucleic acid according to any one of embodiments 24 to 32. Embodiment 34. The vector described in embodiment embodiment 33, further comprising a nucleic acid encoding a cargo nucleotide sequence configured to form a complex with the retrotransposase. Embodiment 35. The vector described in embodiment embodiment 33 or embodiment embodiment 34, wherein the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV) derived virion, or a lentivirus. Embodiment 36. A cell comprising the vector according to any one of embodiments 33 to 35. Embodiment 37. A method for producing a retrotransposase, comprising culturing the cell of embodiment embodiment 36. Embodiment 38. A method for disrupting, binding, nicking, cleaving, marking, or modifying a double-stranded deoxyribonucleic acid polynucleotide comprising a target nucleic acid locus, comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide comprising the target nucleic acid locus with a transposase configured to transpose a cargo nucleotide sequence to the target nucleic acid locus; (b) The method, wherein the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1 to 29, 393 to 735, or 799 to 895. Embodiment 39 The method of embodiment embodiment 38, wherein the retrotransposase is derived from an uncultured microorganism. Embodiment 40 The engineered retrotransposase system described in embodiment embodiment 38 or embodiment embodiment 39, wherein the retrotransposase comprises a reverse transcriptase domain. Embodiment 41. The engineered retrotransposase system according to any one of embodiments 38 to 40, wherein the retrotransposase further comprises one or more zinc finger domains. Embodiment 42. The engineered retrotransposase system according to any one of the embodiments 38 to 41, wherein the retrotransposase further comprises an endonuclease domain. Embodiment 43. The method according to any one of embodiments 38 to 42, wherein the retrotransposase has less than 80% sequence identity to a documented retrotransposase. Embodiment 44. The engineered retrotransposase system according to any one of the embodiments of embodiment 38 to embodiment 43, wherein the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). Embodiment 45. The method according to any one of embodiments 38 to 44, wherein the double-stranded deoxyribonucleic acid polynucleotide is translocated via a ribonucleic acid polynucleotide intermediate. Embodiment 46. The method according to any one of embodiments 38 to 45, wherein the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. Embodiment 47. A method for disrupting or modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus the engineered retrotransposase system described in any one of the embodiments of embodiments 1 to 22, wherein the retrotransposase is configured to transpose a cargo nucleotide sequence to the target nucleic acid locus, and the complex is configured such that upon binding of the complex to the target nucleic acid locus, the complex modifies the target nucleic acid locus. Embodiment 48. The method of embodiment 47, wherein modifying the target nucleic acid locus comprises binding, nicking, cleaving, marking, modifying, or translocating the target nucleic acid locus. Embodiment 49. The method of embodiment 47 or embodiment 48, wherein the target nucleic acid locus comprises deoxyribonucleic acid (DNA). Embodiment 50. The method of embodiment 49, wherein the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. Embodiment 51. The method of any one of embodiments 47 to 50, wherein the target nucleic acid locus is in vitro. Embodiment 52. The method of any one of embodiments 47 to 50, wherein the target nucleic acid locus is intracellular. Embodiment 53. The method of embodiment 52, wherein the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. Embodiment 54 The method of embodiment 52 or embodiment 53, wherein the cells are primary cells. Embodiment 55 The method of embodiment embodiment 54, wherein the primary cells are T cells. Embodiment 56 The method of embodiment embodiment 54, wherein the primary cells are hematopoietic stem cells (HSCs). Embodiment 57. The method of any one of embodiments 47 to 56, wherein delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid of any one of embodiments 24 to 32 or a vector of any of embodiments 33 to 35. Embodiment 58. The method of any one of embodiments 47 to 57, wherein delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding the retrotransposase. Embodiment 59. The method of embodiment embodiment 58, wherein the nucleic acid comprises a promoter to which the open reading frame encoding the retrotransposase is operably linked. Embodiment 60. The method of any one of embodiments 47 to 59, wherein delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a capped mRNA containing the open reading frame encoding the retrotransposase. Embodiment 61. The method of any one of embodiments 47 to 60, wherein delivering the engineered retrotransposase system to the target nucleic acid locus comprises delivering a translated polypeptide. Embodiment 62. The method of any one of embodiments 47 to 61, wherein the retrotransposase does not induce cleavage at or proximal to the target nucleic acid locus. Embodiment 63. A host cell comprising an open reading frame encoding a heterologous retrotransposase having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, or 799-895, or a variant thereof. Embodiment 64. The host cell of embodiment 63, wherein the host cell is an E. coli cell. Embodiment 65. The host cell of embodiment embodiment 64, wherein the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. Embodiment 66. The host cell of embodiment embodiment 64 or embodiment embodiment 65, wherein the E. coli cell has an ompT lon genotype. 67. The open reading frame is selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, BAD The host cell according to any one of the embodiments 63 to 66, wherein the host cell is operably linked to a promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. Embodiment 68. A host cell described in any one of embodiments 63 to 67, wherein the open reading frame comprises a sequence encoding an affinity tag linked in-frame to a sequence encoding the retrotransposase. Embodiment 69. The host cell of embodiment embodiment 68, wherein the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. Embodiment 70. The host cell of embodiment 69, wherein the IMAC tag is a polyhistidine tag. Embodiment 71. The host cell of embodiment embodiment 68, wherein the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. Embodiment 72. A host cell described in any one of embodiments 68 to 71, wherein the affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. Embodiment 73. The host cell of embodiment 72, wherein the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. Embodiment 74. A host cell according to any one of embodiments 63 to 73, wherein the open reading frame is codon-optimized for expression in the host cell. Embodiment 75. A host cell according to any one of embodiments 63 to 74, wherein the open reading frame is provided on a vector. Embodiment 76. A host cell according to any one of embodiments 63 to 74, wherein the open reading frame is integrated into the genome of the host cell. Embodiment 77. A culture comprising a host cell according to any one of the embodiments of embodiments 63 to 76 in a suitable liquid medium. Embodiment 78. A method for producing a retrotransposase, comprising culturing a host cell according to any one of the embodiments of embodiments 63 to 76 in a suitable growth medium. Embodiment 79. The method of embodiment embodiment 78, further comprising inducing expression of the retrotransposase by adding additional chemical agents or increased amounts of nutrients. Embodiment 80. The method of embodiment 79, wherein the additional chemical agent or increased amount of nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. Embodiment 81. The method according to any one of embodiments 78 to 80, further comprising isolating the host cells after said culturing and lysing the host cells to produce a protein extract. Embodiment 82 The method according to embodiment 81, further comprising subjecting the protein extract to IMAC, or ion affinity chromatography. Embodiment 83. The method of embodiment 82, wherein the open reading frame comprises a sequence encoding an IMAC affinity tag linked in-frame to a sequence encoding the retrotransposase. Embodiment 84 The method of embodiment embodiment 83, wherein said IMAC affinity tag is linked in frame to said sequence encoding said retrotransposase via a linker sequence encoding a protease cleavage site. Embodiment 85. The method of embodiment 84, wherein the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. Embodiment 86 The method of embodiment embodiment 84 or embodiment embodiment 85, further comprising cleaving the IMAC affinity tag by contacting the retrotransposase with a protease corresponding to the protease cleavage site. Embodiment 87. The method according to embodiment embodiment 86, further comprising performing subtractive IMAC affinity chromatography to remove said affinity tag from the composition comprising said retrotransposase. Embodiment 88. A method of disrupting a genetic locus in a cell, comprising: (a) a double-stranded nucleic acid comprising a heterologous engineered cargo nucleotide sequence, the cargo nucleotide sequence being configured to interact with a retrotransposase; (b) a retrotransposase, the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus; The retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1 to 29, 393 to 735, or 799 to 895, or a variant thereof; and A method comprising contacting a composition comprising a retrotransposase, wherein the retrotransposase has at least equivalent transposition activity to a demonstrated retrotransposase in the cell. Embodiment 89. The method of embodiment embodiment 88, wherein the transposition activity is measured in vitro by introducing the retrotransposase into a cell containing the target nucleic acid locus and detecting transposition of the target nucleic acid locus in the cell. Embodiment 90 The method of embodiment embodiment 88 or embodiment embodiment 89, wherein the composition comprises 20 pmoles or less of the retrotransposase. Embodiment 91. The method of embodiment 90, wherein the composition comprises 1 pmol or less of the retrotransposase.
Claims
1. 1. An engineered retrotransposase system comprising: (a) a ribonucleic acid (RNA) comprising a heterologous engineered cargo nucleotide sequence, wherein the RNA is configured to interact with a retrotransposase; (b) the retrotransposase, (i) the retrotransposase is configured to transpose the heterologous engineered cargo nucleotide sequence into a target nucleic acid locus; and (ii) a retrotransposase, wherein the retrotransposase comprises a reverse transcriptase (RT) domain or an endonuclease domain.
2. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase further comprises a zinc (Zn)-binding ribbon motif having at least 80% sequence identity to a Zn-binding ribbon motif of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895.
3. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895.
4. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to an RT domain or endonuclease domain of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895.
5. 2. The engineered retrotransposase system of Claim 1, wherein said RNA comprises a sequence 5' to said heterologous engineered cargo nucleotide sequence or a sequence 3' to said heterologous engineered cargo nucleotide sequence, wherein said sequence 5' to said heterologous engineered cargo nucleotide sequence or said sequence 3' to said heterologous engineered cargo nucleotide sequence has at least 80% sequence identity to an RNA cognate of any one of SEQ ID NOs: 769, 770, 779, 780, 761-768, 771-778, 781-798, its complement, or its reverse complement.
6. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
3.
7. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
572.
8. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
822.
9. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
8.
10. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
836.
11. The engineered retrotransposase system of claim 3, wherein the retrotransposase comprises a sequence having at least 80% sequence identity to SEQ ID NO:
602.
12. An engineered retrotransposase system described in any one of claims 1 to 11, further comprising (c) a double-stranded deoxyribonucleic acid (DNA) sequence comprising the target nucleic acid locus.
13. The engineered retrotransposase system of any one of claims 1 to 11, wherein the heterologous engineered cargo nucleotide sequence comprises an expression cassette.
14. The engineered retrotransposase system of any one of claims 1 to 11, wherein the RNA is in vitro transcribed RNA.
15. An engineered deoxyribonucleic acid (DNA) sequence encoding an engineered retrotransposase system according to any one of claims 1 to 11.
16. A method for modifying a target nucleic acid locus, the method comprising contacting the target nucleic acid locus with an engineered retrotransposase system described in any one of claims 1 to 11.
17. An engineered deoxyribonucleic acid (DNA) sequence comprising: (a) a 5′ sequence capable of encoding a ribonucleic acid (RNA) sequence configured to interact with a retrotransposase; (b) a heterologous cargo sequence; and (c) a sequence encoding the retrotransposase, wherein the retrotransposase comprises a reverse transcriptase (RT) domain or an endonuclease domain, and the RT domain or the endonuclease domain comprises a sequence having at least 80% sequence identity to the RT domain or the endonuclease domain of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895; (d) a 3′ sequence capable of encoding an RNA sequence configured to interact with said retrotransposase.
18. The engineered DNA sequence of claim 17, wherein the 5' sequence or the 3' sequence comprises a sequence having at least 80% sequence identity to an RNA cognate of any one of SEQ ID NOs: 769, 770, 779, 780, 761-768, 771-778, 781-798, its complement, or its reverse complement.
19. A method for synthesizing complementary deoxyribonucleic acid (cDNA), comprising: (a) providing a ribonucleic acid (RNA) molecule as a template for cDNA synthesis; (b) providing a nucleic acid primer for priming cDNA synthesis from said RNA molecule; (c) using said template and a reverse transcriptase to synthesize cDNA primed by said nucleic acid primer.
20. The method of claim 19, wherein the reverse transcriptase comprises a sequence having at least 80% sequence identity to the RT domain or endonuclease domain of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895.
21. The method described in claim 20, wherein the nucleic acid primer comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides.
22. A protein comprising a reverse transcriptase (RT) domain or an endonuclease domain, wherein the protein comprises a sequence having at least 80% sequence identity to the RT domain or endonuclease domain of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895, and wherein the sequence is fused or linked at the N-terminus or C-terminus to a non-retrotransposase domain or an affinity tag.
23. The protein described in claim 22, comprising a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895.
24. A nucleic acid encoding an open reading frame, wherein the open reading frame encodes a reverse transcriptase (RT) domain or an endonuclease domain, and the RT domain or the endonuclease domain comprises a sequence having at least 80% sequence identity to the RT domain or the endonuclease domain of any one of SEQ ID NOs: 3, 572, 822, 8, 836, 602, 1, 2, 4-7, 9-29, 393-571, 573-601, 603-735, 799-821, 823-835, or 837-895, and (a) the open reading frame is optimized for expression in an organism, and the organism is different from the source of the RT domain or the endonuclease domain, or (b) the open reading frame comprises a sequence encoding an affinity tag.
25. The nucleic acid described in claim 24, further encoding a retrotransposase.
26. A cell comprising an engineered retrotransposase system described in any one of claims 1 to 11.
27. The cell described in claim 26, which is a eukaryotic cell.
28. A cell comprising an engineered DNA sequence described in claim 17 or 18.
29. The cell described in claim 28, which is a eukaryotic cell.
30. A cell comprising the protein described in claim 22 or 23.
31. The cell described in claim 30, which is a eukaryotic cell.
32. A cell containing the nucleic acid described in claim 24 or 25.
33. The cell described in claim 32, which is a eukaryotic cell.