Retrotransposon compositions and methods of use
The engineered retrotransposase system addresses the underutilization of transposable elements in DNA manipulation by providing a precise and efficient method for nucleic acid modifications, enabling applications like cDNA synthesis and genetic manipulation.
Patent Information
- Application Number
- JP2025533099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-10
- Filing Date
- 2023-12-08
- Publication Date
- 2026-01-06
AI Technical Summary
Existing technologies have not fully exploited the potential of transposable elements for DNA manipulation and gene editing applications, despite their recognized importance in gene function and evolution.
An engineered retrotransposase system comprising a double-stranded nucleic acid with a cargo nucleotide sequence and a retrotransposase that can transpose the cargo sequence to a target nucleic acid, with the retrotransposase having specific amino acid sequences with varying degrees of identity to provided SEQ ID NOs, enabling precise nucleic acid modifications.
The engineered retrotransposase system facilitates efficient and targeted nucleic acid modifications, including binding, nicking, and cleavage, in various cell types, supporting applications such as cDNA synthesis and genetic manipulation.
Smart Images

Figure 2026500188000001_ABST
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 386,865, filed December 9, 2022, U.S. Provisional Patent Application No. 63 / 489,154, filed March 8, 2023, U.S. Provisional Patent Application No. 63 / 491,939, filed March 23, 2023, and U.S. Provisional Patent Application No. 63 / 501,373, filed May 10, 2023, each of which is incorporated herein by reference in its entirety. [Background technology]
[0002] Transposable elements are mobile DNA sequences that play important roles in gene function and evolution. Transposable elements are found in almost all types of life, but their prevalence varies among organisms, and the majority of eukaryotic genomes encode transposable elements. Summary of the Invention
[0003] Although fundamental research on transposable elements was conducted in the 1940s, their potential utility in DNA manipulation and gene editing applications has only recently been recognized.
[0004] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises an amino acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase is encoded by a nucleic acid having at least 75% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the double-stranded nucleic acid comprises a 5' recognition sequence comprising a GG nucleotide sequence and a 3' recognition sequence comprising a TGAC nucleotide sequence. In some embodiments, the 5' and 3' recognition sequences are configured to interact with a retrotransposase. In some embodiments, the double-stranded nucleic acid comprising a cargo nucleotide sequence is RNA. In some embodiments, the RNA is in vitro transcribed RNA.In some embodiments, the RNA comprises a sequence 5' to the cargo sequence or a sequence 3' to the cargo sequence that has at least 80% sequence identity to an RNA homologue, its complement, or its reverse complement of any one of SEQ ID NOs: 761-798, 2161-2164, and 2211-2257.
[0005] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1535-1536, 1542-1543, 1611-1623, 1663-1691, and 1786-1806.
[0006] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to SEQ ID NO:402 or SEQ ID NO:895.
[0007] In certain embodiments, the present disclosure describes an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 388.
[0008] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 389-392 and 1504-1507.
[0009] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 427-439.
[0010] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 356-373, 964-981, and 1003-1019.
[0011] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 66-173, 740-756, 1521-1534, 1539-1541, 1624-1637, 1645-1662, and 1701-1782.
[0012] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 308-309 and 324-325.
[0013] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 310-312, 326-328, 1556-1557, and 1569-1570.
[0014] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 616 or SEQ ID NO: 617. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 313-314 and 329-330.
[0015] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 315-319 and 331-335.
[0016] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 623. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 320 or SEQ ID NO: 336.
[0017] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 321-323, 337-339, and 1785.
[0018] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 321-323, 337-339, and 1785.
[0019] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 627-673, 1039-1475, and 2011-2026. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 174-187 and 1508-1520.
[0020] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 188-197.
[0021] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 198-207.
[0022] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 208-225 and 757-759.
[0023] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 226-235.
[0024] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 236-245 and 759-760.
[0025] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 246-255.
[0026] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 256-277, 1638-1644, and 1693-1700.
[0027] The present disclosure, in certain embodiments, describes an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 278-297.
[0028] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 298-307.
[0029] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1558-1567, 1571-1580, and 1783-1784.
[0030] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 1692.
[0031] The present disclosure describes, in certain embodiments, an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; and (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, the retrotransposase comprising an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 1568 or SEQ ID NO: 1594. In some embodiments, the retrotransposase comprises one or more nuclear localization sequences (NLSs) proximal to the N- or C-terminus of the retrotransposase. In some embodiments, the NLS comprises a sequence at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 1477-1492. In some embodiments, the NLS comprises SEQ ID NO: 1478. In some embodiments, the NLS is proximal to the N-terminus of the retrotransposase. In some embodiments, the NLS comprises SEQ ID NO: 1477. In some embodiments, the NLS is proximal to the C-terminus of the retrotransposase.
[0032] The present disclosure describes, in certain embodiments, polypeptides comprising a reverse transcriptase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266 fused at the N-terminus or C-terminus to a non-retrotransposase domain or affinity tag. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA-binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
[0033] The present disclosure describes, in certain embodiments, nucleic acids encoding the engineered retrotransposase systems described in this disclosure or the polypeptides described in this disclosure.
[0034] In certain embodiments, the present disclosure describes a method for modifying a target nucleic acid sequence, the method comprising contacting the target nucleic acid sequence with an engineered nuclease system described herein. In some embodiments, modifying the target nucleic acid sequence comprises binding, nicking, or cleaving the target nucleic acid sequence. In some embodiments, the target nucleic acid sequence comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA. In some embodiments, the target nucleic acid sequence comprises deoxyribonucleic acid (DNA). In some embodiments, the modification is in vitro. In some embodiments, the modification is in vivo. In some embodiments, the modification is ex vivo.
[0035] In certain embodiments, the present disclosure describes a method for modifying a target nucleic acid sequence in a mammalian cell, comprising contacting the mammalian cell with an engineered nuclease system described in the present disclosure.
[0036] As described herein, certain embodiments describe a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide for initiating cDNA synthesis from the RNA molecule; and (c) synthesizing cDNA from the template initiated by the primer oligonucleotide using a reverse transcriptase comprising a sequence having at least 80% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOS: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266, or a variant thereof. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides.
[0037] In certain embodiments, the present disclosure describes a vector that comprises the nucleic acid described in the present disclosure.In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0038] The present disclosure describes, in certain embodiments, cells comprising the engineered nuclease system or polypeptide described herein. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell. In some embodiments, the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof. In some embodiments, the cells are engineered cells. In some embodiments, the cells are stable cells.
[0039] In some aspects, the disclosure provides a method for the production of a retrotransposase comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus; and (ii) the retrotransposase has at least about 80% affinity to a reverse transcriptase (RT) or endonuclease domain of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, and 1546-1553. and a retrotransposase comprising an RT domain, an endonuclease domain, and a sequence having at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to a target gene, or a variant thereof. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif. In some embodiments, the retrotransposase further comprises a conserved CX [2-3]In some embodiments, the retrotransposase further comprises a C Zn finger motif. In some embodiments, the retrotransposase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 6, 7, 8, 14, and 402. In some embodiments, the system further comprises (c) a double-stranded DNA sequence comprising a target nucleic acid locus. In some embodiments, the double-stranded DNA sequence comprises a 5' recognition sequence and a 3' recognition sequence configured to interact with the retrotransposase, wherein the 5' recognition sequence comprises a GG nucleotide sequence and the 3' recognition sequence comprises a TGAC nucleotide sequence. In some embodiments, the RNA is in vitro transcribed RNA. In some embodiments, the RNA comprises a sequence 5' to or 3' to a cargo sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RNA cognate of any one of SEQ ID NOs: 761-798, 2161-2164, 2211-2257, its complement, or its reverse complement. In some embodiments, the RNA comprises a sequence encoding a retrotransposase. In some embodiments, the heterologous engineered cargo nucleotide sequence comprises an expression cassette.
[0040] In some embodiments, the disclosure provides a method for the preparation of a retrotransposase comprising: (a) a 5′ sequence capable of encoding an RNA sequence configured to interact with a retrotransposase; (b) a heterologous cargo sequence; and (c) a sequence encoding a retrotransposase configured to interact with an RNA cognate of the 5′ sequence, wherein the retrotransposase has a sequence encoding at least about 80%, at least about 81%, at least about 82%, or at least about 83% of the reverse transcriptase (RT) or endonuclease domain of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. Provided is an engineered DNA sequence comprising: a sequence comprising an RT domain or an endonuclease domain comprising a sequence having about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof; and (d) a 3' sequence capable of encoding an RNA sequence configured to interact with a retrotransposase. In some embodiments, the retrotransposase further comprises any one of the Zn-binding ribbon motifs of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif. In some embodiments, the retrotransposase further comprises a CX [2-3]In some embodiments, the retrotransposase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 3, 6, 7, 8, 14, or 402. In some embodiments, the 5' or 3' sequence comprises a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to an RNA cognate of any one of SEQ ID NOs: 761-798, 2161-2164, and 2211-2257, its complement, or its reverse complement.
[0041] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) providing a primer oligonucleotide that is at least about 80%, at least about 81%, at least about 82%, or at least about 83% complementary to a reverse transcriptase domain of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. and synthesizing a primer-primed cDNA from a template using a reverse transcriptase comprising a sequence having about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, cDNA synthesis comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, divalent metal ions, Mg 2+ , or Mn 2+ Further includes:
[0042] In some embodiments, the present disclosure provides a method for identifying a sequence identical to or similar to the sequence of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266, comprising at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, or at least about 89% of the sequence identical to or similar to the sequence of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
[0003] The present invention provides proteins comprising a reverse transcriptase domain comprising a sequence having about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266, or a variant thereof, wherein the sequence is fused at the N- or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
[0043] In some aspects, the disclosure provides nucleic acids encoding any of the polypeptides described herein.
[0044] In some embodiments, the disclosure provides nucleic acids encoding an open reading frame, wherein the open reading frame is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 1109%, at least about 1111%, at least about 1120%, at least about 1130%, at least about 1140%, at least about 1150%, at least about 1160%, at least about 1170%, at least about 1180%, at least about 1190%, at least about 1211%, at least about 1221%, at least about 1230%, at least about 1240%, at least about 1250%, at least about 1260%, at least about 1270%, at least about 1280%, at least about 1290%, at least about 1291%, at least about 1292%, at least about 1293%, at least about 1294%, at least about 1295%, at least about 1295%, at least about 1296%, at least about 1297%, at least about 1298%, at least about 1299%, at least about 1300%, The open reading frame encodes an RT or endonuclease domain, or a variant thereof, having at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, and (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF includes a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT or endonuclease domain of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266, or a variant thereof.
[0045] In some embodiments, the disclosure provides a method for the production of a retrotransposase comprising: (a) an RNA comprising a heterologous engineered cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus; and (ii) the retrotransposase has a nucleotide sequence that is at least about 80%, at least about 81%, at least about 82%, or at least about 83% similar to the reverse transcriptase (RT) or endonuclease domain of SEQ ID NO: 402 or 895. and a retrotransposase comprising an RT domain or endonuclease domain comprising a sequence having at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a conserved CX [2-3] In some embodiments, the system further comprises (c) a double-stranded DNA sequence comprising a target locus. In some embodiments, the RNA is in vitro transcribed RNA. In some embodiments, the RNA comprises a sequence encoding a retrotransposase.
[0046] In some aspects, the present disclosure provides a method for the preparation of a retrotransposase comprising: (a) a 5′ sequence capable of encoding an RNA sequence configured to interact with a retrotransposase; (b) a heterologous cargo sequence; and (c) a sequence encoding a retrotransposase configured to interact with an RNA cognate of the 5′ sequence, wherein the retrotransposase has a nucleotide sequence that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 1109%, at least about 1111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 121%, at least about 122%, at least about 123%, at least about 124%, at least about 125%, at least about 126%, at least about 127%, at least about 128%, at least about 129%, at least about 130%, at least about 131%, at least about 132%, at least about 133%, at least about 134%,
[0013] Provided are engineered DNA sequences comprising: (a) a sequence comprising an RT domain, an endonuclease domain, and (d) a 3' sequence capable of encoding an RNA sequence configured to interact with a retrotransposase; wherein the RT domain comprises a sequence having 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity; and (e) a 3' sequence capable of encoding an RNA sequence configured to interact with a retrotransposase. In some embodiments, the retrotransposase further comprises any of the Zn-binding ribbon motifs of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase further comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif of SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises the conserved CX of SEQ ID NO: 402 or 895. [2-3] C further contains a Zn finger motif.
[0047] In some aspects, the disclosure provides methods for synthesizing complementary DNA (cDNA), the methods comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing cDNA initiated by the primer oligonucleotide from the template using a reverse transcriptase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of SEQ ID NO: 402 or 895. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, divalent metal ions, Mg 2+ , or Mn 2+ Further includes:
[0048] In some aspects, the present disclosure provides polypeptides comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of SEQ ID NO: 402 or 895, wherein the sequence is fused at the N- or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to SEQ ID NO: 402 or 895, or a variant thereof. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
[0049] In some aspects, the disclosure provides nucleic acids encoding an open reading frame, wherein the open reading frame encodes an RT or endonuclease domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT or endonuclease domain of SEQ ID NO: 402 or 895, and wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least 80% sequence identity to SEQ ID NO:402 or 895.
[0050] In some aspects, the disclosure provides methods for synthesizing complementary DNA (cDNA), the methods comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing cDNA initiated by the primer oligonucleotide from the template using a reverse transcriptase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs:555-728. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, and 607. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, and 608. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the primer oligonucleotide comprises at least one phosphorothioate linkage.In some embodiments, cDNA synthesis involves incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture contains dNTPs, a reaction buffer, a divalent metal ion, Mg. 2+ , or Mn 2+ Further includes:
[0051] In some aspects, the present disclosure provides polypeptides comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs:555-728, wherein the sequence is fused at the N- or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, and 607. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, and 608. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA-binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the protein comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 30-32, 40-50, 740-756, 757-760. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-558, 561-567, 569, 570, 575.
[0052] In some aspects, the disclosure provides nucleic acids encoding an open reading frame, wherein the open reading frame encodes an RT or endonuclease domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT or endonuclease domain of any one of SEQ ID NOs:555-728, and wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes a retrotransposase comprising a sequence having at least 80% sequence identity to the RT or endonuclease domain of any of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608, 561, 562, 564, 565, 568, 571, 573, 576-579, 583, 590, 591, 594, 598, 601, 606, and 607. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 555-560, 563, 564, 566, 567, 569, 572, 574, 580-582, 584-588, 592, 593, 596, 602, 604, 605, 608.
[0053] In some aspects, the disclosure provides nucleic acids comprising a sequence comprising an open reading frame (ORF) comprising a sequence encoding a reverse transcriptase domain or maturase domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain or maturase domain of any one of SEQ ID NOs:729-733, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the ORF encodes a protein having at least 80% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the ORF is optimized for expression in a bacterial organism, or the organism is E. coli. In some embodiments, the ORF is optimized for expression in a mammalian organism, or the organism is a primate organism. In some embodiments, the primate organism is Homo sapiens. In some embodiments, the ORF comprises an affinity tag operably linked to a sequence encoding a reverse transcriptase domain or a maturase domain, wherein the ORF has at least 80% sequence identity to any one of SEQ ID NOs: 298-302. In some embodiments, the ORF comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 303-307. In some embodiments, the reverse transcriptase domain or maturase domain comprises the conserved Y[I / L]DD active site motif of any one of SEQ ID NOs: 729-733.
[0054] In some aspects, the disclosure provides methods for synthesizing complementary DNA (cDNA), the methods comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) synthesizing cDNA initiated by the primer oligonucleotide from the template using a reverse transcriptase comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs:440-554. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to any one of SEQ ID NO: 526. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, synthesis of cDNA comprises incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture comprises dNTPs, a reaction buffer, divalent metal ions, Mg 2+ , or Mn 2+ Further includes:
[0055] In some aspects, the present disclosure provides polypeptides comprising a reverse transcriptase domain comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the reverse transcriptase domain of any one of SEQ ID NOs: 440-554, wherein the sequence is fused at the N- or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 526. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA-binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the sequence is fused to an affinity tag at the N-terminus or C-terminus.
[0056] In some aspects, the disclosure provides nucleic acids encoding an open reading frame, wherein the open reading frame encodes an RT domain having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to the RT domain of any one of SEQ ID NOs:440-554, and wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF includes a sequence encoding an affinity tag. In some embodiments, the nucleic acid further encodes an RT having at least 80% sequence identity to any one of SEQ ID NOs: 518-522, 524-527, and 529-532. In some embodiments, the reverse transcriptase comprises a sequence having at least 80% sequence identity to SEQ ID NO: 526. In some embodiments, the open reading frame comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 356-373.
[0057] In some aspects, the disclosure provides a method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to initiate cDNA synthesis from the RNA molecule; and (c) providing a primer oligonucleotide that is at least about 80%, at least about 81%, at least about 82%, at least about 83%, or at least about 84% complementary to a reverse transcriptase domain of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, 627-673, 1544-1545, and 1555. and synthesizing a primer-primed cDNA from a template using a reverse transcriptase comprising a sequence having at least about 3%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof. In some embodiments, the reverse transcriptase domain comprises the conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, 627-673, 1544-1545, and 1555. In some embodiments, the reverse transcriptase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, and 633. In some embodiments, the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides. In some embodiments, the primer oligonucleotide comprises at least six consecutive nucleotides having at least 80% sequence identity to any one of SEQ ID NOs: 340-355, 1582-1594, and 1842-1849.In some embodiments, cDNA synthesis involves incubating a template RNA molecule, a primer oligonucleotide, and a reverse transcriptase in a reaction mixture under conditions suitable for elongating a DNA sequence from the RNA template. In some embodiments, the reaction mixture contains dNTPs, a reaction buffer, a divalent metal ion, Mg. 2+ , or Mn 2+ Further includes:
[0058] In some embodiments, the present disclosure provides a method for identifying a sequence encoding at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 1109%, at least about 1111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 1211%, at least about 1221%, at least about 1231%, at least about 1232%, at least about 1241%, at least about 1242%, at least about 1251%, at least about 1252%, at least about 1261%, at least about 1263%, at least about 1275%, at least about 1276%, at least about 1280%, at least about 1282%, at least about 1284%, at least about 1286%, at least about 1288%, at least about 1290%, at least about 1292%, at least about 1294%, at least about 12
[0013] Polypeptides are provided comprising a reverse transcriptase domain comprising a sequence having 9%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, or a variant thereof, wherein the sequence is fused at the N- or C-terminus to a non-retrotransposase domain or an affinity tag. In some embodiments, the reverse transcriptase domain comprises a conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, 627-673, 1544-1545, and 1555. In some embodiments, the reverse transcriptase domain comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, 633. In some embodiments, the non-retrotransposase domain is an RNA-binding protein domain. In some embodiments, the RNA-binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain. In some embodiments, the sequence is fused at the N-terminus or C-terminus to an affinity tag.
[0059] In some aspects, the disclosure provides nucleic acids encoding an open reading frame (ORF) optimized for expression in an organism, wherein the ORF has at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, at least about 100%, at least about 101%, at least about 102%, at least about 103%, at least about 104%, at least about 105%, at least about 106%, at least about 107%, at least about 108%, at least about 109%, at least about 1109%, at least about 1111%, at least about 112%, at least about 113%, at least about 114%, at least about 115%, at least about 116%, at least about 117%, at least about 118%, at least about 119%, at least about 120%, at least about 1211%, at least about 1221%, at least about 1231%, at least about 1232, at least about 1241, at least about 1242, at least about 1251, at least about 1252, at least about 1263, at least about 1264, at least about 1275, at least about 1276, at least about 1287, at least about 1288, at least about 1290, at least about 1292, at least about 1294, at least about 1296, at least about 1298 In some embodiments, the reverse transcriptase domain encodes an RT domain or variant thereof having at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity, wherein (a) the open reading frame is optimized for expression in an organism that is different from the source of the RT or endonuclease domain, or (b) the ORF comprises a sequence encoding an affinity tag. In some embodiments, the reverse transcriptase domain comprises the conserved xxDD, [F / Y]XDD, NAxxH, or VTG motif of any one of SEQ ID NOs: 609-610, 611-615, 616-617, 618-622, 623, 624-626, or 627-673, 1544-1545, or 1555. In some embodiments, the nucleic acid further encodes an RT having at least 80% sequence identity to any one of SEQ ID NOs: 612-613, 616-619, 622, 624, 627-630, 633. In some embodiments, the ORF comprises a sequence encoding an affinity tag.In some embodiments, the open reading frame comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 66-119, 174-180, 188-192, 198-202, 208-216, 226-230, 236-240, 246-250, 308-309, 310-312, 313-314, 315-319, 320, 321-323, 363-373, 1569-1570, 1571-1580, and 1581. In some embodiments, the organism is different from the source of the RT domain. In some embodiments, the ORF comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
[0060] In some aspects, the present disclosure provides synthetic oligonucleotides comprising at least six consecutive nucleotides having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 340-355, 1582-1594, and 1842-1849. In some embodiments, the synthetic oligonucleotide comprises DNA nucleotides. In some embodiments, the oligonucleotide further comprises at least one phosphorothioate linkage.
[0061] In some embodiments, the present disclosure provides a vector comprising a sequence having at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of SEQ ID NOs: 340-355, 1582-1594, and 1842-1849.
[0062] In some aspects, the disclosure provides a vector comprising any of the nucleic acids described herein.
[0063] In some aspects, the disclosure provides a host cell comprising any of the nucleic acids described herein. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen, or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the nucleic acid comprises an open reading frame (ORF) encoding a retrotransposase, a fragment thereof, or a reverse transcriptase domain, wherein the open reading frame is selected from the group consisting of a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a ... BAD In some embodiments, the open reading frame is operably linked to a promoter, a strong left-handed promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to a sequence encoding a retrotransposase, a fragment thereof, or a reverse transcriptase domain.
[0064] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.
[0065] In some aspects, the present disclosure provides methods for producing a retrotransposase, a fragment thereof, or a reverse transcriptase domain, comprising culturing any of the host cells described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the retrotransposase, a fragment thereof, or a reverse transcriptase domain by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to affinity chromatography or ion affinity chromatography specific for the affinity tag.
[0066] In some aspects, the present disclosure provides in vitro transcribed mRNA comprising an RNA homolog of any of the nucleic acids described herein.
[0067] In some aspects, the present disclosure provides an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase is derived from an uncultured microorganism. In some embodiments, the cargo nucleotide sequence is engineered. In some embodiments, the cargo nucleotide sequence is heterologous. In some embodiments, the cargo nucleotide sequence does not have the sequence of a wild-type genomic sequence present in an organism. In some embodiments, the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to retrotransposases described in the literature. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. In some embodiments, the retrotransposase comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the retrotransposase. In some embodiments, the NLS comprises a sequence at least 80% identical to a sequence selected from the group consisting of SEQ ID NOs: 1477-1492.In some embodiments, sequence identity is determined by CLUSTALW using the parameters of BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithms. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at presence of 11 and extension of 1, and using a conditional composition score matrix adjustment.
[0068] In some aspects, the present disclosure provides an engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus, and (ii) the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase is derived from an uncultured microorganism. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to a literature-described retrotransposase. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate. In some embodiments, sequence identity is determined by CLUSTALW using the parameters of BLASTP, CLUSTALW, MUSCLE, MAFFT, or Smith-Waterman homology search algorithms. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using the BLOSUM62 scoring matrix with parameters of word length (W) of 3, expectation (E) of 10, and gap costs of 11 presences and 1 extensions, with a conditional composition score matrix adjustment.
[0069] In some aspects, the present disclosure provides a deoxyribonucleic acid polynucleotide encoding an engineered retrotransposase system according to any one of the aspects or embodiments described herein.
[0070] In some aspects, the disclosure provides a nucleic acid comprising an engineered nucleic acid sequence optimized for expression in an organism, wherein the nucleic acid encodes a retrotransposase, wherein the retrotransposase is derived from an uncultivated microorganism, and wherein the organism is not an uncultivated microorganism. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, and 1546-1553. In some embodiments, the retrotransposase comprises a sequence encoding one or more nuclear localization sequences (NLSs) proximal to the N- or C-terminus of the retrotransposase. In some embodiments, the NLS comprises a sequence selected from SEQ ID NOs: 1477-1492. In some embodiments, the NLS comprises SEQ ID NO: 1478. In some embodiments, the NLS is proximal to the N-terminus of the retrotransposase. In some embodiments, the NLS comprises SEQ ID NO: 1477. In some embodiments, the NLS is proximal to the C-terminus of the retrotransposase. In some embodiments, the organism is a prokaryote, a bacterium, a eukaryote, a fungus, a plant, a mammal, a rodent, or a human.
[0071] In some aspects, the present disclosure provides a vector comprising the nucleic acid of any one of the aspects or embodiments described herein. In some embodiments, the vector further comprises a nucleic acid encoding a cargo nucleotide sequence configured to form a complex with a retrotransposase. In some embodiments, the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
[0072] In some aspects, the present disclosure provides a cell comprising the vector of any one of the aspects or embodiments described herein.
[0073] In some aspects, the present disclosure provides a method of producing a retrotransposase, comprising culturing a cell according to any of the aspects or embodiments described herein.
[0074] In some aspects, the disclosure provides methods for binding, nicking, cleaving, marking, modifying, or transposing a double-stranded deoxyribonucleic acid polynucleotide, the method comprising: (a) contacting the double-stranded deoxyribonucleic acid polynucleotide with a retrotransposase configured to transpose a cargo nucleotide sequence to a target nucleic acid locus, wherein the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase is derived from an uncultured microorganism. In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease domain. In some embodiments, the retrotransposase has less than 80% sequence identity to a retrotransposase described in the literature. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR). In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is transposed via a ribonucleic acid polynucleotide intermediate. In some embodiments, the double-stranded deoxyribonucleic acid polynucleotide is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide.
[0075] In some aspects, the present disclosure provides a method of modifying a target nucleic acid locus, the method comprising delivering to the target nucleic acid locus an engineered retrotransposase system of any one of the aspects or embodiments described herein, wherein the retrotransposase is configured to transpose a cargo nucleotide sequence to the target nucleic acid locus, and wherein the complex is configured to modify the target nucleic acid locus upon binding of the complex to the target nucleic acid locus. In some embodiments, modifying the target nucleic acid locus comprises binding, nicking, cleaving, marking, modifying, or transposing the target nucleic acid locus. In some embodiments, the target nucleic acid locus comprises deoxyribonucleic acid (DNA). In some embodiments, the target nucleic acid locus comprises genomic DNA, viral DNA, or bacterial DNA. In some embodiments, the target nucleic acid locus is in vitro. In some embodiments, the target nucleic acid locus is within a cell. In some embodiments, the cell is a prokaryotic cell, a bacterial cell, a eukaryotic cell, a fungal cell, a plant cell, an animal cell, a mammalian cell, a rodent cell, a primate cell, a human cell, or a primary cell. In some embodiments, the cell is a primary cell. In some embodiments, the primary cell is a T cell. In some embodiments, the primary cell is a hematopoietic stem cell (HSC).
[0076] In some aspects, the present disclosure provides a method of any one of the aspects or embodiments described herein, wherein delivering an engineered retrotransposase system to a target nucleic acid locus comprises delivering a nucleic acid described in any one of the aspects or embodiments described herein, or a vector described in any of the aspects or embodiments described herein. In some embodiments, delivering an engineered retrotransposase system to a target nucleic acid locus comprises delivering a nucleic acid comprising an open reading frame encoding a retrotransposase. In some embodiments, the nucleic acid comprises a promoter to which an open reading frame encoding the retrotransposase is operably linked. In some embodiments, delivering an engineered retrotransposase system to a target nucleic acid locus comprises delivering a capped mRNA containing an open reading frame encoding the retrotransposase. In some embodiments, delivering an engineered retrotransposase system to a target nucleic acid locus comprises delivering a translated polypeptide. In some embodiments, the retrotransposase does not induce cleavage at or proximal to the target nucleic acid locus.
[0077] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding a heterologous retrotransposase having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the host cell is an E. coli cell. In some embodiments, the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is encoded by a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a ... araP promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a rhaBAD promoter sequence, a r BADIn some embodiments, the open reading frame is operably linked to a promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to the sequence encoding the retrotransposase. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose-binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in a host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of a host cell.
[0078] In some aspects, the present disclosure provides a culture comprising a host cell according to any one of the aspects or embodiments described herein in a compatible liquid medium.
[0079] In some aspects, the present disclosure provides a method for producing a retrotransposase, comprising culturing a host cell described in any one of the aspects or embodiments described herein in a compatible growth medium. In some embodiments, the method further comprises inducing expression of the retrotransposase by adding an additional chemical agent or an increased amount of a nutrient. In some embodiments, the additional chemical agent or increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract. In some embodiments, the method further comprises subjecting the protein extract to IMAC or ion affinity chromatography. In some embodiments, the open reading frame comprises a sequence encoding an IMAC affinity tag linked in-frame to the sequence encoding the retrotransposase. In some embodiments, the IMAC affinity tag is linked in-frame to the sequence encoding the retrotransposase via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the IMAC affinity tag is removed by contacting the retrotransposase with a protease corresponding to the protease cleavage site. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the retrotransposase.
[0080] In some aspects, the disclosure provides a method of disrupting a genetic locus in a cell, the method comprising contacting the cell with a composition comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence, wherein the cargo nucleotide sequence is configured to interact with a retrotransposase; and (b) a retrotransposase, wherein: (i) the retrotransposase is configured to transpose the cargo nucleotide sequence to a target nucleic acid locus; (ii) the retrotransposase comprises a sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266; and (iii) the retrotransposase has transposition activity in the cell that is at least equivalent to a retrotransposase described in the literature. In some embodiments, the transposition activity is measured in vitro by introducing a retrotransposase into a cell containing a target nucleic acid locus and detecting transposition of the target nucleic acid locus in the cell. In some embodiments, the composition comprises 20 pmoles or less of the retrotransposase. In some embodiments, the composition comprises 1 pmol or less of the retrotransposase.
[0081] In some aspects, the disclosure provides a host cell comprising an open reading frame encoding any of the proteins or polypeptides described herein. In some embodiments, the host cell is an E. coli cell or a mammalian cell. In some embodiments, the host cell is an E. coli cell, wherein the E. coli cell is a λDE3 lysogen or the E. coli cell is a BL21(DE3) strain. In some embodiments, the E. coli cell has an ompT lon genotype. In some embodiments, the open reading frame is encoded by a T7 promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, an araP promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a araP promoter sequence, a T7-lac promoter sequence, a lac promoter sequence, a tac promoter sequence, a trc promoter sequence, a ParaBAD promoter sequence, a PrhaBAD promoter sequence, a T5 promoter sequence, a cspA promoter sequence, a araP promoter sequence, a rAp ... BADIn some embodiments, the open reading frame is operably linked to a promoter, a strong leftward promoter from phage lambda (pL promoter), or any combination thereof. In some embodiments, the open reading frame comprises a sequence encoding an affinity tag linked in-frame to the protein-encoding sequence. In some embodiments, the affinity tag is an immobilized metal affinity chromatography (IMAC) tag. In some embodiments, the IMAC tag is a polyhistidine tag. In some embodiments, the affinity tag is a myc tag, a human influenza hemagglutinin (HA) tag, a maltose-binding protein (MBP) tag, a glutathione S-transferase (GST) tag, a streptavidin tag, a strep tag, a FLAG tag, or any combination thereof. In some embodiments, the affinity tag is linked in-frame to the protein-encoding sequence via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site is a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the open reading frame is codon-optimized for expression in a host cell. In some embodiments, the open reading frame is provided on a vector. In some embodiments, the open reading frame is integrated into the genome of a host cell.
[0082] In some aspects, the disclosure provides a culture comprising any of the host cells described herein in a compatible liquid medium.
[0083] In some aspects, the present disclosure provides methods for producing any of the proteins described herein, comprising culturing any of the host cells described herein encoding any of the proteins described herein in a suitable growth medium. In some embodiments, the method further comprises inducing expression of the protein. In some embodiments, the induction of nuclease expression is by the addition of an additional chemical agent or an increased amount of a nutrient, or by an increased or decreased temperature. In some embodiments, the additional chemical agent or increased amount of a nutrient comprises isopropyl β-D-1-thiogalactopyranoside (IPTG) or an additional amount of lactose. In some embodiments, the method further comprises isolating the host cells after culturing and lysing the host cells to produce a protein extract comprising the protein. In some embodiments, the method further comprises isolating the protein. In some embodiments, the isolating comprises subjecting the protein extract to IMAC, ion exchange chromatography, anion exchange chromatography, or cation exchange chromatography. In some embodiments, the host cell comprises a nucleic acid comprising an open reading frame comprising a sequence encoding an affinity tag linked in-frame to a sequence encoding the protein. In some embodiments, the affinity tag is linked in-frame to the protein-encoding sequence via a linker sequence encoding a protease cleavage site. In some embodiments, the protease cleavage site comprises a tobacco etch virus (TEV) protease cleavage site, a PreScission® protease cleavage site, a thrombin cleavage site, a factor Xa cleavage site, an enterokinase cleavage site, or any combination thereof. In some embodiments, the method further comprises cleaving the affinity tag by contacting the protein with a protease corresponding to the protease cleavage site. In some embodiments, the affinity tag is an IMAC affinity tag. In some embodiments, the method further comprises performing subtractive IMAC affinity chromatography to remove the affinity tag from the composition comprising the protein.
[0084] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive. [Brief explanation of the drawings]
[0085] The novel features of the present disclosure are set forth with particularity in the appended claims. The features and advantages of the present disclosure will be better understood by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosure are utilized, and the accompanying drawings.
[0086] [Figure 1] Figure 1 shows the genomic context of bacterial retrotransposons. MG140-1 is a predicted retrotransposase (arrow) that encodes a zinc finger DNA-binding domain and a reverse transcriptase domain. Regions flanking the retrotransposase display secondary structures that may represent binding sites for the retrotransposase (secondary structure box and zoomed image). Regions of similarity with other homologs indicate putative target sites where the retrotransposon integrated. [Figure 2] Figure 2 shows that microbial MG retrotransposases (thick black branch on clade 4) are more closely related to eukaryotes than viral retrotransposases (thick black branch on clade 6). Clade 1: telomerase reverse transcriptases, class 2: group II intron reverse transcriptases, class 3: eukaryotic R1 retrotransposases, clade 4: microbial and eukaryotic R2 retrotransposases, clade 5: eukaryotic retrovirus-associated reverse transcriptases, and class 6: viral reverse transcriptases. [Figure 3]Figure 3 shows clades 3 and 4 from the phylogenetic gene tree from Figure 2. Some microbial MG retrotransposases contain multiple Zn finger motifs (vertical rectangles), a conserved RVT_1 reverse transcriptase domain, and an APE / RLE or other endonuclease domain (top and bottom panels). Some microbial MG retrotransposases lack the endonuclease domain (middle panel). [Figure 4] Figure 4 shows a phylogenetic tree inferred from a multiple sequence alignment of reverse transcriptase domains from diverse enzymes. RT sequences are derived from DNA and RNA ensembles. A reference RT was included in the tree for classification purposes. [Figure 5A] Figure 5A shows a phylogenetic tree inferred from a multiple sequence alignment of identified RT domains from a family of non-LTR retrotransposases (MG140, MG146, and MG147) and a related RT (MG148). [Figure 5B] Figure 5B shows data demonstrating that non-LTR retrotransposases (MG140, MG146, and MG147) contain an RT domain, an endonuclease domain (Endo), and multiple zinc-binding ribbon motifs, while the MG148 RT family lacks an endonuclease domain. [Figure 6A] Figure 6A shows data demonstrating that the MG140 R2 retrotransposase contains RT and endonuclease (EN) domains, as well as multiple zinc fingers, and shares 24%-26% average amino acid identity (AAI) with the reference Danio rerio R2 retrotransposase (R2Dr). [Figure 6B] Figure 6B shows data demonstrating that the MG140-47 R2 retrotransposon integrates into the 28S rRNA gene. Alignment of the MG140-47 contig to the reference (GQ398061) ribosomal RNA operon shows a large gap in the reference 28S rDNA gene due to the integration of the R2 element (dotted box) into the MG140-47 28S rDNA gene. [Figure 7] Figure 7A shows the genomic context of the MG145-45 retrotransposon. The enzyme contains an RT domain and a zinc finger domain. A partial 18S rDNA gene at the 5' end and a polyA tail at the 3' end likely delineate the boundaries of the transposon. [Figure 8A] FIG. 8A shows the contig encoding the MG146-1 retrotransposase with the RT and endonuclease domains. [Figure 8B] Figure 8B shows the MG140-17-R2 retrotransposon, which encodes three genes predicted to be involved in mobilization: an RNA recognition motif gene (RRM), an endonuclease enzyme, and a reverse transcriptase with RT and RNAse H domains. [Figure 9A] Figure 9A shows the genomic context of two members of the MG148 family of RTs. Predicted genes not associated with RT are shown as white arrows. [Figure 9B] Figure 9B shows a nucleotide sequence alignment of five members of the MG148 family (annotated arrow above the consensus sequence) showing the conserved region upstream of the RT (box below the sequence). [Figure 10] Figure 10 shows the in vitro activity screening of the RTns family of enzymes (MG140) by qPCR. Activity was detected by qPCR using primers that amplify full-length cDNA products derived from primer extension reactions containing each RT. Samples were derived from RT reactions containing 100 nM substrate. Negative control: water control without template in in vitro expression reactions; Positive control 1: R2Tg (Taeniopygia guttata); Positive control 2: R2Bm (Bombyx mori). Two positive controls are R2 retrotransposons, as described in the literature. Active candidates, defined as signals at least 10-fold higher than the negative controls, are marked with a diagonal line; inactive candidates under these conditions are marked with a white bar. [Figure 11]Figure 11 shows the in vitro activity screening of the RTns family of enzymes (MG146, MG147, and MG148) by qPCR. Activity was detected by qPCR using primers amplifying full-length cDNA products derived from primer extension reactions containing the respective RTs. Samples were derived from RT reactions containing 100 nM substrate. Negative control: water control without template in in vitro expression reactions. Positive control 1: R2Tg (Taeniopygia guttata), an R2 retrotransposon described in the literature. Active candidates, defined as signals at least 10-fold higher than the negative control, are marked with a diagonal line; inactive candidates under these conditions are marked with a white bar. [Figure 12] Figure 12 shows an assay for assessing the fidelity of R2 and R2-like candidates by next-generation sequencing. The cDNA products obtained from the primer extension reactions were PCR-amplified and libraries were prepared for NGS. Trimmed reads were aligned to the reference sequence, and the frequency of misincorporation was calculated. Background: Water control without template in in vitro expression reactions. Positive control 1: R2Tg (Taeniopygia guttata). [Figure 13A] Figure 13A shows a phylogenetic tree inferred from a multiple sequence alignment of full-length group II intron RTs identified from novel families from diverse classes. [Figure 13B] Figure 13B shows a summary table of the MG families of group II introns. AAI: average pairwise amino acid identity of the MG family to the reference group II intron sequence. [Figure 14A]Figures 14A-14D show the in vitro activity screening of GII intron class C candidates MG153-1 through MG153-21 and MG153-25 through MG153-27 by primer extension assay. For Figures 14A-14C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - Marathon RT control RT. Bold numbers correspond to gel lanes with active candidates. Results are representative of two independent experiments. Lanes 5-14 in Figure 14A correspond to candidates MG153-1 through MG153-10. Arrows in Figures 14A-14C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 14B] Figures 14A-14D show the in vitro activity screening of GII intron class C candidates MG153-1 through MG153-21 and MG153-25 through MG153-27 by primer extension assay. For Figures 14A-14C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - Marathon RT control RT. Bold numbers correspond to gel lanes with active candidates. Results are representative of two independent experiments. Lane numbers 5-14 in Figure 14B correspond to candidates MG153-11 through MG153-20. Arrows in Figures 14A-14C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 14C]Figures 14A-14D show the in vitro activity screening of GII intron class C candidates MG153-1 through MG153-21 and MG153-25 through MG153-27 by primer extension assay. For Figures 14A-14C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4 - Marathon RT control RT. Bold numbers correspond to gel lanes with active candidates. Results are representative of two independent experiments. Lane numbers 5-8 in Figure 14C correspond to candidates MG153-21, MG153-25, MG153-26, and MG153-27, respectively. Arrows in Figures 14A-14C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 14D] Figures 14A-14D show the in vitro activity screening of GII intron class C candidates MG153-1 to MG153-21 and MG153-25 to MG153-27 by primer extension assay. Figure 14D shows the detection of full-length cDNA production by qPCR. The shaded bars correspond to RTs that generate products at least 10-fold above background. Results were measured in two technical replicates. [Figure 15A] Figures 15A-15D show the in vitro activity screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 by primer extension assay. For Figures 15A-15C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4-13 in Figure 15A correspond to candidates MG153-28 to MG153-37. Arrows in Figures 15A-15C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 15B]Figures 15A-15D show the in vitro activity screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 by primer extension assay. For Figures 15A-15C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4-13 in Figure 15B correspond to candidates MG153-39 to MG153-48. Arrows in Figures 15A-15C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 15C] Figures 15A-15D show the in vitro activity screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 by primer extension assay. For Figures 15A-15C, lane numbers correspond to the following: 1 - PURExpress (in vitro expression) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT. Bold numbering corresponds to gel lanes. Lane numbers 4-13 in Figure 15C correspond to candidates MG153-49 to MG153-57. Arrows in Figures 15A-15C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 15D] Figures 15A-15D show the in vitro activity screening of GII intron class C candidates MG153-28 to MG153-37 and MG153-39 to MG153-57 by primer extension assay. Figure 15D shows the detection of full-length cDNA production by qPCR. The shaded bars correspond to RTs that generate products at least 10-fold above background. Results were measured in two technical replicates. [Figure 16A]Figures 16A-B show the in vitro activity screening of the GII intron class D MG165 family of reverse transcriptases by primer extension assay. For Figure 16A, lane numbers correspond to the following: 1 - PURExpress (in vitro expressed) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4-12 - candidate MG165-1-9. Bold numbering corresponds to the gel lanes with active candidates. Arrows in Figure 16A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 16B] Figures 16A-B show the in vitro activity screening of the GII intron class D MG165 family of reverse transcriptases by primer extension assay. Figure 16B shows the quantification of full-length cDNA production by qPCR. The shaded bars correspond to RTs that generate products at least 10-fold above background. Results were measured in duplicate. [Figure 17A] Figures 17A-B show the in vitro activity screening of the GII intron class F MG167 family of reverse transcriptases by primer extension assay. For Figure 17A, lane numbers correspond to the following: 1 - PURExpress (in vitro expressed) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4-8 - candidate MG167-1. Bold numbers correspond to gel lanes with active candidates. Arrows in Figure 17A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 17B] Figures 17A-B show the in vitro activity screening of the GII intron class F MG167 family of reverse transcriptases by primer extension assay. Figure 17B shows the quantification of full-length cDNA production by qPCR. The shaded bars correspond to RTs that generate products at least 10-fold above background. Results were measured in two technical replicates. [Figure 18]Figure 18 shows an assay for assessing the fidelity of GII intron class C RT candidates from the MG153 family by next-generation sequencing. The cDNA products obtained from the primer extension reactions were PCR-amplified and library-prepared for NGS. Trimmed reads were aligned to the reference sequence, and the frequency of misincorporation was calculated. Results were measured in two technical replicates. [Figure 19A] Figures 19A-19C show screens to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells. Figure 19A shows the detection of 542 bp (top) and 100 bp (bottom) PCR products by agarose gel analysis. Lanes not relevant to the experiments described in Figures 19A and 19B are highlighted in white boxes. [Figure 19B] Figures 19A-19C show screens to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells. Figure 19B shows detection of 542 bp (top) and 100 bp (bottom) PCR products using a D1000 TapeStation. Lanes not relevant to the experiments described in Figures 19A and 19B are highlighted in white boxes. [Figure 19C] Figures 19A-19C show a screen to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells. Figure 19C shows detection of a 542 bp PCR product by D1000 TapeStation for additional candidates. [Figure 20A] Figure 20A shows a phylogenetic tree of full-length G2L4-like RTs. The reference G2L4 sequence and the MG172 candidate (dots) are highlighted. [Figure 20B] FIG. 20B shows data demonstrating that columns 277-280 of the reference and MG172 RT represent catalytic residues involved in reverse transcriptase function. [Figure 21A] Figure 21A shows a phylogenetic tree of full-length LTR RTs. The reference LTR RT sequence and the MG151 candidate (dots) are highlighted. [Figure 21B] Figure 21B shows the genomic context of MG151-82 RT (labeled ORF 7). Predicted domains are shown as boxes, and long terminal repeats (LTRs) are shown as arrows flanking the LTR transposon. [Figure 21C] FIG. 21C shows the 3D structure prediction of MG151-82 showing the protease domain, RT domain, RNAse H domain, and integrase domain. [Figure 22] Figure 22 shows a multiple sequence alignment of the full-length pol protein sequence to highlight the protease domain, RT-RNAse H domain, and integrase domain. The catalytic residues of the RT, RNAse H, and integrase domains of MMLV RT are indicated by bars under each domain. The protease domain of the MMLV reference sequence is not shown in the alignment. [Figure 23A] Figures 23A-C show the in vitro activity screening of virus candidates MG151-80 through MG151-97 by primer extension assay. For Figure 23A, lane numbers correspond to the following: 1—RNA template annealed to primer, 2—MMLV control RT, 3—Ty3 control RT, 4-9—candidates MG151-80 through MG151-85, 10—RT control. Arrows in Figures 23A-C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 23B] Figures 23A-C show the in vitro activity screening of virus candidates MG151-80 to MG151-97 by primer extension assay. For Figure 23B, lane numbers correspond to the following: 1—RNA template annealed to primer, 2–12—candidates MG151-87–97, 13—MMLV control RT. Arrows in Figures 23A-C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 23C]Figures 23A-C show the screening of virus candidates MG151-80 to MG151-97 for in vitro activity by primer extension assay. Figure 23C shows testing of the in vitro activity of Ty3 control RT in different buffer conditions. Lane numbers correspond to the following: 1 - PURExpress (in vitro expressed) no template control, 2 - Buffer A (40 mM Tris-HCl pH 7.5, 0.2 M NaCl, 10 mM MgCl, 1 mM TCEP), 3 - Buffer B (20 mM Tris pH 7.5, 150 mM KCl, 5 mM MgCl, 1 mM TCEP, 2% PEG-8000), 4 - Buffer C (10 mM Tris-HCl pH 7.5, 80 mM NaCl, 9 mM MgCl, 1 mM TCEP, 0.01% (v / v) Triton X-100), 5 - Buffer D (10 mM Tris pH 7.5, 130 mM NaCl, 9 mM MgCl, 1 mM TCEP, 10% glycerol). Arrows in Figures 23A-C indicate full-length cDNA products (arrows near the top of the gel) and examples of cDNA drop-off (lower arrows). [Figure 24A]Figures 24A-B show testing of in vitro RT processivity and priming parameters for candidate MG151-89, MG151-92, and MG151-97 on structured RNA templates. For Figures 24A and 24B, lane 1: 6, 10, and 16 nucleotide oligomarkers (arrows); lane 2: 8, 13, and 20 nucleotide oligomarkers; lane 3: 43 and 55 nucleotide oligomarkers; lanes 4 and 10: 6 nucleotide primers; lanes 5 and 11: 8 nucleotide primers; lanes 6 and 12: 10 nucleotide primers; lanes 7 and 13: 13 nucleotide primers; lanes 8 and 14: 16 nucleotide primers; and lanes 9 and 15: 20 nucleotide primers. Lanes 4-9 in Figure 24A correspond to reverse transcription reactions containing MMLV with various primer lengths. MMLV reverse transcribes through the structured RNA hairpin. Lanes 10-15 correspond to reverse transcription reactions containing MG151-89 with various primer lengths. MG151-89 prefers primer lengths of 16 and 20 nucleotides and appears to terminate reverse transcription at structured RNA hairpins. [Figure 24B]Figures 24A-B show testing of in vitro RT processivity and priming parameters of candidate MG151-89, MG151-92, and MG151-97 on structured RNA templates. For Figures 24A and 24B, lane 1: 6, 10, and 16 nucleotide oligomarkers (arrows); lane 2: 8, 13, and 20 nucleotide oligomarkers; lane 3: 43 and 55 nucleotide oligomarkers; lanes 4 and 10: 6 nucleotide primers; lanes 5 and 11: 8 nucleotide primers; lanes 6 and 12: 10 nucleotide primers; lanes 7 and 13: 13 nucleotide primers; lanes 8 and 14: 16 nucleotide primers; and lanes 9 and 15: 20 nucleotide primers. Lanes 4-9 in Figure 24B correspond to reverse transcription reactions containing MG151-92 with various primer lengths. Lanes 10-15 correspond to reverse transcription reactions containing MG151-97 with various primer lengths. Neither MG151-92 nor MG151-97 appear to be active under these experimental conditions. [Figure 25] Figure 25 shows a phylogenetic analysis of 2407 retron RTs, highlighting the initial candidates selected for downstream characterization in vitro. Nine of the 16 experimentally validated retrons from the literature were added and highlighted in the tree. Stars represent candidate MG154-MG159 and MG173 family members. [Figure 26] Figure 26 shows the genomic context of the MG157-1 retron (RT labeled with an arrow on a line). Retron non-coding RNAs (ncRNAs) are highlighted with dotted boxes. [Figure 27A] Figure 27A shows an inset depicting the MG157-1 retron ncRNA with its flanking inverted repeats. [Figure 27B] Figure 27B shows the predicted structure of the MG157-1 retron ncRNA. [Figure 28A] Figure 28A shows the genomic context of the MG160-3 retron-like single-domain RT. The region upstream from the RT (dotted box) is conserved across MG160 members. [Figure 28B] FIG. 28B shows the 3D structure prediction of MG160-3, showing the RT domain aligned to the group II intron cryo-EM structure. [Figure 28C] Figure 28C shows the predicted structures of the 5'UTRs of five MG160 members. [Figure 29A] Figures 29A-B show the screening of retron-like candidates MG160-1 through MG160-6 and MG160-8 for in vitro activity by primer extension assay. Lane numbers in Figure 29A correspond to the following samples: 1 - PURExpress (in vitro expressed) no-template control, 2 - MMLV control RT, 3 - TGIRT-III control RT, 4-10 - candidates MG160-1 through MG160-6 and MG160-8. Bold numbers correspond to gel lanes with active candidates. Arrows in Figure 29A indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (lower arrow). [Figure 29B] Figures 29A-B show the screening of retron-like candidates MG160-1 to MG160-6 and MG160-8 for in vitro activity by primer extension assay. Figure 29B shows quantification of full-length cDNA production by qPCR. Shaded bars correspond to RTs that generate products at least 10-fold above background. Results were measured in two technical replicates. [Figure 30A] Figures 30A-30C show cell-free expression of retron RT candidates and the generation of retron ncRNA by in vitro transcription. Figure 30A shows confirmation of retron RT protein production in a cell-free expression system. Lanes correspond to the following: 1: ladder, 2: no-template control, 3: MG156-1 (39 kDa), 4: MG156-2 (40 kDa), 5: MG157-1 (38 kDa). [Figure 30B]Figures 30A-30C show cell-free expression of retron RT candidates and the generation of retron ncRNA by in vitro transcription. Figure 30B shows confirmation of retron RT protein production in a cell-free expression system. Lanes correspond to the following: 1: ladder, 2: no-template control, 3: MG157-2 (37 kDa), 4: MG157-5 (43 kDa), 5: MG159-1 (53 kDa), and 6: Ec86 (38 kDa, positive control retron RT). [Figure 30C] Figures 30A-30C show cell-free expression of retron RT candidates and the generation of retron ncRNAs by in vitro transcription. Figure 30C shows the generation of retron ncRNA templates by in vitro transcription. Lanes correspond to the following ncRNAs: 1: MG154-1, 2: MG154-2, 3: MG155-1, 4: MG155-2, 5: MG155-3, 6: MG156-1, 7: MG156-2, 8: MG157-1, 9: MG157-2, 10: MG157-5, 11: MG158-1, 12: MG159-1, 13: Ec86, 14: MG155-4, 15: MG173-1, and 16: MG155-5. [Figure 31] Figure 31 shows the domain architecture demonstrating that the MG140-1 R2 retrotransposon integrates into the 28S rRNA gene. The R2 retrotransposase (lightly shaded bar) contains multiple zinc fingers, as well as an RT domain and an endonuclease domain. MG140-1 is flanked by the 5'UTR and 3'UTR, which define the transposon boundaries. MG140-1 integrates precisely between the G and T nucleotides in the target site motif GGTAGC. [Figure 32]FIG. 32 shows testing of RT activity by primer extension using DNA oligos containing phosphorothioate linkage modifications. Lane numbers correspond to the following: 1: PURExpress (in vitro expressed) no-template control with PS-modified primer 1; 2: PURExpress (in vitro expressed) no-template control with PS-modified primer 2; 3: PURExpress (in vitro expressed) no-template control with PS-modified primer 3; 4: MMLV RT with unmodified primer; 5: MMLV RT with PS-modified primer 1; 6: MMLV RT with PS-modified primer 2; 7: MMLV RT with PS-modified primer 3; 8: TGIRT-III with unmodified primer; 9: TGIRT-III with PS-modified primer 1; 10: TGIRT-III with PS-modified primer 2; 11: TGIRT-III with PS-modified primer 3; 12: MG153-9 with unmodified primer; 13: MG153-9 with PS-modified primer 1; 14: MG153-9 with PS-modified primer 2; 15: MG153-9 with PS-modified primer 3. MMLV RT and TGIRT-III are control RTs. [Figure 33] Figure 33 shows the screening of retron RT activity on RNA templates by primer extension assay. Lane numbers correspond to the following: 1: PURExpress (in vitro expressed) no template control, 2: MMLV control RT, 3: MG154-1, 4: MG155-1, 5: MG155-2, 6: MG155-3, 7: MG156-2, 8: MG157-1, 9: MG157-2, 10: MG157-5, 11: MG158-1, 12: MG159-1, 13: Ec86 control retron RT, 14: Sa163 control retron RT, 15: St85 control retron RT. Bold lanes correspond to retron RTs that exhibit primer extension activity on the tested substrates. [Figure 34]Figure 34 shows the screening of the ability of MG153 GII-derived RT to synthesize cDNA in mammalian cells. Detection of a 542-bp cDNA synthesis PCR product was assayed by Taqman qPCR. cDNA activity was normalized to an active TGIRT control, with TGIRT representing a value of 1. The Y-axis is shown on a log 10 scale. [Figure 35A] Figures 35A-C show protein expression of MG153 GII-derived RTs by immunoblot. Cells were transfected with plasmids containing candidate RTs, and protein expression was assessed by immunoblot to detect the HA peptide fused to the N-terminus of the RT. All lanes were normalized to total protein concentration. The white arrow indicates a band 2X the expected molecular size of the protein, indicating a protein dimer. Lanes not relevant to the experiments described in Figures 35A and 35B are covered with white boxes. [Figure 35B] Figures 35A-C show protein expression of MG153 GII-derived RTs by immunoblot. Cells were transfected with plasmids containing candidate RTs, and protein expression was assessed by immunoblot to detect the HA peptide fused to the N-terminus of the RT. All lanes were normalized to total protein concentration. The white arrow indicates a band 2X the expected molecular size of the protein, indicating a protein dimer. Lanes not relevant to the experiments described in Figures 35A and 35B are covered with white boxes. [Figure 35C] Figures 35A-C show protein expression of MG153 GII-derived RTs by immunoblot. Cells were transfected with plasmids containing candidate RTs, and protein expression was assessed by immunoblot to detect the HA peptide fused to the N-terminus of the RTs. All lanes were normalized to total protein concentration. The white arrow indicates a band 2X the expected molecular size of the protein, indicating a protein dimer. Figure 35C: Multiple sequence alignment of GII-derived RTs. The region shown corresponds to positions 196-201 of the alignment. The dimerization motif CACQQ (SEQ ID NO: 2267) is highlighted. [Figure 36] Figure 36 shows the relative activity of GII-derived RT normalized to protein expression. cDNA synthesis was detected by Taqman qPCR, and protein expression was detected by immunoblot. Activity against TGIRT was normalized to total protein concentration. The Y-axis is shown on a linear scale. [Figure 37A] Figures 37A-E show retroviral RTs for cDNA synthesis. Figure 37A shows a phylogenetic tree of full-length LTR RTs, highlighting the MG151 candidate (gray branch) and a new group of RTs belonging to betaretroviruses (stars). [Figure 37B] Figures 37A-E show retroviral RT for cDNA synthesis. Figure 37B shows the structural alignment of the MG RT domain (dark gray) to reference RT domains from simian retrovirus and mouse mammary tumor virus (light gray). [Figure 37C] Figures 37A-37E show retroviral RTs for cDNA synthesis. Figure 37C shows the screening of the MG151 family of retroviral RTs for in vitro cDNA synthesis activity. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: MG151-98; Lane 4: MG151-99; Lane 5: MG151-100; Lane 6: MG151-101; Lane 7: MG151-102; Lane 8: MG151-103; Lane 9: MG151-104; Lane 10: MG151-105. Bold numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrows near the top of the gel), and lines indicate examples of cDNA drop-off. [Figure 37D]Figures 37A-37E show retroviral RTs for cDNA synthesis. Figure 37D shows the screening of the MG151 family of retroviral RTs for in vitro cDNA synthesis activity. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: MG151-106; Lane 4: MG151-107; Lane 5: MG151-108; Lane 6: MG151-109; Lane 7: MG151-110; Lane 8: MG151-111; Lane 9: MG151-112; Lane 10: MG151-113; Lane 11: MG151-114; Lane 12: MG151-115; Lane 13: MG151-116; Lane 14: MG151-117. Bold numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrows near the top of the gel) and lines indicate examples of cDNA drop-off. [Figure 37E]Figures 37A-E show retroviral RTs for cDNA synthesis. Figure 37E shows screening of the in vitro cDNA synthesis activity of the MG151 family of retroviral RTs using unmodified and modified RNA substrates. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) non-template control using a uridine-containing RNA (U-RNA) substrate; Lane 2: PURExpress (in vitro expressed) non-template control using an N1-methylpsuedouridine-containing RNA (m1Ψ-RNA) substrate; Lane 3: MMLV control RT using a U-RNA substrate; Lane 4: MMLV control RT using a m1Ψ-RNA substrate; Lane Lane 5: MG151-118 using U-RNA substrate; Lane 6: MG151-118 using m1Ψ-RNA substrate; Lane 7: MG151-119 using U-RNA substrate; Lane 8: MG151-119 using m1Ψ-RNA substrate; Lane 9: MG151-120 using U-RNA substrate; Lane 10: MG151-120 using m1Ψ-RNA substrate; Lane 11: MG151-121 using U-RNA substrate; Lane 12: m1Ψ-RNA substrate Lane 13: MG151-122 with U-RNA substrate; Lane 14: MG151-122 with m1Ψ-RNA substrate; Lane 15: MG151-123 with U-RNA substrate; Lane 16: MG151-123 with m1Ψ-RNA substrate; Lane 17: MG151-124 with U-RNA substrate; Lane 18: MG151-124 with m1Ψ-RNA substrate; Lane 19: MG151-124 with U-RNA substrate Lane 1: MG151-125 using mΨ-RNA substrate; Lane 20: MG151-125 using mΨ-RNA substrate; Lane 21: MG151-126 using U-RNA substrate; Lane 22: MG151-126 using mΨ-RNA substrate; Lane 23: MG151-127 using U-RNA substrate; Lane 24: MG151-127 using mΨ-RNA substrate; Lane 25: MG151-128 using U-RNA substrate; Lane 26: MG151-128 using mΨ-RNA substrate. Bold numbers correspond to the gel lanes with active candidates.Arrows indicate full-length cDNA products (arrows near the top of the gel) and lines indicate examples of cDNA drop-off. [Figure 38A] Figure 38A shows a phylogenetic tree of full-length retrons and the MG160 RT. The MG160 candidate (gray dot) is highlighted within a long branch within the retron clade. [Figure 38B] Figure 38B shows the structural alignment of MG160 RT (dark gray) to a reference retron RT from E. coli (Ec86, light gray). The additional N-terminus in Ec86 is boxed. [Figure 38C] Figure 38C shows a multiple sequence alignment of the full-length MG160 RT to the reference Ec86 lectin RT. The N-terminal region, RT domain, and C-terminal region are shown as bars below the reference sequence, and catalytic residues are highlighted in boxes. [Figure 38D-1] Figure 38D shows a multiple sequence alignment region of the active MG160 RT versus the group II intron and retron reference sequences. Enzyme-specific motifs are highlighted in boxes below the following sequences: MG160-specific motifs AXXXH and GX(3)Y[V / L]XXVN (SEQ ID NO: 2268), retron-specific motifs NAXXH and VTG, and group II intron-specific motifs GXXXY (partially shared with the MG160 enzyme) and FLG. The conserved histidine residue and motif [N / S]XXK found in most RTs are also highlighted. [Figure 38D-2] Continued from Figure 38D-1. [Figure 38D-3] Continued from Figure 38D-1 2. [Figure 38E]Figure 38E shows the screening of retron-like RTs and retron RTs of the MG154, MG155, MG156, MG157, MG158, MG159, and MG160 families for in vitro cDNA synthesis activity. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: MG160-28; Lane 4: MG160-31; Lane 5: MG160-37; Lane 6: MG160-40; Lane 7: MG160-51; Lane 8: MG160-52; Lane 9: MG160-53; Lane 10: MG160-54; Lane 11: MG160-55; Lane 12: MG160-56; Lane 13: MG160-57; Lane 14: MG160-58; Lane 15: MG160-59; Lane 20: MG160-60; Lane 21: MG160-61; Lane 22: MG160-62; Lane 23: MG160-63; Lane 24: MG160-64; Lane 25: MG160-65; Lane 26: MG160-66; Lane 27: MG160-67; Lane 28: MG160-68; Lane 29: MG160-70; Lane 30: MG160-71; Lane 31: MG160-72; Lane 32: MG160-73; Lane 33: MG160-74 Lane 5: MG160-59; Lane 16: MG160-60; Lane 17: MG160-61; Lane 18: unrelated lane; Lane 19: MG160-63; Lane 20: MG160-64; Lane 21: MG160-65; Lane 22: MG160-66; Lane 23: MG160-67; Lane 24: MG155-4; Lane 25: MG155-5; Lane 26: MG173-1. Lanes 3-23 correspond to the retron-like MG160 family of RTs. Lanes 24-26 correspond to retron RTs. Bold numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrow below). [Figure 38F]Figure 38F shows the in vitro cDNA synthesis activity screening of retron RTs from the MG154, MG155, MG156, MG157, MG158, and MG159 families. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expression) no template control; Lane 2: MMLV control RT; Lane 3: MG154-1; Lane 4: MG155-1; Lane 5: MG155-2; Lane 6: MG155-3; Lane 7: MG156-2; Lane 8: MG157-1; Lane 9: MG157-2; Lane 10: MG157-5; Lane 11: MG158-1; Lane 12: MG159-1; Lane 13: Ec86 control retron RT; Lane 14: Sa163 control retron RT; Lane 15: St85 control retron RT. Bold numbers correspond to the gel lanes containing active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrow below). [Figure 38G] Figure 38G shows the in vitro cDNA synthesis activity screening of retron RTs and retron-like RTs from the MG154, MG155, MG156, MG157, MG158, MG159, MG160, and MG173 families. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: TGIRT-III control RT; Lanes 4-7: unrelated MG RTs; Lane 8: MG160-17; Lane 9: MG154-2; Lane 10: MG156-1; Lane 11: MG157-3; Lane 12: MG157-4; Lane 13: MG159-2; Lane 14: MG159-3; Lane 15: MG173-2. Lane 8 corresponds to the retron-like MG160 family of RTs. Lanes 9-15 correspond to MG retron RTs. Bold numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrowhead or vertical line below). [Figure 39A]Figures 39A-39D show screening of the in vitro cDNA synthesis activity of GII intron RTs. Lane numbers in Figure 39A correspond to the following samples: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: TGIRT-III control RT; Lane 4: MG153-38; Lanes 5-9: MG163-1 to MG163-5; Lanes 10-13: MG166-2 to MG166-5. [Figure 39B] Figures 39A-39D show screening of the in vitro cDNA synthesis activity of GII intron RTs. For Figure 39B, lane numbers correspond to the following: Lane 1: PURExpress (in vitro expressed) no template control; Lane 2: MMLV control RT; Lane 3: TGIRT-III control RT; Lanes 4-14: MG169-1 to MG169-11. For both panels, bold lane numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrow below). [Figure 39C] Figures 39A-39D show screening of the in vitro cDNA synthesis activity of GII intron RTs. Figure 39C shows screening of the in vitro activity of GII intron classes C, A, B, E, G, ML, and CL (MG153, MG163, MG164, MG166, MG168, MG169, and MG170). Quantification of full-length cDNA production by qPCR. Lightly hatched bars correspond to RTs that generate enough cDNA for gel detection. Darkly shaded bars correspond to RTs that have detectable activity only by qPCR and generate product at least 10-fold above background. [Figure 39D] Figures 39A-39D show screening of in vitro cDNA synthesis activity of GII intron RTs. Figure 39D shows a summary of in vitro GII intron class A-G, class ML, and class CL cDNA synthesis activity. RT activity normalized to TGIRT was determined from quantification of full-length cDNA products after primer extension using a 202-nt RNA template. [Figure 40]Figure 40 shows the screening of R2 MG140 and MG146 families for in vitro activity by primer extension assay with quantification of full-length cDNA production by qPCR. Active RTs are those that generated product at least 10-fold above background (Purex) (dotted line). Results were measured in two technical replicates. Purex is PURExpress (in vitro expression), a no-template control, and MMLV and Tg R2 are control RTs. [Figure 41A] Figures 41A-B show the primer extension activity of the GII intron RT in vitro on a 4.1 kb RNA template. Figure 41A shows a schematic of the primer extension assay and detection of the cDNA product by Taqman qPCR. The RNA template contains an MS2 loop located 3' of the DNA priming oligo. The full-length cDNA product obtained from the RNA template is 4.1 kb. Taqman probes and primers are designed to quantify amplification of the first (FAM) and last (HEX) 100 bp amplicons of the cDNA. [Figure 41B] Figures 41A-B show the primer extension activity of GII intron RT in vitro on a 4.1 kb RNA template. Figure 41B shows the ratio of product to start (FAM) corresponding to the end of the cDNA (HEX) quantified for MG RT. TGIRT is the GII class C control RT, and MMLV is the retroviral control RT. [Figure 42] Figure 42 shows a schematic diagram illustrating the methodology used to detect cDNA synthesis in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using Taqman-based qPCR. [Figure 43A]Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class A MG163 candidate (Figure 43A). [Figure 43B] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class B MG164 candidate (Figure 43B). [Figure 43C] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class C MG153 candidate (Figure 43C). [Figure 43D] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class D MG165 candidate (Figure 43D). [Figure 43E]Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class E MG166 candidate (Figure 43E). [Figure 43F] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class F MG167 candidate (Figure 43F). [Figure 43G] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class G MG168 candidate (Figure 43G). [Figure 43H] Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class ML MG169 candidate (Figure 43H). [Figure 43I]Figures 43A-I show screens to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the following GII intron RTs: class CL MG170 candidate (Figure 43I). [Figure 44] Figure 44 shows a screen to assess the ability of the indicated control RT and R2 RT candidates to synthesize cDNA in mammalian cells. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the indicated R2 RT candidates. [Figure 45A] Figures 45A-45B show screening of the ability of the indicated Family II intron and R2 RT candidates to synthesize cDNA in mammalian cells with and without the MCP tag. Taqman qPCR was used to detect the first (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from RNA templates by the indicated Family II intron and R2 RT candidates, as well as the control TGIRT Family II intron and R2Tg R2 RT. R2 [Figure 45B] Figures 45A-45B show screening of the ability of the indicated Family II intron and R2 RT candidates to synthesize cDNA in mammalian cells with and without the MCP tag. Taqman qPCR was used to detect the first (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from RNA templates by the indicated Family II intron and R2 RT candidates, as well as the control TGIRT Family II intron and R2Tg R2 RT. R2 [Figure 46A]Figure 46A shows primer conversion activity of the MG151 family of RTs on standard (U) versus modified (m1Ψ) RNA templates. RT primer extension activity is normalized to the control retroviral RT, MMLV. [Figure 46B] Figure 46B shows the primer extension activity of various RTs against standard and m1Ψ-modified RNA templates. Lane numbers correspond to the following samples: Lane 1: PURExpress (in vitro expressed) NTC with standard RNA template; Lane 2: PURExpress (in vitro expressed) NTC with m1Ψ-modified RNA template; Lane 3: MMLV control RT with standard RNA template; Lane 4: MMLV control RT with m1Ψ-modified RNA template; Lane 5: TGIRT control RT with standard RNA template; Lane 6: TGIRT control RT with m1Ψ-modified RNA template; Lane 7: MG153-18 with standard RNA template; Lane 8: MG153-18 with m1Ψ-modified RNA template; Lane 9: MG153-20 with standard RNA template; Lane 10: MG153-20 with m1Ψ-modified RNA template; Lane 11: MG153-20 with standard RNA template. Lane 11: MG153-51 with mΨ-modified RNA template; Lane 12: MG153-51 with mΨ-modified RNA template; Lane 13: MG153-56 with standard RNA template; Lane 14: MG153-56 with mΨ-modified RNA template; Lane 15: MG170-1 with standard RNA template; Lane 16: MG170-1 with mΨ-modified RNA template; Lane 17: MG140-3 with standard RNA template; Lane 18: MG140-3 with mΨ-modified RNA template; Lane 19: MG140-8 with standard RNA template; Lane 20: MG140-8 with mΨ-modified RNA template; Lane 21: MG140-46 with standard RNA template; Lane 22: MG140-46 with mΨ-modified RNA template; Lane 23: Tg with standard RNA template Lane 24: Tg R2 control RT with m1Ψ-modified RNA template; Lane 25: MG160-4 with standard RNA template; Lane 26: MG160-4 with m1Ψ-modified RNA template. Arrows indicate full-length cDNA product (arrow near the top of the gel) and an example of cDNA drop-off (arrow below). [Figure 46C]Figures 46C-46D show quantification of RT activity on standard versus modified templates for various RTs. Figure 46C shows quantification of primer conversion by gel analysis. Results were measured in two technical replicates. [Figure 46D] Figures 46C-46D show quantification of RT activity on standard versus modified templates for various RTs. Figure 46D shows quantification of full-length cDNA production by qPCR performed on candidates with little or no detectable primer conversion on a denaturing gel. Technically, qPCR was performed twice to measure results. [Figure 47A] Figures 47A-C show a screen to assess the ability of the indicated control RT and GII intron class C candidates to synthesize cDNA in mammalian cells. Figure 47A shows a schematic diagram illustrating the methodology used to detect cDNA synthesis in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using TaqMan-based qPCR. [Figure 47B] Figures 47A-C show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using Taqman-based qPCR. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG148 family of non-LTR retrotransposon-derived RTs (Figure 47B) and the MG160 family of tretron-like RTs (Figure 47C). [Figure 47C]Figures 47A-C show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using Taqman-based qPCR. Taqman qPCR was used to detect the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG148 family of non-LTR retrotransposon-derived RTs (Figure 47B) and the MG160 family of tretron-like RTs (Figure 47C). [Figure 48] Figure 48 shows screening of rationally engineered variants of the optimal RT candidates MG153-18 and MG153-20 for their ability to synthesize cDNA in mammalian cells. Taqman qPCR detection of the first (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the indicated control and selected RT candidates. The MG153-18 variant showed a 5-fold increased activity compared to its WT counterpart, while the MG153-20 variant did not have improved activity. [Figure 49] Figure 49 shows screening of the indicated control RTs and putative inactivating mutants of the optimal group II intron-derived and R2 RT candidates for their ability to synthesize cDNA in mammalian cells. Taqman qPCR detection of the first (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from RNA templates by the indicated control and selected RT candidates. [Figure 50] Figure 50 shows a schematic diagram of the mechanism of retrons producing multiple copy single-stranded DNA (msDNA). [Figure 51]Figure 51 shows SDS-PAGE analysis of the expression of the MG173 and MG192 families from PURExpress. Protein expression is marked with arrows. Lane numbers correspond to the following: Lane 1: protein ladder; Lane 2: no template control (NTC); Lane 3: MG173-3; Lane 4: MG173-4; Lane 5: MG173-5; Lane 6: skip; Lane 7: MG173-6; Lane 8: MG173-7; Lane 9: protein ladder; Lane 10: no template control (NTC); Lane 11: MG173-8; Lane 12: MG173-9; Lane 13: MG173-10; Lane 14: MG192-1. [Figure 52] Figure 52 shows the general in vitro cDNA synthesis activity screening of retron RTs from the MG173 and MG192 families. Lane numbers correspond to the following samples: Lane 1: PURExpress no-template control, RT reaction containing no reverse transcriptase; Lane 2: positive control retroviral RT MMLV; Lane 3: positive control retron RT Ec86; Lanes 4-11: MG173-3 to MG173-10; Lane 12: MG192-1. Bold numbers correspond to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrowhead or vertical line below). [Figure 53A] Figures 53A and 53B show the in vitro primer extension activity of retron RT on a 4.1 kb RNA template. Figure 53A shows a schematic of the primer extension assay and detection of cDNA product by Taqman qPCR. The RNA template is annealed to a priming oligo before the initiation of the cDNA synthesis reaction. The full-length cDNA product obtained from the RNA template is 4.1 kb. Taqman probes and primers are designed to quantify amplification of the first (FAM) and final (HEX) 100 bp amplicons of cDNA. [Figure 53B]Figures 53A and 53B show the in vitro primer extension activity of retron RTs on a 4.1 kb RNA template. Figure 53B shows the ratio of product to start (FAM) corresponding to the end of the cDNA (HEX) quantified for MG RTs. TGIRT is the GII class C control RT, MMLV is the retrovirus control RT, and Ec86 is the retron control RT. [Figure 54] FIG. 54 shows the RT error substitution rates for the GII intron positive control RT TGIRT and the MG GII intron RTs MG153-5, MG153-18, MG153-20, MG153-51, and MG153-53 on standard and modified (N1-methylpseudouridine, m1Ψ) RNA templates. [Figure 55A] Figures 55A-55D show screens to assess the ability of the indicated control RTs and engineered candidate RTs to synthesize cDNA in mammalian cells. Figure 55A shows a schematic diagram illustrating the methodology used to detect cDNA synthesis in mammalian cells. The first (FAM) and last (HEX) 100 bp of a 4.1 kb RNA template were detected using TaqMan-based qPCR. [Figure 55B] Figures 55A-55D show screens to assess the ability of the indicated control RTs and genetically engineered candidate RTs to synthesize cDNA in mammalian cells. Figures 55B-55D show Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100-bp PCR products amplified from cDNA synthesized from RNA templates by the MG140-3 and MG140-8 non-LTR retrotransposon-derived RT variants (Figure 55B), the MG153-5, MG153-51, and MG169-1 variants of the GII intron RT (Figure 55C), and the MG153-18 and MG153-20 variants of the GII intron RT (Figure 55D). [Figure 55C]Figures 55A-55D show screens to assess the ability of the indicated control RTs and genetically engineered candidate RTs to synthesize cDNA in mammalian cells. Figures 55B-55D show Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100-bp PCR products amplified from cDNA synthesized from RNA templates by the MG140-3 and MG140-8 non-LTR retrotransposon-derived RT variants (Figure 55B), the MG153-5, MG153-51, and MG169-1 variants of the GII intron RT (Figure 55C), and the MG153-18 and MG153-20 variants of the GII intron RT (Figure 55D). [Figure 55D] Figures 55A-55D show screens to assess the ability of the indicated control RTs and genetically engineered candidate RTs to synthesize cDNA in mammalian cells. Figures 55B-55D show Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100-bp PCR products amplified from cDNA synthesized from RNA templates by the MG140-3 and MG140-8 non-LTR retrotransposon-derived RT variants (Figure 55B), the MG153-5, MG153-51, and MG169-1 variants of the GII intron RT (Figure 55C), and the MG153-18 and MG153-20 variants of the GII intron RT (Figure 55D). [Figure 56A] Figures 56A-56D show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Figures 56A-56D show Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG140 family of non-LTR retrotransposon-derived RTs (Figure 56A), the MG169 family of GII intron-derived RTs (Figure 56B), the MG153 family of GII intron-derived RTs (Figure 56C), and the retron RT (Figure 56D). [Figure 56B]Figures 56A-56D show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Figures 56A-56D show Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG140 family of non-LTR retrotransposon-derived RTs (Figure 56A), the MG169 family of GII intron-derived RTs (Figure 56B), the MG153 family of GII intron-derived RTs (Figure 56C), and the retron RT (Figure 56D). [Figure 56C] Figures 56A-56D show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Figures 56A-56D show Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG140 family of non-LTR retrotransposon-derived RTs (Figure 56A), the MG169 family of GII intron-derived RTs (Figure 56B), the MG153 family of GII intron-derived RTs (Figure 56C), and the retron RT (Figure 56D). [Figure 56D] Figures 56A-56D show a screen to assess the ability of the indicated control RTs and GII intron class C candidates to synthesize cDNA in mammalian cells. Figures 56A-56D show Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by the MG140 family of non-LTR retrotransposon-derived RTs (Figure 56A), the MG169 family of GII intron-derived RTs (Figure 56B), the MG153 family of GII intron-derived RTs (Figure 56C), and the retron RT (Figure 56D). [Figure 57A]Figures 57A-B show analysis of protein expression of selected RT candidates by Western blot. Western blot analysis of the MG153-18 and MG153-20 variants of the GII intron RT (Figure 57A). Blots showing anti-HA (top) and anti-cyclophilin (bottom). * indicates a nonspecific band detected in the anti-HA blot. [Figure 57B] Figures 57A-B show analysis of protein expression of selected RT candidates by Western blot. Selected candidates for GII intron Rts and R2 RTs have high cDNA synthesis activity and processivity (Figure 57B). Blots showing anti-HA (top) and anti-cyclophilin (bottom). * indicates a nonspecific band detected in the anti-HA blot. [Figure 58A] Figures 58A-B show a screen to assess the ability of the indicated control and trimmed candidate RTs to synthesize cDNA in mammalian cells. Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from RNA templates by trimmed variants of the MG140-3 and MG140-8 families of non-LTR retrotransposon-derived RTs (Figure 58A), and the MG140-74 and MG140-88 families of non-LTR retrotransposon-derived RTs (Figure 58B), compared with the activity of endonuclease domain (ED)-inactivated RTs and / or reverse transcriptase (RT) domain-inactivated RTs. [Figure 58B]Figures 58A-B show a screen to assess the ability of the indicated control and trimmed candidate RTs to synthesize cDNA in mammalian cells. Taqman qPCR detection of the first (FAM probe) and last (HEX probe) 100 bp PCR products amplified from cDNA synthesized from RNA templates by trimmed variants of the MG140-3 and MG140-8 families of non-LTR retrotransposon-derived RTs (Figure 58A), and the MG140-74 and MG140-88 families of non-LTR retrotransposon-derived RTs (Figure 58B), compared with the activity of endonuclease domain (ED)-inactivated RTs and / or reverse transcriptase (RT) domain-inactivated RTs. [Figure 59A] Figures 59A-59B show RT substitution error rates. Figure 59A shows the RT substitution error rate calculated from the consensus UMI sequences as mismatch / (mapping + mismatch) for standard (U) and modified (m1Ψ) RNA templates. Bar graphs show the mean, and the upper and lower bars show the 95% CI determined by Bayesian analysis. Data were derived from two independent experiments, each technically repeated three times. [Figure 59B] Figures 59A-59B show the RT substitution error rate. Figure 59B shows the theoretical length of a substitution-free cDNA molecule for each RT, calculated as 1 / (substitution error rate), for standard and modified RNA templates. MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 60-1] Figure 60 shows RT error types (mismatch, insertion, or deletion) by position along a canonical RNA template. The arrow indicates a substitution (or mismatch) hotspot shared between RTs at position 78. The inset image shows a portion of the predicted RNA template fold, with position 78 located within a putative hairpin. MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 60-2] Figure 60-1 continued 1. [Figure 60-3]Continued from Figure 60-1 2. [Figure 61-1] Figure 61 shows the RT error type (mismatch, insertion, or deletion) by position along the modified RNA template. MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control, respectively. [Figure 61-2] Figure 61-1 continued 1. [Figure 61-3] Continued from Figure 61-1 2. [Figure 62-1] Figure 62 shows RT substitution selectivity in standard or modified (RNA templates shown as a confusion matrix comparing reference nucleotides with observed nucleotide identities). MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 62-2] Figure 62-1 continued 1. [Figure 62-3] Continued from Figure 62-12. [Figure 63-1] Figure 63 shows RT indel analysis of standard RNA templates, showing the frequency and size of each observed insertion (number of positives) or deletion (number of negatives). MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 63-2] Figure 63-1 continued 1. [Figure 63-3] Continued from Figure 63-1 2. [Figure 63-4] Figure 63-1 continued 3. [Figure 64-1] Figure 64 shows the RT indel analysis on the modified RNA templates, showing the frequency and size of each observed insertion (number of positives) or deletion (number of negatives). MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 64-2] Figure 64-1 continued 1. [Figure 64-3] Continued from Figure 64-1 2. [Figure 64-4] Figure 64-1 continued 3. [Figure 65-1]Figure 65 shows the RT distribution of cDNA lengths on standard templates, showing cDNA drop-off products, full length, and non-template addition (NTA). MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 65-2] Figure 65-1 continued 1. [Figure 65-3] Continued from Figure 65-1 2. [Figure 66-1] Figure 66 shows the RT distribution of cDNA lengths on the modified templates, showing cDNA drop-off products, full-length, and non-templated addition (NTA). MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 66-2] Figure 66-1 continued 1. [Figure 66-3] Continued from Figure 66-12. [Figure 67-1] Figure 67 shows an analysis of RT non-templated addition (NTA) nucleotide incorporation selectivity on standard RNA templates. MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 67-2] Figure 67-1 continued 1. [Figure 67-3] Continued from Figure 67-12. [Figure 68-1] Figure 68 shows an analysis of RT non-templated addition (NTA) nucleotide incorporation selectivity on modified RNA templates. MMLV, TGIRT, and MarathonRT are referred to as Control 1, Control 2, and Control 3, respectively. [Figure 68-2] Figure 68-1 continued 1. [Figure 68-3] Continued from Figure 68-12. [Figure 69]Figure 69 shows an expression screen of MG140-8c5. A medium-throughput heterologous expression screen of MG140-8c5 in E. coli. The construct is expressed in small-scale culture flasks and induced at various temperatures in different growth media. Purification is performed in 24-deep-well plates, and the eluate is run on a gel for analysis. The data show that MG140-8c5 can be purified from the pMGE expression vector with the SUMO fusion, while expression in the pMGD vector with the MBP fusion does not produce the full-length protein. [Figure 70A] Figures 70A-70D show large-scale expression and purification of MG140-8c5. MG140-8c5 was induced overnight at either 16°C (Figures 70A-70B) or purified at 23.5°C for 5 hours (Figures 70C-70D) and eluted with an imidazole gradient on a 5 mL HisTrap. In elution profiles monitoring A280 and A260, significantly higher A260 levels were observed in the 23.5°C purification (presumably due to more nucleic acid contamination, Figure 70C) than in the 16°C purification (Figure 70A). Sample runs on gels revealed elution of the protein of interest (approximately 116 kDa) at relatively high imidazole concentrations. [Figure 70B] Figures 70A-70D show large-scale expression and purification of MG140-8c5. MG140-8c5 was induced overnight at either 16°C (Figures 70A-70B) or purified at 23.5°C for 5 hours (Figures 70C-70D) and eluted with an imidazole gradient on a 5 mL HisTrap. In elution profiles monitoring A280 and A260, significantly higher A260 levels were observed in the 23.5°C purification (presumably due to more nucleic acid contamination, Figure 70C) than in the 16°C purification (Figure 70A). Sample runs on gels revealed elution of the protein of interest (approximately 116 kDa) at relatively high imidazole concentrations. [Figure 70C]Figures 70A-70D show large-scale expression and purification of MG140-8c5. MG140-8c5 was induced overnight at either 16°C (Figures 70A-70B) or purified at 23.5°C for 5 hours (Figures 70C-70D) and eluted with an imidazole gradient on a 5 mL HisTrap. In elution profiles monitoring A280 and A260, significantly higher A260 levels were observed in the 23.5°C purification (presumably due to more nucleic acid contamination, Figure 70C) than in the 16°C purification (Figure 70A). Sample runs on gels revealed elution of the protein of interest (approximately 116 kDa) at relatively high imidazole concentrations. [Figure 70D] Figures 70A-70D show large-scale expression and purification of MG140-8c5. MG140-8c5 was induced overnight at either 16°C (Figures 70A-70B) or purified at 23.5°C for 5 hours (Figures 70C-70D) and eluted with an imidazole gradient on a 5 mL HisTrap. In elution profiles monitoring A280 and A260, significantly higher A260 levels were observed in the 23.5°C purification (presumably due to more nucleic acid contamination, Figure 70C) than in the 16°C purification (Figure 70A). Sample runs on gels revealed elution of the protein of interest (approximately 116 kDa) at relatively high imidazole concentrations. [Figure 71]Figure 71 shows the primer extension activity of MG140-8c5 on standard and modified RNA templates. Gel lane numbers correspond to the following: Lane 1: no RT control, standard template; Lane 2: no RT control, modified template; Lane 3: MMLV Control Enzyme 1, standard template, replicate 1; Lane 4: MMLV Control Enzyme 1, modified template, replicate 1; Lane 5: 140-8c5, standard template, replicate 1; Lane 6: 140-8c5, modified template, replicate 1; Lane 7: AccuScript Control Enzyme 4, standard template, replicate 1; Lane 8: AccuScript Control Enzyme 4, modified template, replicate 1; Lane 9: MMLV Control Enzyme 1, standard template, replicate 2; Lane 10: MMLV Control Enzyme 1, modified template, replicate 2; Lane 11: 140-8c5, standard template, replicate 2; Lane 12: 140-8c5, modified template, replicate 2; Lane 13: AccuScript Control Enzyme 4, standard template, replicate 2; and Lane 14: AccuScript Control Enzyme 4, modified template, replicate 2. [Figure 72A] Figures 72A-B show the use of fluorescence anisotropy to detect strand displacement during cDNA synthesis. Figure 72A: Substrate template RNA is annealed to a priming oligo and a displacement oligo conjugated to a FAM fluorophore. In the annealed state, the fluorophore tumbles slowly and the emitted light is not depolarized. Upon strand displacement, the much smaller oligo-conjugated FAM tumbles in solution much faster, depolarizing the light after emission. [Figure 72B] Figures 72A-72B show the use of fluorescence anisotropy to detect strand displacement during cDNA synthesis. Figure 72B: A reaction containing annealed substrate template and purified MG140-8c5 enzyme is initiated by the addition of dNTPs, allowing the enzyme to polymerize the cDNA and displace the FAM-labeled oligo, which is detectable by depolarization of the emitted light. In comparison, a substrate template without dNTPs emits polarized light due to the lack of displacement, while a FAM-oligo-only control emits highly depolarized light. [Figure 73A]Figures 73A-73B show the use of fluorescent unquenching to detect strand displacement during second-strand synthesis. Figure 73A: Substrate template ssDNA synthesized with a 5' FAM fluorophore conjugation is annealed to a priming oligo and a displacement oligo bearing a 3' quencher moiety. In the annealed state, the fluorophore is quenched by the quencher molecule of the displacement oligo, and fluorescence is low. Upon strand displacement, the template FAM is no longer quenched, and an increase in fluorescence is observed. [Figure 73B] Figures 73A-73B show the use of fluorescence unquenching to detect strand displacement during second-strand synthesis. Figure 73B: Reactions containing annealed substrate template and purified MG140-8c5 enzyme are initiated by the addition of dNTPs, allowing the enzyme to polymerize second-strand DNA, displacing the quenched oligonucleotide and resulting in an increase in fluorescence. In contrast, reactions without added dNTPs do not show the same gradual increase in fluorescence over time. [Figure 74A] Figures 74A-B show the use of strand displacement to measure enzyme activity from PURExpress. Figure 74A: A 1004 nt ssDNA template was generated by PCR amplification followed by lambda exonuclease digestion. [Figure 74B] Figures 74A-74B show the use of strand displacement to measure enzyme activity from PURExpress. Figure 74B: A fluorescence unquenched assay was set up using templates annealed to the priming and displacement oligos. The data show a rapid increase in fluorescence when PURExpress products MG153-5 and MG153-51 were added, but not when the no-template control (NTC) was added. [Figure 75]Figure 75 shows a schematic diagram of the template switching assay. RT initiates cDNA production at the 3' end of the donor RNA template. Acceptor RNA templates were used as an equimolar mixture of templates with different 3'-terminal nucleotides (NN-UU, AA, CC, GG) unless otherwise specified. cDNA products resulting from initiation (FAM probe) and template switching (HEX probe) are quantified by multiplexed Taqman qPCR. Template switching efficiency (TS%) is calculated as the percentage of cDNA detected by HEX divided by FAM. [Figure 76A] Figures 76A-B show template switching of GII intron RT using acceptor RNA with a terminal 3' UU nucleotide. Figure 76A: Amount of cDNA (nM) determined by FAM and HEX signals quantified by Taqman qPCR. RT is derived from a cell-free expression system. NTC is a no-template control, in which no RT-expressing template is provided to the cell-free expression system. Full is a control template in which the acceptor and donor sequences are ligated and FAM and HEX signals are expected to be equivalent. 10X A:D indicates that the acceptor RNA template (containing a 3' terminal UU nucleotide in this experiment) was used in 10-fold molar excess over the donor RNA template. [Figure 76B] Figures 76A-B show template switching of the GII intron RT using an acceptor RNA with a terminal 3' UU nucleotide. Figure 76B: The template switching efficiency of each RT is calculated as shown in Figure 75. In both figure panels, TGIRT (GII intron), MMLV (retrovirus), and MarathonRT (GII intron) are referred to as Control 1, 2, and 3, respectively. [Figure 77A]Figures 77A-77B show template switching of GII intron RT using acceptor RNA with mixed 3'-terminal nucleotides. Figure 77A: Amount of cDNA (nM) determined by FAM and HEX signals quantified by Taqman qPCR. RT is derived from a cell-free expression system. NTC is a no-template control; no RT-expressing template is provided to the cell-free expression system. Full is a control template in which the acceptor and donor sequences are ligated and FAM and HEX signals are expected to be equivalent. 10X A:D indicates that the acceptor RNA template (in this experiment, containing mixed 3'-terminal nucleotides, the preparation of which is described in the text) was used in a 10-fold molar excess over the donor RNA template. [Figure 77B] Figures 77A-77B show template switching of GII intron RT using acceptor RNAs with mixed 3'-terminal nucleotides. Figure 77B: The template switching efficiency of each RT is calculated as shown in Figure 75. In both figure panels, TGIRT (GII intron), MMLV (retrovirus), and MarathonRT (GII intron) are referred to as Control 1, 2, and 3, respectively. [Figure 78A]Figures 78A-78B show template switching of R2 MG140-8c5 using acceptor titration. Figure 78A: The amount of cDNA (nM) determined by FAM and HEX signals quantified by Taqman qPCR. Two buffers were tested to assess whether the buffer composition affected template switching activity. The composition of Buffer 1 is specified in this method because it is the primary buffer used in the template switching reaction. Buffer 2 consists of 40 mM Tris-HCl (pH 7.5), 0.2 M NaCl, 10 mM MgCl2, 1 mM TCEP, RNase inhibitor, and 0.5 mM dNTPs. No RT control reactions were performed for each buffer to establish signal background. MG140-8c5 was tested as purified protein. Full is a control template in which the acceptor and donor sequences are ligated and FAM and HEX signals are expected to be equivalent. 10X, 5X, and 1X A:D indicate that the acceptor RNA template (in this experiment containing mixed 3'-terminal nucleotides, the preparation of which is described in the text) was used at a 10-fold, 5-fold, or 1-fold molar excess over the donor RNA template. [Figure 78B] Figures 78A-78B show template switching of R2 MG140-8c5 using acceptor titration. Figure 78B: Template switching efficiency for 140-8c5 at 10-, 5-, and 1-fold molar excess of acceptor over donor in Buffer 1 or Buffer 2 is calculated as shown in Figure 75. [Figure 79-1]Figure 79 shows RT-primed and unprimed cDNA synthesis quantified by qPCR. The RNA template design is indicated at the top of the figure. The first two templates contain a 22-nt polyA sequence at the 3' end and are referred to as "A" in the bar graph below. The last two templates have an MS2 hairpin instead of polyA and are referred to as "MS2" in the bar graph below. The polyA and MS2 templates were tested with either a free 3' hydroxyl (denoted as 3'OH) or a free 3'OH block (denoted as 3'B, IDT 3' C3 spacer / 3SpC3 / ). Each template was also tested primed (P), meaning annealed to a 20-nt priming DNA oligo, or unprimed (UP). The dashed line represents the amount of cDNA 10-fold above the highest background negative control, PURExpress with no RT expression template (PUREx NTC). MMLV and TGIRT are referred to as Control 1 and Control 2, respectively. [Figure 79-2] Continuation of Figure 79-1. [Figure 80] Figure 80 shows primed cDNA synthesis activity divided by unprimed activity for each template listed in Figure 79, where A-3OH refers to a polyA sequence with a free 3' hydroxyl, A-3B refers to a polyA sequence with a blocked 3' OH, MS2-3OH refers to an MS2 sequence with a free 3' hydroxyl, and MS2-3B refers to an MS2 sequence with a blocked 3' OH. MMLV and TGIRT are referred to as Control 1 and Control 2, respectively. [Figure 81] Figure 81 shows primed cDNA synthesis activity divided by unprimed activity averaged across all four templates, as described in Figure 79. RTs are ordered by their preference for primed RNA templates. MMLV and TGIRT are referred to as Control 1 and Control 2, respectively. [Figure 82A]Figures 82A-B show assessment of primed and unprimed activity in MG140-8c5. Figure 82A: A 5'-labeled 100 nt RNA template annealed to quench the displacement oligo either in the presence or absence of the priming oligo was used as a substrate in reactions containing purified MG140-8c5 and initiated by the addition of dNTPs. The data show an increase in fluorescence for both cDNA synthesis and strand displacement, and for both primed and unprimed substrates with extension. [Figure 82B] Figures 82A-B show assessment of primed and unprimed activity in MG140-8c5. Figure 82B: A 5'-labeled 100 nt ssDNA template annealed to quench the displacement oligo either in the presence or absence of the priming oligo was used as a substrate in reactions containing purified MG140-8c5 and initiated by the addition of dNTPs. The data show an increase in fluorescence and both second-strand synthesis and strand displacement by extension for both primed and unprimed substrates. [Figure 83] Figure 83 shows a cladogram of reconstructed ancestral variants of the MG160 family of retron-like RTs. A phylogenetic tree was generated. [Figure 84] Figure 84 shows the typical in vitro cDNA synthesis activity of the MG157 retron RT. Lane numbers correspond to the following samples: 1: PURExpress no-template control, RT reaction containing no reverse transcriptase. 2: Positive control Group II intron RT TGIRT. Lanes 3-9 correspond to MG157 family candidates. Bold numbering corresponds to gel lanes with active candidates. Arrows indicate full-length cDNA products (arrow near the top of the gel) and examples of cDNA drop-off (arrowhead or vertical line below). [Figure 85A]Figures 85A-B show graphs depicting the results of screening the ability of directed and engineered candidate RTs to synthesize cDNA in mammalian cells. Figure 85A: Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100 bp PCR products amplified from cDNA synthesized from an RNA template by group II intron RTs of the MG165, MG166, MG167, and MG169 families. [Figure 85B] Figures 85A-85B show graphs depicting the results of screening the ability of directed RTs and engineered candidate RTs to synthesize cDNA in mammalian cells. Figure 85B: Taqman qPCR detection of the initial (FAM probe) and final (HEX probe) 100-bp PCR products amplified from cDNA synthesized from a 4.1-kb RNA template by rationally engineered variants of the non-LTR retrotransposon RTs MG140-74 and MG140-88. The lower dotted line indicates background (no RT control), and the upper dotted line indicates the maximum cDNA synthesis activity of the positive control RT TGIRT. Additional positive control RTs include MMLV WT and engineered RTs, as well as R2Tg. The variants of the MG140 family of RTs exhibiting the highest levels of cDNA synthesis activity are highlighted with a star. [Figure 86A] Figures 86A-86C show a schematic diagram and results of the template switching assay. Figure 86A shows a schematic diagram of the template switching assay. RT initiates cDNA production at the 3' end of the donor RNA template. Acceptor RNA templates were used as an equimolar mixture of templates with different 3'-terminal nucleotides (NN-UU, AA, CC, GG) unless otherwise specified. cDNA products resulting from initiation (FAM probe) and template switching (HEX probe) are quantified by multiplexed Taqman qPCR. Template switching efficiency (TS%) is calculated as the percentage of cDNA detected by HEX divided by FAM. [Figure 86B]Figures 86A-86C show a schematic and results of the template switching assay. Figure 86B shows a graph depicting the assay results as % template switch. Template switching efficiency was quantified for the non-LTR retrotransposase variants MG140-8c5, MG140-8(Endodead), and MG140-8, which has a dead endonuclease domain with a D451A mutation, as well as the group II intron RTs MG153-18 and MG153-51. The enzymes were purified prior to template switching experiments. [Figure 86C] Figures 86A-86C show a schematic diagram and results of the template switching assay. Figure 86C shows a graph depicting the assay results as % template switch. Template switching efficiency was quantified for the group II intron RTs MG153-18 and MG153-51, as well as for the non-LTR retrotransposase variant MG140-3 with a dead endonuclease domain (Endodead), and for MG140-3 with a dead endonuclease domain and D451A, F702A, and L698A mutations. The enzymes were expressed in a cell-free expression system.
[0087] Brief Description of Sequence Listing The Sequence Listing submitted herein provides exemplary polynucleotide and polypeptide sequences for use in the methods, compositions, and systems according to the present disclosure. Below are exemplary descriptions of the sequences therein.
[0088] MG140 SEQ ID NOs: 1-29, 393-401, 1476, 1850-1926, and 2165-2210 show the full-length peptide sequences of the MG140 translocation protein.
[0089] SEQ ID NOs: 374 to 386 show the nucleotide sequences of genes encoding HA-His-tagged MG140 reverse transcriptase proteins.
[0090] SEQ ID NOs: 761-798, 2161-2164, and 2211-2232 show the nucleotide sequences of MG140 UTRs.
[0091] SEQ ID NOs: 799 to 894 show the full-length peptide sequences of the MG140 reverse transcriptase protein.
[0092] SEQ ID NOs: 1535-1536, 1611-1623, 1663-1691, and 1786-1806 show the nucleotide sequences of genes encoding the MG140 reverse transcriptase protein, optimized for expression in mammalian cells.
[0093] SEQ ID NOs: 1542 to 1543 show the nucleotide sequence of the gene encoding the dead mutant MG140 reverse transcriptase protein, optimized for expression in mammalian cells.
[0094] MG146 SEQ ID NOs: 402 and 895 show the full-length peptide sequence of the MG146 translocating protein.
[0095] SEQ ID NO: 387 shows the nucleotide sequence of the gene encoding the HA-His tagged MG146 reverse transcriptase protein.
[0096] MG147 SEQ ID NO: 388 shows the nucleotide sequence of the gene encoding the HA-His tagged MG147 reverse transcriptase protein.
[0097] MG148 SEQ ID NOs: 403 to 426 show the full-length peptide sequences of the MG148 reverse transcriptase protein.
[0098] SEQ ID NOs: 389 to 392 show the nucleotide sequences of genes encoding HA-His-tagged MG148 reverse transcriptase proteins.
[0099] SEQ ID NOs: 1504 to 1507 show the nucleotide sequence of the gene encoding the MG148 reverse transcriptase protein, optimized for expression in mammalian cells.
[0100] MG149 SEQ ID NOs: 427 to 439 show the full-length peptide sequences of the MG149 reverse transcriptase protein.
[0101] MG151 SEQ ID NOs: 440 to 554 and 1020 to 1037 show the full-length peptide sequences of the MG151 reverse transcriptase protein.
[0102] SEQ ID NOs: 356 to 362 show the nucleotide sequences of genes encoding TwinStrep-tagged MG151 reverse transcriptase proteins.
[0103] SEQ ID NOs: 363 to 373 show the nucleotide sequences of genes encoding strep-tagged MG151 reverse transcriptase proteins.
[0104] SEQ ID NOs: 964 to 981 and 1003 to 1019 show the nucleotide sequences of the genes encoding the MG151 reverse transcriptase protein, optimized for expression in mammalian cells and cloned into a non-ligand plasmid.
[0105] MG153 SEQ ID NOs: 555-608 and 1927-2010 show the full-length peptide sequences of the MG153 reverse transcriptase protein.
[0106] SEQ ID NOs: 30 to 32 and 40 to 50 show the nucleotide sequences of fusion proteins containing the MG153 reverse transcriptase protein and the MS2 coat protein (MCP).
[0107] SEQ ID NOs: 66 to 119 show the nucleotide sequences of genes encoding strep-tagged MG153 reverse transcriptase proteins.
[0108] SEQ ID NOs: 120 to 173 show the nucleotide sequences of the E. coli codon-optimized gene encoding the MG153 reverse transcriptase protein.
[0109] SEQ ID NOs: 740 to 756 show the nucleotide sequences of genes encoding MCP-tagged MG153 reverse transcriptase proteins.
[0110] SEQ ID NOs: 1521-1534, 1624-1637, 1645-1662, and 1701-1782 show the nucleotide sequences of the genes encoding the MG153 reverse transcriptase protein, optimized for expression in mammalian cells and cloned into a non-ligand plasmid.
[0111] SEQ ID NOs: 1539 to 1541 show the nucleotide sequence of the gene encoding the dead mutant MG153 reverse transcriptase protein, optimized for expression in mammalian cells.
[0112] SEQ ID NOs: 2233 to 2257 show the nucleotide sequences of MG153 UTRs.
[0113] MG154 SEQ ID NOs: 609-610 and 1555 show the full-length peptide sequences of the MG154 reverse transcriptase protein.
[0114] SEQ ID NOs: 308 to 309 show the nucleotide sequences of the genes encoding the strep-tagged MG154 reverse transcriptase proteins.
[0115] SEQ ID NOs: 324-325 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG154 reverse transcriptase protein.
[0116] SEQ ID NOs: 340 to 341 show the nucleotide sequences of ncRNAs compatible with MG154 nuclease.
[0117] MG155 SEQ ID NOs: 611-615 and 1544-1545 show the full-length peptide sequences of the MG155 reverse transcriptase protein.
[0118] SEQ ID NOs: 310 to 312 and 1569 to 1570 show the nucleotide sequences of genes encoding strep-tagged MG155 reverse transcriptase proteins.
[0119] SEQ ID NOs: 326-328 and 1556-1557 show the nucleotide sequences of the E. coli codon-optimized gene encoding the MG155 reverse transcriptase protein.
[0120] SEQ ID NOs: 342-344 and 1582-1583 show the nucleotide sequences of ncRNAs compatible with MG155 nuclease.
[0121] MG156 SEQ ID NOs: 616-617 show the full-length peptide sequence of the MG156 reverse transcriptase protein.
[0122] SEQ ID NOs: 313 to 314 show the nucleotide sequences of the genes encoding the strep-tagged MG156 reverse transcriptase proteins.
[0123] SEQ ID NOs: 329-330 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG156 reverse transcriptase protein.
[0124] SEQ ID NOs: 345 to 346 show the nucleotide sequence of an ncRNA compatible with MG156 nuclease.
[0125] MG157 SEQ ID NOs: 618-622 and 2258-2266 show the full-length peptide sequences of the MG157 reverse transcriptase protein.
[0126] SEQ ID NOs: 315 to 319 show the nucleotide sequences of the genes encoding strep-tagged MG157 reverse transcriptase proteins.
[0127] SEQ ID NOs: 331 to 335 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG157 reverse transcriptase protein.
[0128] SEQ ID NOs: 347-351 and 1842-1849 show the nucleotide sequences of ncRNAs compatible with MG157 nuclease.
[0129] MG158 SEQ ID NO: 623 shows the full-length peptide sequence of the MG158 reverse transcriptase protein.
[0130] SEQ ID NO: 320 shows the nucleotide sequence of the gene encoding the strep-tagged MG158 reverse transcriptase protein.
[0131] SEQ ID NO: 336 shows the nucleotide sequence of the E. coli codon-optimized gene encoding the MG158 reverse transcriptase protein.
[0132] SEQ ID NO: 352 shows the nucleotide sequence of an ncRNA compatible with MG158 nuclease.
[0133] MG159 SEQ ID NOs: 624 to 626 show the full-length peptide sequences of the MG159 reverse transcriptase protein.
[0134] SEQ ID NOs: 321 to 323 show the nucleotide sequences of genes encoding strep-tagged MG159 reverse transcriptase proteins.
[0135] SEQ ID NOs: 337 to 339 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG159 reverse transcriptase protein.
[0136] SEQ ID NOs: 353 to 355 show the nucleotide sequences of ncRNAs compatible with MG159 nuclease.
[0137] SEQ ID NO: 1785 shows the nucleotide sequence of the target of the gene encoding the MG159 reverse transcriptase protein optimized for expression in mammalian cells.
[0138] MG160 SEQ ID NOs: 627-673, 1039-1475 and 2011-2026 show the full-length peptide sequences of the MG160 reverse transcriptase protein.
[0139] SEQ ID NOs: 174 to 180 show the nucleotide sequences of genes encoding strep-tagged MG160 reverse transcriptase proteins.
[0140] SEQ ID NOs: 181 to 187 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG160 reverse transcriptase protein.
[0141] SEQ ID NOs: 982-1002 show the nucleotide sequences of the genes encoding the MG160 reverse transcriptase protein, optimized for expression in mammalian cells and cloned into the combined spCas9(H840A) plasmid.
[0142] SEQ ID NOs: 1508 to 1520 show the nucleotide sequence of the gene encoding the MG160 reverse transcriptase protein, optimized for expression in mammalian cells.
[0143] MG163 SEQ ID NOs: 674 to 678 show the full-length peptide sequences of the MG163 reverse transcriptase protein.
[0144] SEQ ID NOs: 188 to 192 show the nucleotide sequences of genes encoding strep-tagged MG163 reverse transcriptase proteins.
[0145] SEQ ID NOs: 193 to 197 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG163 reverse transcriptase protein.
[0146] MG164 SEQ ID NOs: 679 to 683 show the full-length peptide sequences of the MG164 reverse transcriptase protein.
[0147] SEQ ID NOs: 198 to 202 show the nucleotide sequences of genes encoding strep-tagged MG164 reverse transcriptase proteins.
[0148] SEQ ID NOs: 203-207 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG164 reverse transcriptase protein.
[0149] MG165 SEQ ID NOs: 684-692 and 2027-2046 show the full-length peptide sequences of the MG165 reverse transcriptase protein.
[0150] SEQ ID NOs: 208 to 216 show the nucleotide sequences of genes encoding strep-tagged MG165 reverse transcriptase proteins.
[0151] SEQ ID NOs: 217-225 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG165 reverse transcriptase protein.
[0152] SEQ ID NOs: 757 to 759 show the nucleotide sequences of genes encoding MCP-tagged MG165 reverse transcriptase proteins.
[0153] MG166 SEQ ID NOs: 693-697 and 2047-2090 show the full-length peptide sequences of the MG166 reverse transcriptase protein.
[0154] SEQ ID NOs: 226 to 230 show the nucleotide sequences of genes encoding strep-tagged MG166 reverse transcriptase proteins.
[0155] SEQ ID NOs: 231-235 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG166 reverse transcriptase protein.
[0156] MG167 SEQ ID NOs: 698-702 and 2091-2120 show the full-length peptide sequences of the MG167 reverse transcriptase protein.
[0157] SEQ ID NOs: 236 to 240 show the nucleotide sequences of the genes encoding strep-tagged MG167 reverse transcriptase proteins.
[0158] SEQ ID NOs: 241-245 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG167 reverse transcriptase protein.
[0159] SEQ ID NOs: 759 to 760 show the nucleotide sequences of the genes encoding the MCP-tagged MG167 reverse transcriptase proteins.
[0160] MG168 SEQ ID NOs: 703 to 707 show the full-length peptide sequences of the MG168 reverse transcriptase protein.
[0161] SEQ ID NOs: 246 to 250 show the nucleotide sequences of genes encoding strep-tagged MG168 reverse transcriptase proteins.
[0162] SEQ ID NOs: 251-255 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG168 reverse transcriptase protein. MG169
[0163] SEQ ID NOs: 708-718 and 2121-2159 show the full-length peptide sequences of the MG169 reverse transcriptase protein.
[0164] SEQ ID NOs: 256 to 266 show the nucleotide sequences of genes encoding strep-tagged MG169 reverse transcriptase proteins.
[0165] SEQ ID NOs: 267-277 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG169 reverse transcriptase protein.
[0166] SEQ ID NOs: 1638-1644 and 1693-1700 show the nucleotide sequences of the genes encoding the MG169 reverse transcriptase protein, optimized for expression in mammalian cells.
[0167] MG170 SEQ ID NOs: 719 to 728 show the full-length peptide sequences of the MG170 reverse transcriptase protein.
[0168] SEQ ID NOs: 278 to 287 show the nucleotide sequences of genes encoding strep-tagged MG170 reverse transcriptase proteins.
[0169] SEQ ID NOs: 288-297 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG170 reverse transcriptase protein.
[0170] MG172 SEQ ID NOs: 729 to 733 show the full-length peptide sequences of the MG172 reverse transcriptase protein.
[0171] SEQ ID NOs: 298 to 302 show the nucleotide sequences of genes encoding strep-tagged MG172 reverse transcriptase proteins.
[0172] SEQ ID NOs: 303-307 show the nucleotide sequence of the E. coli codon-encoding gene encoding the optimized MG172 reverse transcriptase protein.
[0173] MG173 SEQ ID NOs: 734-735 and 1546-1553 show the full-length peptide sequences of the MG173 reverse transcriptase protein.
[0174] SEQ ID NOs: 1571 to 1580 show the nucleotide sequences of genes encoding strep-tagged MG173 reverse transcriptase proteins.
[0175] SEQ ID NOs: 1558 to 1567 show the nucleotide sequence of the E. coli codon-optimized gene encoding the MG173 reverse transcriptase protein.
[0176] SEQ ID NOs: 1584 to 1593 show the nucleotide sequences of ncRNAs compatible with MG173 nuclease.
[0177] SEQ ID NOs: 1783-1784 show the nucleotide sequence of the gene encoding the MG173 reverse transcriptase protein optimized for expression in mammalian cells.
[0178] MG176 SEQ ID NOs: 1038 and 2160 show the full-length peptide sequence of the MG176 retrotransposition protein.
[0179] SEQ ID NO: 1692 shows the nucleotide sequence of the gene encoding the MG176 reverse transcriptase protein optimized for expression in mammalian cells.
[0180] MG192 SEQ ID NO: 1554 shows the full-length peptide sequence of the MG192 reverse transcriptase protein.
[0181] SEQ ID NO: 1581 shows the nucleotide sequence of the gene encoding the strep-tagged MG192 reverse transcriptase protein.
[0182] SEQ ID NO: 1568 shows the nucleotide sequence of the E. coli codon-optimized gene encoding the MG192 reverse transcriptase protein.
[0183] SEQ ID NO: 1594 shows the nucleotide sequence of an ncRNA compatible with the MG192 nuclease.
[0184] Other arrays SEQ ID NOs: 736-738, 897-900, 927-928, 952-955, 1494-1497, 1595-1599, 1601-1604, 1809-1810, 1812-1815, and 1818-1819 show the nucleotide sequences of the primers.
[0185] SEQ ID NOs: 739, 901-902, 1498-1499, and 1605-1606 show the nucleotide sequences of Taqman probes for qPCR.
[0186] SEQ ID NOs: 896, 1493, and 1600 show the nucleotide sequences of RNA templates for cDNA synthesis.
[0187] SEQ ID NOs: 903 to 926 and 934 to 951 show the full-length sequences of chemically modified guide RNAs.
[0188] SEQ ID NOs: 929 and 932-933 show the nucleotide sequences of the cDNAs encoding the gene targets.
[0189] SEQ ID NO: 930 shows the nucleotide sequence of an exemplary RT-nickase linker.
[0190] SEQ ID NO: 931 shows the nucleotide sequence of MG3-6(H586A).
[0191] SEQ ID NOs: 956 to 963 show the nucleotide sequences of the reverse transcriptase cloned into the combined MG3-6(H586A) plasmid.
[0192] SEQ ID NOs: 1500-1502 and 1607-1610 show the nucleotide sequences of genes encoding reverse transcriptase proteins, optimized for expression in mammalian cells.
[0193] SEQ ID NOs: 1537-1538 show the nucleotide sequence of the gene encoding the dead mutant control reverse transcriptase protein optimized for expression in mammalian cells. DETAILED DESCRIPTION OF THE INVENTION
[0194] While various embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed.
[0195] The practice of some of the methods disclosed herein employs, unless otherwise indicated, immunological, biochemical, chemical, molecular biology, microbiology, cell biology, genomics, and recombinant DNA techniques.
[0196] As used in this disclosure, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Furthermore, to the extent the terms "including," "includes," "having," "has," "with," or variants thereof, are used in either the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0197] The term "about" or "approximately" refers to a range of acceptable error for a particular value as determined by one skilled in the art, which depends in part on how the value is measured or determined, for example, on the limitations of the measurement system. For example, "about" may refer to within one or more standard deviations, as is customary in the art. Alternatively, "about" may refer to a range of up to 20%, up to 15%, up to 10%, up to 5%, or up to 1% of a given value.
[0198] As used in this disclosure, the term "nucleotide" refers to a base-sugar-phosphate combination. Contemplated nucleotides include naturally occurring and synthetic nucleotides. Nucleotides are monomeric units of nucleic acid sequences (e.g., deoxyribonucleic acid (DNA) and ribonucleic acid (RNA)). The term nucleotide includes the ribonucleoside triphosphates adenosine triphosphate (ATP), uridine triphosphate (UTP), cytosine triphosphate (CTP), guanosine triphosphate (GTP), and deoxyribonucleoside triphosphates, such as dATP, dCTP, dITP, dUTP, dGTP, dTTP, or derivatives thereof. Such derivatives include, for example, [αS]dATP, 7-deaza-dGTP, and 7-deaza-dATP, as well as nucleotide derivatives that confer nuclease resistance to nucleic acid molecules containing them. As used in this disclosure, the term nucleotide encompasses dideoxyribonucleoside triphosphates (ddNTPs) and their derivatives. Illustrative examples of ddNTP include, but are not limited to, ddATP, ddCTP, ddGTP, ddITP and ddTTP.Nucleotide can be unlabeled or can be detectably labeled, for example, by using an optically detectable moiety (for example, fluorophore) or a moiety that includes quantum dot.Detectable label includes, for example, radioisotope, fluorescent label, chemiluminescent label, bioluminescent label and enzyme label. Fluorescent labels for nucleotides include, but are not limited to, fluorescein, 5-carboxyfluorescein (FAM), 2'7'-dimethoxy-4'5-dichloro-6-carboxyfluorescein (JOE), rhodamine, 6-carboxyrhodamine (R6G), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), 6-carboxy-X-rhodamine (ROX), 4-(4'dimethylaminophenylazo)benzoic acid (DABCYL), Cascade Blue, Oregon Green, Texas Red, cyanine, and 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS).Specific examples of fluorescently labeled nucleotides include [R6G]dUTP, [TAMRA]dUTP, [R110]dCTP, [R6G]dCTP, [TAMRA]dCTP, [JOE]ddATP, [R6G]ddATP, [FAM]ddCTP, [R110]ddCTP, [TAMRA]ddGTP, [ROX]ddTTP, [dR6G]ddATP, [dR110]ddCTP, [dTAMRA]ddGTP, and [dROX]ddTTP available from Perkin Elmer (Foster City, Calif.); FluoroLink DeoxyNucleotides, FluoroLink Cy3-dCTP, FluoroLink Cy5-dCTP, FluoroLink Fluor X-dCTP, FluoroLink Cy3-dUTP, and FluoroLink available from Amersham (Arlington Heights, IL). Cy5-dUTP, fluorescein-15-dATP, fluorescein-12-dUTP, tetramethyl-rhodamine-6-dUTP, IR770-9-dATP, fluorescein-12-ddUTP, fluorescein-12-UTP, and fluorescein-15-2'-dATP available from Boehringer (Mannheim, Indianapolis, Ind.), and Molecular Examples of chromosomal labeled nucleotides available from Probes (Eugene, Oreg.) include BODIPY-FL-14-UTP, BODIPY-FL-4-UTP, BODIPY-TMR-14-UTP, BODIPY-TMR-14-dUTP, BODIPY-TR-14-UTP, BODIPY-TR-14-dUTP, Cascade Blue-7-UTP, Cascade Blue-7-dUTP, Fluorescein-12-UTP, Fluorescein-12-dUTP, Oregon Green 488-5-dUTP, Rhodamine Green-5-UTP, Rhodamine Green-5-dUTP, Tetramethylrhodamine-6-UTP, Tetramethylrhodamine-6-dUTP, Texas Red-5-UTP, Texas Red-5-dUTP, and Texas Red-12-dUTP. The term nucleotide encompasses chemically modified nucleotides. An exemplary chemically modified nucleotide is biotin-dNTP.Non-limiting examples of biotinylated dNTPs include biotin-dATP (e.g., bio-N6-ddATP, biotin-14-dATP), biotin-dCTP (e.g., biotin-11-dCTP, biotin-14-dCTP), and biotin-dUTP (e.g., biotin-11-dUTP, biotin-16-dUTP, biotin-20-dUTP).
[0199] The terms "polynucleotide," "oligonucleotide," and "nucleic acid" are used interchangeably to refer to polymeric forms of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof, in single-, double-, or multiple-stranded form. Contemplated polynucleotides include genes or fragments thereof. Exemplary polynucleotides include, but are not limited to, DNA, RNA, coding or non-coding regions of genes or gene fragments, multiple loci (single locus) defined from binding analysis, exons, introns, messenger RNA (mRNA), transfer RNA (tRNA), ribosomal RNA (rRNA), short interfering RNA (siRNA), short hairpin RNA (shRNA), micro-RNA (miRNA), ribozymes, cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, cell-free polynucleotides, including cell-free DNA (cfDNA) and cell-free RNA (cfRNA), nucleic acid probes, and primers. When referring to T, in polynucleotides, T refers to U (uracil) in RNA and T (thymine) in DNA. Polynucleotides can be exogenous or endogenous to cells and / or exist in a cell-free environment. The term polynucleotide encompasses modified polynucleotides (e.g., modified backbones, sugars, or nucleobases). When present, modifications to the nucleotide structure are imparted before or after assembly of the polymer. Non-limiting examples of modifications include 5-bromouracil, peptide nucleic acids, heterologous nucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-conjugated nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. The sequence of nucleotides can be interrupted by non-nucleotide components.
[0200] The term "transfection" or "transfected" refers to the introduction of a nucleic acid into a cell by non-viral or viral-based methods. The nucleic acid molecule may be a gene sequence encoding an entire protein or a functional portion thereof.
[0201] As used in this disclosure, "non-naturally occurring" refers to a nucleic acid or polypeptide sequence that does not occur in nature. Non-naturally occurring refers to a nucleic acid or polypeptide sequence that does not occur in nature, including modifications such as mutations, insertions, or deletions. The term non-naturally occurring encompasses fusion nucleic acids or polypeptides that encode or exhibit an activity (e.g., an enzymatic activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) of the nucleic acid or polypeptide sequence to which the non-naturally occurring sequence is fused. Non-naturally occurring nucleic acid or polypeptide sequences include those that are linked by genetic engineering to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) to generate a chimeric nucleic acid or polypeptide sequence that encodes the chimeric nucleic acid or polypeptide.
[0202] As used in this disclosure, "non-naturally occurring" can also refer to a nucleic acid or polypeptide sequence that is not found in a natural nucleic acid or protein. Non-naturally occurring can refer to an affinity tag. Non-naturally occurring can refer to a fusion. Non-naturally occurring can refer to a naturally occurring nucleic acid or polypeptide sequence that includes a mutation, insertion, or deletion. A non-naturally occurring sequence can exhibit or encode an activity (e.g., an enzyme activity, a methyltransferase activity, an acetyltransferase activity, a kinase activity, a ubiquitination activity, etc.) that can also be exhibited by the nucleic acid or polypeptide sequence to which the non-naturally occurring sequence is fused. A non-naturally occurring nucleic acid or polypeptide sequence can be linked to a naturally occurring nucleic acid or polypeptide sequence (or a variant thereof) by genetic engineering to generate a chimeric nucleic acid or polypeptide sequence that encodes a chimeric nucleic acid or polypeptide.
[0203] As used in this disclosure, the term "promoter" refers to a regulatory DNA region that controls the transcription or expression of a polynucleotide (e.g., a gene) and may be located adjacent to or overlapping the nucleotide or region of nucleotide at which RNA transcription is initiated. A promoter may contain specific DNA sequences that bind protein factors, often referred to as transcription factors, that facilitate the binding of RNA polymerase to DNA, resulting in gene transcription. Basal promoters in eukaryotes typically, but not necessarily, contain a TATA box and / or a CAAT box.
[0204] The term "expression," as used in this disclosure, refers to the process by which a nucleic acid sequence or polynucleotide is transcribed from a DNA template (e.g., into mRNA or other RNA transcript) and / or the process by which a transcribed mRNA is subsequently translated into a peptide, polypeptide, or protein. The transcript and the encoded polypeptide may be collectively referred to as a "gene product." If the polynucleotide is derived from genomic DNA, expression may include splicing of the mRNA in a eukaryotic cell.
[0205] As used in this disclosure, "operably linked," "operable linkage," "operatively linked," or their grammatical equivalents refer to the arrangement of genetic elements, e.g., promoters, enhancers, polyadenylation sequences, etc., such that the operation (e.g., movement or activation) of a first genetic element has some effect on a second genetic element. The effect on the second genetic element can be, but need not be, of the same type as the operation of the first genetic element. For example, two genetic elements are operably linked if movement of the first element causes activation of the second element. A regulatory element, which may include, for example, a promoter and / or enhancer sequence, is operably linked with a coding region if the regulatory element helps to initiate transcription of the coding sequence. There may be intervening residues between the regulatory element and the coding region, so long as this functional relationship is maintained.
[0206] As used in this disclosure, "vector" refers to a polymer or a polymeric association that contains or is associated with a polynucleotide and mediates the delivery of the polynucleotide to a cell. Examples of vectors include nucleic acid-based vectors (e.g., plasmids and viral vectors) and liposomes. Exemplary nucleic acid-based vectors generally include genetic elements, such as regulatory elements, operably linked to a gene to promote the expression of the gene in a target.
[0207] As used in this disclosure, "expression cassette" and "nucleic acid cassette" are used interchangeably to refer to a component of a vector that contains a combination of nucleic acid sequences or elements (e.g., therapeutic gene, promoter, and terminator) that are expressed together or operably linked for expression. The term encompasses expression cassettes that contain a combination of one or more genes with regulatory elements that are operably linked for expression.
[0208] A "functional fragment" of a DNA or protein sequence refers to a fragment that retains a biological activity (either functional or structural) substantially similar to that of the full-length DNA or protein sequence. The biological activity of a DNA sequence includes the ability to affect expression in a manner attributable to the full-length sequence.
[0209] The terms "engineered," "synthetic," and "artificial" are used interchangeably herein to refer to entities modified by human intervention. For example, the terms refer to polynucleotides or polypeptides that do not occur in nature. Engineered peptides may, but need not, have low sequence identity (e.g., less than 50% sequence identity, less than 25% sequence identity, less than 10% sequence identity, less than 5% sequence identity, less than 1% sequence identity) to naturally occurring human proteins. For example, the VPR domain and the VP64 domain are synthetic transactivation domains. Non-limiting examples include: nucleic acids that are modified by changing their sequence to one that does not occur in nature; nucleic acids that are modified by ligating to a nucleic acid with which they are not naturally associated such that the ligated product possesses a function not present in the original nucleic acid; engineered nucleic acids that are synthesized in vitro using a sequence that does not occur in nature; proteins that are modified by changing their amino acid sequence to a sequence that does not occur in nature; and engineered proteins that acquire new functions or properties. An "engineered" system includes at least one engineered component.
[0210] As used in this disclosure, the term "transposable element" refers to a DNA sequence that can move from one location to another within a genome (i.e., they can "transpose"). Transposable elements can be broadly divided into two classes: Class I transposable elements, or "retrotransposons," are transposed via transcription and translation of an RNA intermediate and then reintegrated into the genome at their new location via reverse transcription (a process mediated by reverse transcriptase). Class II transposable elements, or "DNA transposons," are transposed via a complex of single- or double-stranded DNA flanked on both sides by transposase.
[0211] As used in this disclosure, the term "retrotransposon" refers to a class I transposable element that functions according to a bipartite "copy-and-paste" mechanism involving an RNA intermediate. "Retrotransposase" refers to an enzyme involved in the transposition of retrotransposons. Retrotransposases can contain a reverse transcriptase domain, one or more zinc finger domains, an endonuclease domain, or a combination thereof.
[0212] As used in this disclosure, the terms "gene editing" and "genome editing" can be used interchangeably. Gene editing or genome editing refers to changing the nucleic acid sequence of a gene or genome. Genome editing can include, for example, insertion, deletion, and mutation. Genome editing can be performed by a gene editing system, for example, a retrotransposase.
[0213] As used in this disclosure, the term "complex" refers to the joining of at least two components. Each of the two components may retain the properties / activities it had before forming the complex, or may acquire properties as a result of forming the complex. Joining may be by, but is not limited to, covalent bonding, non-covalent bonding (i.e., hydrogen bonding, ionic interactions, van der Waals interactions, and hydrophobic bonding), the use of a linker, fusion, or any other suitable method. Contemplated components of the complex include polynucleotides, polypeptides, or combinations thereof. For example, the complex may include an endonuclease and a guide polynucleotide.
[0214] The term "sequence identity" or "percent identity" in the context of two or more nucleic acid or polypeptide sequences refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences that are identical or have a specified percentage of identical amino acid residues or nucleotides when compared and aligned for maximum correspondence over a local or global comparison window, as measured using a sequence comparison algorithm. Suitable sequence comparison algorithms for polypeptide sequences include, for example, BLASTP using the BLOSUM62 scoring matrix with parameters of a word length (W) of 3, an expectation (E) of 10, and gap costs set at 11, an extension of 1, and a conditional composition score matrix adjustment for polypeptide sequences longer than 30 residues; BLASTP using parameters of a word length (W) of 2, an expectation (E) of 1,000,000, and PAM30 scoring with gap costs set at 9 for open gaps and 1 for extended gaps for sequences shorter than 30 residues (these are the default parameters for BLASTP in the BLAST suite available at https: / / blast.ncbi.nlm.nih.gov); CLUSTALW with Smith-Waterman homology search algorithm parameters of 2 matches, -1 mismatches, and -1 gaps; MUSCLE with default parameters; MAFFT with parameters of 2 retries and 1000 maximum repeats; Novafold with default parameters; and HMMER hmmalign with default parameters.
[0215] In the context of two or more nucleic acid or polypeptide sequences, the term "optimally aligned" refers to two (e.g., in a pairwise alignment) or more (e.g., in a multiple sequence alignment) sequences aligned for maximum amino acid residue or nucleotide correspondence, as determined, for example, by the alignment producing the highest or "optimized" percent identity score.
[0216] The term "open reading frame" or "ORF" refers to a nucleotide sequence that can encode a protein, or a portion of a protein. An open reading frame can begin with a start codon (e.g., represented in standard codes as AUG for RNA molecules and ATG for DNA molecules) and can be read with codon triplets until the frame ends with a stop codon (e.g., represented in standard codes as UAA, UGA, or UAG for RNA molecules and TAA, TGA, or TAG for DNA molecules).
[0217] Any variant of the enzyme described in this disclosure that has one or more conservative amino acid substitutions is included in the present disclosure.Such conservative substitutions can be made in the amino acid sequence of a polypeptide without disrupting the three-dimensional structure or function of the polypeptide.Conservative substitutions can be achieved by substituting amino acids with similar hydrophobicity, polarity, and R chain length.In addition, or alternatively, by comparing the aligned sequences of homologous proteins from different species, conservative substitutions can be identified by finding amino acid residues (e.g., non-conserved residues) that are mutated between species without changing the basic function of the encoded protein. Such conservatively substituted variants may include variants having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity to any one of the retrotransposase protein sequences described herein (e.g., an MG140 family retrotransposase described herein, or any other family retrotransposase described herein). In some embodiments, such conservatively substituted variants are functional variants. Such functional variants can include sequences with substitutions that do not disrupt the activity of one or more important active site residues of retrotransposase. In some embodiments, a functional variant of any of the proteins described in the present disclosure lacks substitution of at least one conserved or functional residue.In some embodiments, a functional variant of any of the proteins described in this disclosure lacks all conserved or functional residue substitutions.
[0218] The present disclosure also includes variants (e.g., reduced activity variants) of any of the enzymes described herein that have substitutions of one or more catalytic residues to reduce or eliminate the activity of the enzyme. In some embodiments, reduced activity variants of the proteins described herein include at least one, at least two, or all three catalytic residues.
[0219] Conservative substitution tables resulting in functionally similar amino acids are available in various references (see, for example, Creighton, Proteins: Structures and Molecular Properties (W.H. Freeman & Co.; 2nd edition (December 1993)). The following eight groups each contain amino acids that are conservative substitutions for one another: 1) alanine (A), glycine (G); 2) aspartic acid (D), glutamic acid (E); 3) asparagine (N), glutamine (Q); 4) arginine (R), lysine (K); 5) isoleucine (I), leucine (L), methionine (M), valine (V); 6) phenylalanine (F), tyrosine (Y), tryptophan (W); 7) serine (S), threonine (T); and 8) Cysteine (C), methionine (M).
[0220] Variants of any of the nucleic acid sequences described herein that have one or more substitutions, deletions, or insertions are also included in the present disclosure. In some embodiments, such variants have a sequence that has at least about 80%, at least about 81%, at least about 82%, at least about 83%, at least about 84%, at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or 100% sequence identity with any one of the nucleic acid sequences described herein.
[0221] Some of the protein sequences described herein involve the determination of specific domains (e.g., reverse transcriptase or RT domains) from the sequence of a selected larger protein (e.g., a retrotransposase). In such cases, multiple sequence alignment (MSA) with a reference larger protein (e.g., a retrotransposase) whose domains have been validated (e.g., in a 3D structure) is used to identify domain boundaries by aligning the selected protein with the larger protein containing the validated domains. When sequences are highly divergent and the MSA is inconclusive, the 3D structure of the larger protein is determined, and the structural domains are compared to known domains to define the boundaries. These boundaries can be further validated by ensuring the presence of critical catalytic residues for the domain within the domain boundaries.
[0222] As used in this disclosure, the term "LINE retrotransposase" refers to a class of autonomous non-LTR retrotransposons (Long Interspersed Elements). As used in this disclosure, the terms "R2 retrotransposase" or "R4 retrotransposase" refer to a subclass of LINE retrotransposases that share a similar domain architecture but differ in that R2 retrotransposases can be site-specific (e.g., integrate at specific sites in rRNA genes), while R4 retrotransposons can integrate at both rRNA genes as well as other non-specific sites containing repeats.
[0223] overview The discovery of new transposable elements with unique functionality and structure could further disrupt deoxyribonucleic acid (DNA) editing technologies, offering the potential to improve their speed, specificity, functionality, and ease of use. Compared to the predicted prevalence of transposable elements in microorganisms and indeed in a wide variety of microbial species, there are relatively few functionally characterized transposable elements in the literature. This is in part due to the inability to readily cultivate the vast number of microbial species under laboratory conditions. Metagenomic sequencing from natural environmental niches containing large numbers of microbial species could dramatically increase the number of novel transposable elements described in the literature, potentially accelerating the discovery of novel oligonucleotide editing functions.
[0224] Transposable elements are deoxyribonucleic acid sequences that can change position within a genome, often resulting in the generation or improvement of mutations. In eukaryotes, the majority of the genome and the majority of the cellular DNA mass are attributable to transposable elements. Although transposable elements are "selfish genes" that propagate themselves at the expense of other genes, they perform a variety of important functions and have been found to be important in genome evolution. Based on their mechanism, transposable elements are classified as either class I "retrotransposons" or class II "DNA transposons."
[0225] Class I transposable elements, also known as retrotransposons, function according to a two-part "copy-and-paste" mechanism involving an RNA intermediate. First, the retrotransposon is transcribed. The resulting RNA is then converted back into DNA by a reverse transcriptase (generally encoded by the retrotransposon itself), and the reverse-transcribed retrotransposon is integrated into its new location within the genome by an integrase. Retrotransposons are further classified into three lineages. Long terminal repeat ("LTR") retrotransposons encode reverse transcriptase and are flanked by long stretches of repetitive DNA. Long interspersed nucleotide sequences ("LINE") retrotransposons encode reverse transcriptase, lack LTRs, and are transcribed by RNA polymerase II. Short interspersed nucleotide sequences ("SINE") retrotransposons are transcribed by RNA polymerase III but lack reverse transcriptase and instead rely on the reverse transcription machinery of other transposable elements (e.g., LINEs).
[0226] Class II transposable elements, also known as DNA transposons, function according to a mechanism that does not involve an RNA intermediate. Many DNA transposons exhibit a "cut-and-paste" mechanism, in which a transposase binds to terminal inverted repeats ("TIRs") flanking the transposon, cleaves the transposon from the donor region, and inserts it into the target region of the genome. Others, called "helitrons," involve a single-stranded DNA intermediate and exhibit a "rolling circle" mechanism mediated by an undocumented protein understood to possess HUH endonuclease function and 5' to 3' helicase activity. First, a circular strand of DNA is nicked to create two single DNA strands. The protein remains attached to the 5' phosphate of the nicked strand, leaving the 3' hydroxyl end of the complementary strand exposed, thus allowing polymerase to replicate the unnicked strand. Once replication is complete, the new strand dissociates and replicates itself alongside the original template strand. Yet another DNA transposon, "pollinton," is theorized to undergo a "self-synthesizing" mechanism. Transposition is initiated by integrase excision of single-stranded extrachromosomal Pollinator elements, which form racket-like structures. Pollinator elements undergo replication by DNA polymerase B, and double-stranded Pollinator elements are inserted into the genome by integrase. In addition, some DNA transposons, such as those of the IS200 / IS605 family, proceed via a "peel and paste" mechanism in which TnpA excises a piece of single-stranded DNA (as a circular "transposon junction") from the lagging strand template of the donor gene and reinserts it into the replication fork of the target gene.
[0227] While transposable elements have found some use as biological tools, the transposable elements described in the literature do not encompass the full range of possible biodiversity and targeting possibilities, and may not represent all possible activities. Here, we extracted thousands of genome fragments from multiple metagenomes for transposable elements. The diversity of transposable elements described in the literature may be expanded, and systems may be developed into highly targetable, compact, and precise gene editing agents.
[0228] Retrons are bacterial retroelements that produce single-stranded reverse-transcribed DNA (RT-DNA), a key part of the newly discovered phage defense system. Retrons possess the unique ability to generate multicopy single-stranded DNA (msDNA) consisting of a single strand of structured RNA, 'msr,' bound to a single strand of DNA, and 'msd,' flanked by two inverse-complementary repeats (5' IRa1 and 3' IRa2; Figure 50). msr and msd are encoded in a compact, contiguous transcription cassette that also contains a dedicated reverse transcriptase (RT; approximately 300–400 amino acids); this cassette is referred to as the entire retron (Figure 50). RT initiates reverse transcription using the base-paired 5' and 3' stems (IRa1 and IRa2) of msr-msd as primers, as well as a conserved priming guanosine within the conserved AGC sequence at the 3' end of msr. The msr and msd molecules are linked by a 2'-5' phosphodiester bond between the priming guanosine and the phosphate at the 5' end of the msd, covalently linking the RNA and DNA strands into a single branched molecule (Figure 50). The exact mechanism of termination is not yet understood. RT extends the reverse transcript to a defined position where reverse transcription stops via an unknown mechanism. RT has been observed to terminate at similar sites in vitro and in cells, strongly suggesting that RNA structure can induce RT termination (Shimamoto T. et al., 1995; Simon A. et al., 2019). Interestingly, retron reverse transcriptase (RT) typically lacks an RNase H domain and therefore relies on endogenous RNase H1 to remove the RNA template from RT-DNA. Concurrent with reverse transcription, cellular RNAse H1 activity degrades the msd template, except for approximately 5-10 RNA bases at its 5' end. This segment of RNA remains hybridized to the complementary reverse transcript and is considered to be part of both the msr and msd forms of mature msDNA (Figure 50).
[0229] Retrons can be a powerful tool for genome editing because they can produce high copy numbers of intracellular DNA molecules in the host. Previous experiments have shown that a retron from E. coli (Ec67) msr and RT can successfully reverse transcribe another retron (Ec73) msd. This experiment demonstrated that while the msr and associated RT of a particular retron are always paired and essential for initiating reverse transcription, the msd can be variable and can encode in situ DNA with a desired artificial sequence. This important discovery may enable the repurposing of retrons for biotechnological and therapeutic applications.
[0230] MG enzyme In some aspects, the present disclosure provides retrotransposases. In some embodiments, the retrotransposase is MG140, MG146, MG147, MG148, MG149, MG151, MG153, MG154, MG155, MG156, MG157, MG158, MG159, MG160, MG163, MG164, MG165, MG166, MG167, MG168, MG169, MG170, MG172, MG173, or MG176 retrotransposase (FIG. 1). In some embodiments, these retrotransposases are less than about 1,400 amino acids in length. In some embodiments, these retrotransposases simplify delivery and expand therapeutic applications.
[0231] In some embodiments, the present disclosure provides engineered retrotransposase systems discovered through metagenomic sequencing. In some embodiments, metagenomic sequencing is performed on samples. In some embodiments, the samples are collected from various environments. In some embodiments, such environments are human microbiomes, animal microbiomes, hot environments, and cold environments. In some embodiments, the environments include sediments.
[0232] In some embodiments, the present disclosure provides an engineered retrotransposase system comprising a retrotransposase derived from an uncultured microorganism. In some embodiments, the retrotransposase is configured to bind to a 3' untranslated region (UTR). In some embodiments, the retrotransposase binds to a 5' untranslated region (UTR).
[0233] In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having 100% identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
[0234] In some embodiments, the retrotransposase is MG140 retrotransposase (i.e., SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210.In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210.
[0235] In some embodiments, the retrotransposase is MG146 retrotransposase (i.e., SEQ ID NO:402 or SEQ ID NO:895). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO:402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to SEQ ID NO:402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 75% identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to SEQ ID NO: 402 or 895.In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to SEQ ID NO: 402 or 895. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to SEQ ID NO: 402 or 895.
[0236] In some embodiments, the retrotransposase is MG148 retrotransposase (i.e., SEQ ID NOS: 403-426). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOS: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOS: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 403-426.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 403-426. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 403-426.
[0237] In some embodiments, the retrotransposase is MG149 retrotransposase (i.e., SEQ ID NOS: 427-439). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOS: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOS: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 427-439.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 427-439. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 427-439.
[0238] In some embodiments, the retrotransposase is MG151 retrotransposase (i.e., SEQ ID NOs: 440-554 and 1020-1037). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037.
[0239] In some embodiments, the retrotransposase is MG153 retrotransposase (i.e., SEQ ID NOs: 555-608 and 1927-2010). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010.
[0240] In some embodiments, the retrotransposase is MG154 retrotransposase (i.e., SEQ ID NOs: 609-610 and 1555). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 609-610 and 1555.In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 609-610 and 1555. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 609-610 and 1555.
[0241] In some embodiments, the retrotransposase is an MG155 retrotransposase (i.e., SEQ ID NOs: 611-615 and 1544-1545). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 611-615 and 1544-1545.
[0242] In some embodiments, the retrotransposase is MG156 retrotransposase (i.e., SEQ ID NO:616 or SEQ ID NO:617). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to SEQ ID NO:616 or SEQ ID NO:617. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to SEQ ID NO:616 or SEQ ID NO:617.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to SEQ ID NO: 616 or SEQ ID NO: 617. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to SEQ ID NO: 616 or SEQ ID NO: 617. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to SEQ ID NO: 616 or SEQ ID NO: 617.
[0243] In some embodiments, the retrotransposase is MG157 retrotransposase (i.e., SEQ ID NOs: 618-622 and 2258-2266). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266.
[0244] In some embodiments, the retrotransposase is MG158 retrotransposase (i.e., SEQ ID NO: 623). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 70% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 80% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 85% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 90% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 95% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 96% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 97% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 98% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having at least about 99% identity to SEQ ID NO: 623. In some embodiments, the retrotransposase comprises a sequence having 100% identity to SEQ ID NO: 623.
[0245] In some embodiments, the retrotransposase is MG159 retrotransposase (i.e., SEQ ID NOS: 624-626). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOS: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOS: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 624-626.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 624-626. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 624-626.
[0246] In some embodiments, the retrotransposase is MG160 retrotransposase (i.e., SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 70% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 75% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 80% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 85% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 90% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026.In some embodiments, the retrotransposase comprises a sequence having at least about 95% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 96% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 97% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 98% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having at least about 99% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026. In some embodiments, the retrotransposase comprises a sequence having 100% identity to any one of SEQ ID NOs: 627-673, and 1039-1475, and 2011-2026.
[0247] In some embodiments, the retrotransposase is MG163 retrotransposase (i.e., SEQ ID NOs: 674-678). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 674-678.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 674-678. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 674-678.
[0248] In some embodiments, the retrotransposase is MG164 retrotransposase (i.e., SEQ ID NOs: 679-683). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 679-683.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 679-683. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 679-683.
[0249] In some embodiments, the retrotransposase is MG165 retrotransposase (i.e., SEQ ID NOs: 684-692 and 2027-2046). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046.
[0250] In some embodiments, the retrotransposase is MG166 retrotransposase (i.e., SEQ ID NOs: 693-697 and 2047-2090). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 693-697 and 2047-2090.
[0251] In some embodiments, the retrotransposase is MG167 retrotransposase (i.e., SEQ ID NOs: 698-702 and 2091-2119). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119.
[0252] In some embodiments, the retrotransposase is MG168 retrotransposase (i.e., SEQ ID NOS: 703-707). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOS: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOS: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 703-707.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 703-707. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 703-707.
[0253] In some embodiments, the retrotransposase is MG169 retrotransposase (i.e., SEQ ID NOs: 708-718 and 2121-2159). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 708-718 and 2121-2159.
[0254] In some embodiments, the retrotransposase is MG170 retrotransposase (i.e., SEQ ID NOs: 719-728). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs:719-728.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 719-728. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 719-728.
[0255] In some embodiments, the retrotransposase is MG172 retrotransposase (i.e., SEQ ID NOs: 729-733). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 729-733.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 729-733. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 729-733.
[0256] In some embodiments, the retrotransposase is MG173 retrotransposase (i.e., SEQ ID NOs: 734-735 and 1546-1553). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553.In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553.
[0257] In some embodiments, the retrotransposase is MG176 retrotransposase (i.e., SEQ ID NO: 1038 or SEQ ID NO: 2160). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to SEQ ID NO:1038 or SEQ ID NO:2160.In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160. In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160.
[0258] In some embodiments, the retrotransposase is MG192 retrotransposase (i.e., SEQ ID NO: 1554). In some embodiments, the retrotransposase comprises a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 70% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 75% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 80% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 85% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 90% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 95% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 96% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 97% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 98% sequence identity to SEQ ID NO: 1554. In some embodiments, the retrotransposase comprises a sequence having at least about 99% sequence identity to SEQ ID NO:1554.In some embodiments, the retrotransposase comprises a sequence having 100% sequence identity to SEQ ID NO:1554.
[0259] In some embodiments, the retrotransposase is encoded by a codon-optimized nucleic acid sequence. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence that is codon-optimized for expression in mammalian cells. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence that is codon-optimized for expression in mammalian cells. In some embodiments, the retrotransposase is at least about 20%, at least about 25%, or at least about 30% identical to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. , at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 70% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 75% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 85% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 95% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, and 1556-1568. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 96% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, and 1556-1568. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 97% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 98% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.In some embodiments, the retrotransposase is encoded by a nucleic acid sequence having at least 99% sequence identity to any one of the nucleic acid sequences set forth in SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806. In some embodiments, the retrotransposase is encoded by any one of the nucleic acid sequences of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
[0260] In some embodiments, the retrotransposase comprises a reverse transcriptase domain. In some embodiments, the retrotransposase further comprises one or more zinc finger domains. In some embodiments, the retrotransposase further comprises an endonuclease finger domain. In some embodiments, the retrotransposase further comprises a conserved catalytic D, QG, [Y / F]XDD, or LG motif. In some embodiments, the retrotransposase further comprises a conserved CX [2-3] C further contains a Zn finger motif.
[0261] In some embodiments, the retrotransposase has less than about 90%, less than about 85%, less than about 80%, less than about 75%, less than about 70%, less than about 65%, less than about 60%, less than about 55%, less than about 50%, less than about 45%, less than about 40%, less than about 35%, less than about 30%, less than about 25%, less than about 20%, less than about 15%, less than about 10%, or less than about 5% sequence identity to retrotransposases described in the literature.
[0262] In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0263] In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a single-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence as a double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the retrotransposase is configured to transpose the cargo nucleotide sequence via a ribonucleic acid polynucleotide intermediate.
[0264] In some embodiments, the retrotransposase comprises one or more nuclear localization sequences (NLSs), hi some embodiments, the NLSs are proximal to the N-terminus or C-terminus of the retrotransposase. In some embodiments, an NLS is attached to the N-terminus or C-terminus of the retrotransposase and comprises any one of SEQ ID NOs: 1477-1492, or a sequence having at least about 20%, at least about 25%, at least about 30%, at least about 35%, at least about 40%, at least about 45%, at least about 50%, at least about 55%, at least about 60%, at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 91%, at least about 92%, at least about 93%, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99% sequence identity to any one of SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 80% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 85% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 90% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 91% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 92% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 93% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 94% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 95% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 96% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 97% identity to SEQ ID NOs: 1477-1492.In some cases, the NLS comprises a sequence having at least about 98% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having at least about 99% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having 100% identity to SEQ ID NOs: 1477-1492. In some cases, the NLS comprises a sequence having 100% identity to SEQ ID NO: 1477. In some cases, the NLS comprises a sequence having 100% identity to SEQ ID NO: 1478. [Table 1-1] [Table 1-2]
[0265] In some embodiments, the retrotransposase comprises a tag. In some embodiments, the tag is an affinity tag. Exemplary affinity tags include, but are not limited to, His tag, Flag tag, Myc tag, MBP tag, and GST tag.
[0266] In some embodiments, the retrotransposase comprises a protease cleavage site. Exemplary protease cleavage sites include, but are not limited to, a TEV site, a C3 site, a factor Xa site, and an enterokinase site.
[0267] In some embodiments, the retrotransposase is tethered to a site-specific nuclease. In some embodiments, the retrotransposase is fused to a site-specific nuclease. In some embodiments, the retrotransposase is recruited to the site-specific nuclease. In some embodiments, the site-specific nuclease is a modifying endonuclease. In some embodiments, the site-specific nuclease is a Cas nuclease. In some embodiments, the Cas nuclease is an RNA-guided CRISPR Cas9 nuclease. In some embodiments, the site-specific nuclease is a dead nuclease or a nickase. In some embodiments, the site-specific nuclease brings the retrotransposase to a site adjacent to the target site to be modified.
[0268] Guide nucleic acid In some embodiments, the retrotransposase system further comprises a site-specific nuclease and a guide RNA (e.g., gRNA). When referring to T, in polynucleotides, T means U (uracil) in RNA and T (thymine) in DNA. In some embodiments, the retrotransposase system described in the present disclosure comprises a means for directing the site-specific nuclease to a specific position in the target nucleic acid.
[0269] In some embodiments, the guide RNA comprises synthetic or modified nucleotides. In some embodiments, the guide RNA comprises one or more internucleoside linkers modified from natural phosphodiester. In some embodiments, the internucleoside linkers of the guide RNA, or all of its contiguous nucleotide sequences, are modified. For example, in some embodiments, the internucleoside linkages comprise sulfur (S), such as phosphorothioate internucleoside linkages.
[0270] In some embodiments, the guide RNA comprises a modification to the ribose sugar or nucleobase. In some embodiments, the guide RNA comprises one or more nucleosides comprising a modified sugar moiety, where the modified sugar moiety is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acids (LNA)), or an unlinked ribose ring that typically lacks a bond between the C2 and C3 carbons (e.g., UNA). In some embodiments, the sugar-modified nucleoside comprises a bicyclohexose nucleic acid or a tricyclic nucleic acid. In some embodiments, the modified nucleoside comprises a nucleoside in which the sugar moiety is replaced with a non-sugar moiety, such as a peptide nucleic acid (PNA) or morpholino nucleic acid.
[0271] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, sugar modifications include modifications made by changing the substituent on the ribose ring to a group other than hydrogen or to the 2'-OH group naturally found in DNA and RNA nucleosides. In some embodiments, the substituent is introduced at the 2', 3', 4', or 5' position, or a combination thereof. In some embodiments, nucleosides having modified sugar moieties include 2'-modified nucleosides, e.g., 2'-substituted nucleosides. 2'-sugar-modified nucleosides, in some embodiments, are nucleosides having a substituent other than -H or -OH at the 2' position (2'-substituted nucleosides) or include a 2'-linked biradical, and include 2'-substituted nucleosides and LNA (2'-4' biradical bridged) nucleosides. Examples of 2'-substituted modified nucleosides include, but are not limited to, 2'-O-alkyl-RNA, 2'-O-methyl-RNA, 2'-alkoxy-RNA, 2'-O-methoxyethyl-RNA (MOE), 2'-amino-DNA, 2'-fluoro-RNA, and 2'-F-ANA nucleosides. In some embodiments, the modification in the ribose group comprises a modification at the 2' position of the ribose group. In some embodiments, the modification at the 2' position of the ribose group is selected from the group consisting of 2'-O-methyl, 2'-fluoro, 2'-deoxy, and 2'-O-(2-methoxyethyl).
[0272] In some embodiments, the guide RNA comprises one or more modified sugars. In some embodiments, the guide RNA comprises only modified sugars. In certain embodiments, the guide RNA comprises more than about 10%, 25%, 50%, 75%, or 90% modified sugars. In some embodiments, the modified sugar is a bicyclic sugar. In some embodiments, the modified sugar comprises a 2'-O-methoxyethyl group. In some embodiments, the guide RNA comprises both an internucleoside linker modification and a nucleoside modification.
[0273] In some cases, the guide RNA comprises a sequence complementary to a eukaryotic, fungal, plant, mammalian, or human genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a eukaryotic genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a fungal genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a plant genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a mammalian genomic polynucleotide sequence. In some cases, the guide RNA comprises a sequence complementary to a human genomic polynucleotide sequence.
[0274] In some cases, the guide RNA is 30-400 nucleotides in length. In some cases, the guide RNA is 85-245 nucleotides in length. In some cases, the guide RNA is more than 90 nucleotides in length. In some cases, the guide RNA is less than 245 nucleotides in length. In some embodiments, the guide RNA is 30, 40, 50, 60, 70, 80, 90, 100, 120, 140, 160, 180, 200, 220, 240, or more than 240 nucleotides in length. In some embodiments, the guide RNAs are about 30 to about 40, about 30 to about 50, about 30 to about 60, about 30 to about 70, about 30 to about 80, about 30 to about 90, about 30 to about 100, about 30 to about 120, about 30 to about 140, about 30 to about 160, about 30 to about 180, about 30 to about 200, about 30 to about 220, about 30 to about 240, about 50 to about 60, about 50 to about 70, about 50 to about 80, about 50 to about 90, about 50 to about 100, about 50 The length is about 120, about 50 to about 140, about 50 to about 160, about 50 to about 180, about 50 to about 200, about 50 to about 220, about 50 to about 240, about 100 to about 120, about 100 to about 140, about 100 to about 160, about 100 to about 180, about 100 to about 200, about 100 to about 220, about 100 to about 240, about 160 to about 180, about 160 to about 200, about 160 to about 220, or about 160 to about 240 nucleotides.
[0275] In some embodiments, the gRNA is encoded by any one of the nucleic acid sequences of SEQ ID NOs: 903-926 and 934-951, a sequence having at least about 80%, 85%, 90%, 95%, 97%, 98%, or 99% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 80% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 85% sequence identity to any one of the nucleic acid sequences of SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 90% sequence identity to any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 95% sequence identity to any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 97% sequence identity to any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 98% sequence identity to any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence having at least about 99% sequence identity to any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof. In some embodiments, the guide RNA is encoded by a sequence set forth in any one of the nucleic acid sequences in SEQ ID NOs: 903-926 and 934-951, or a reverse complement thereof.
[0276] In some embodiments, sequences may be determined by the BLASTP, CLUSTALW, MUSCLE, or MAFFT algorithms, or the CLUSTALW algorithm using Smith-Waterman homology search algorithm parameters. In some embodiments, sequence identity is determined by the BLASTP homology search algorithm using a BLOSUM62 scoring matrix setting parameters of word length (W) of 3, expectation (E) of 10, and gap costs at 11 presence and 1 extension, and using a conditional composition score matrix adjustment.
[0277] Cargo Nucleic Acid In some embodiments, the retrotransposase system comprises a loading nucleic acid or polynucleotide. In some embodiments, the loading nucleic acid is comprised in a double-stranded deoxyribonucleic acid. In some embodiments, the cargo nucleic acid is a eukaryotic, plant, fungal, mammalian, rodent, or human double-stranded deoxyribonucleic acid polynucleotide. In some embodiments, the cargo nucleotide sequence is flanked by a 3' untranslated region (UTR) and a 5' untranslated region (UTR).
[0278] In some embodiments, the cargo nucleic acid comprises synthetic or modified nucleotides. In some embodiments, the cargo nucleic acid comprises one or more internucleoside linkers modified from natural phosphodiester. In some embodiments, the internucleoside linkers of the cargo nucleic acid, or all of its contiguous nucleotide sequences, are modified. For example, in some embodiments, the internucleoside linkages comprise sulfur (S), such as phosphorothioate internucleoside linkages.
[0279] In some embodiments, the cargo nucleic acid comprises a modification to the ribose sugar or nucleobase. In some embodiments, the cargo nucleic acid comprises one or more nucleosides comprising a modified sugar moiety, which is a modification of the sugar moiety compared to the ribose sugar moiety found in deoxyribose nucleic acids (DNA) and RNA. In some embodiments, the modification is within the ribose ring structure. Exemplary modifications include, but are not limited to, replacement with a hexose ring (HNA), a bicyclic ring having a biradical bridge between the C2 and C4 carbons on the ribose ring (e.g., locked nucleic acids (LNA)), or an unlinked ribose ring that typical...
Claims
1. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
2. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase comprises an amino acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
3. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase comprises an amino acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
4. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase comprises an amino acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266.
5. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase is encoded by a nucleic acid having at least 75% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
6. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
7. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 90% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
8. 2. The engineered retrotransposase system of Claim 1, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 95% sequence identity to any one of SEQ ID NOs: 120-173, 181-187, 193-197, 203-207, 217-225, 231-235, 241-245, 251-255, 267-277, 288-297, 303-307, 324-339, 964-981, 1003-1019, 1504-1520, 1521-1536, 1539-1543, 1556-1568, and 1611-1806.
9. 9. The engineered retrotransposase system of any one of claims 1 to 8, wherein the double-stranded nucleic acid comprises a 5' recognition sequence comprising a GG nucleotide sequence and a 3' recognition sequence comprising a TGAC nucleotide sequence.
10. 10. The engineered retrotransposase system of claim 9, wherein the 5' recognition sequence and the 3' recognition sequence are configured to interact with the retrotransposase.
11. 11. The engineered retrotransposase system of any one of claims 1 to 10, wherein the double-stranded nucleic acid comprising a cargo nucleotide sequence is RNA.
12. 12. The engineered retrotransposase system of claim 11, wherein the RNA is in vitro transcribed RNA.
13. 13. The engineered retrotransposase system of any one of Claims 11-12, wherein the RNA comprises a sequence 5' to the cargo sequence or a sequence 3' to the cargo sequence that has at least 80% sequence identity to an RNA cognate of any one of SEQ ID NOs: 761-798, 2161-2164, and 2211-2257, its complement, or its reverse complement.
14. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-401, 799-894, 1476, 1850-1926, and 2165-2210.
15. 15. The engineered retrotransposase system of claim 14, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1535-1536, 1542-1543, 1611-1623, 1663-1691, and 1786-1806.
16. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO:402 or SEQ ID NO:
895.
17. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO:
388.
18. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:403-426.
19. 19. The engineered retrotransposase system of Claim 18, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 389-392 and 1504-1507.
20. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:427-439.
21. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 440-554 and 1020-1037.
22. 22. The engineered retrotransposase system of Claim 21, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 356-373, 964-981, and 1003-1019.
23. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 555-608 and 1927-2010.
24. 24. The engineered retrotransposase system of Claim 23, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 66-173, 740-756, 1521-1534, 1539-1541, 1624-1637, 1645-1662, and 1701-1782.
25. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 609-610 and 1555.
26. 26. The engineered retrotransposase system of Claim 25, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 308-309 and 324-325.
27. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:611-615 and 1544-1545.
28. 28. The engineered retrotransposase system of Claim 27, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 310-312, 326-328, 1556-1557, and 1569-1570.
29. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO:616 or SEQ ID NO:
617.
30. 30. The engineered retrotransposase system of Claim 29, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:313-314 and 329-330.
31. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 618-622 and 2258-2266.
32. 32. The engineered retrotransposase system of Claim 31, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 315-319 and 331-335.
33. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO:
623.
34. 34. The engineered retrotransposase system of Claim 33, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 320 or SEQ ID NO:
336.
35. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:624-626.
36. 36. The engineered retrotransposase system of Claim 35, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 321-323, 337-339, and 1785.
37. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:624-626.
38. 36. The engineered retrotransposase system of Claim 35, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 321-323, 337-339, and 1785.
39. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 627-673, 1039-1475, and 2011-2026.
40. 40. The engineered retrotransposase system of Claim 39, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:174-187 and 1508-1520.
41. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:674-678.
42. 42. The engineered retrotransposase system of Claim 41, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 188-197.
43. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:679-683.
44. 44. The engineered retrotransposase system of Claim 43, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 198-207.
45. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 684-692 and 2027-2046.
46. 46. The engineered retrotransposase system of Claim 45, wherein the retrotransposase further comprises a sequence having at least 80% sequence identity to any one of SEQ ID NOs: 208-225 and 757-759, or a variant thereof.
47. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:693-697 and 2047-2090.
48. 48. The engineered retrotransposase system of Claim 47, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:226-235.
49. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 698-702 and 2091-2119.
50. 50. The engineered retrotransposase system of Claim 49, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:236-245 and 759-760.
51. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:703-707.
52. 52. The engineered retrotransposase system of Claim 51, wherein said retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:246-255.
53. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:708-718 and 2121-2159.
54. 54. The engineered retrotransposase system of Claim 53, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:256-277, 1638-1644, and 1693-1700.
55. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:719-728.
56. 56. The engineered retrotransposase system of Claim 55, wherein said retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:278-297.
57. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs:729-733.
58. 58. The engineered retrotransposase system of Claim 57, wherein said retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs:298-307.
59. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 734-735 and 1546-1553.
60. 60. The engineered retrotransposase system of Claim 59, wherein the retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to any one of SEQ ID NOs: 1558-1567, 1571-1580, and 1783-1784.
61. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 1038 or SEQ ID NO: 2160.
62. 62. The engineered retrotransposase system of Claim 61, wherein said retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 1692.
63. 1. An engineered retrotransposase system comprising: (a) a double-stranded nucleic acid comprising a cargo nucleotide sequence configured to form a complex with a retrotransposase; (b) a retrotransposase configured to transpose the cargo nucleotide sequence to a target nucleic acid sequence, wherein the retrotransposase comprises an amino acid sequence having at least 75% sequence identity to SEQ ID NO: 1554.
64. 64. The engineered retrotransposase system of Claim 63, wherein said retrotransposase is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 1568 or SEQ ID NO: 1594.
65. 65. The engineered retrotransposase system of any one of claims 1-64, wherein the retrotransposase comprises one or more nuclear localization sequences (NLS) proximal to the N-terminus or C-terminus of the retrotransposase.
66. 66. The engineered retrotransposase system of Claim 65, wherein the NLS comprises a sequence that is at least 80% identical to a sequence from the group consisting of SEQ ID NOs: 1477-1492.
67. 66. The engineered retrotransposase system of Claim 65, wherein said NLS comprises SEQ ID NO: 1478.
68. 66. The engineered retrotransposase system of Claim 65, wherein the NLS is proximal to the N-terminus of the retrotransposase.
69. 66. The engineered retrotransposase system of Claim 65, wherein said NLS comprises SEQ ID NO: 1477.
70. 66. The engineered retrotransposase system of Claim 65, wherein the NLS is proximal to the C-terminus of the retrotransposase.
71. A polypeptide comprising a reverse transcriptase comprising an amino acid sequence having at least 75% sequence identity to any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210 and 2258-2266 fused at the N-terminus or C-terminus to a non-retrotransposase domain or affinity tag.
72. 72. The polypeptide of Claim 71, wherein the non-retrotransposase domain is an RNA-binding protein domain.
73. 73. The polypeptide of claim 72, wherein the RNA-binding protein domain comprises a bacteriophage MS2 coat protein (MCP) domain.
74. 74. A nucleic acid encoding the engineered retrotransposase system of any one of claims 1 to 64 or the polypeptide of any one of claims 71 to 73.
75. 65. A method for modifying a target nucleic acid sequence, the method comprising contacting said target nucleic acid sequence with the engineered nuclease system of any one of claims 1 to 64.
76. 76. The method of claim 75, wherein modifying the target nucleic acid sequence comprises binding, nicking, or cleaving the target nucleic acid sequence.
77. 77. The method of any one of claims 75-76, wherein the target nucleic acid sequence comprises genomic DNA, viral DNA, viral RNA, or bacterial DNA.
78. 77. The method of any one of claims 75-76, wherein the target nucleic acid sequence comprises deoxyribonucleic acid (DNA).
79. 79. The method of any one of claims 75 to 78, wherein the modification is in vitro.
80. 79. The method of any one of claims 75 to 78, wherein the modification is in vivo.
81. 79. The method of any one of claims 75 to 78, wherein the modification is ex vivo.
82. 65. A method of modifying a target nucleic acid sequence in a mammalian cell, the method comprising contacting said mammalian cell with the engineered nuclease system of any one of claims 1 to 64.
83. 1. A method for synthesizing complementary DNA (cDNA), comprising: (a) providing an RNA molecule as a template for cDNA synthesis; (b) providing a primer oligonucleotide to prime cDNA synthesis from said RNA molecule; (c) synthesizing a cDNA primed by the primer oligonucleotide from the template using a reverse transcriptase comprising a sequence having at least 80% sequence identity to a reverse transcriptase domain of any one of SEQ ID NOs: 1-29, 393-735, 799-895, 1020-1476, 1544-1554, 1850-2160, 2165-2210, and 2258-2266, or a variant thereof.
84. 84. The method of claim 83, wherein the primer oligonucleotide comprises an oligo(dT) sequence or a degenerate sequence of at least six oligonucleotides.
85. A vector comprising the nucleic acid of claim 74.
86. 86. The vector of claim 85, wherein the vector is a plasmid, a minicircle, a CELiD, an adeno-associated virus (AAV)-derived virion, or a lentivirus.
87. 74. A cell comprising the engineered nuclease system of any one of claims 1 to 64, or the polypeptide of any one of claims 71 to 73.
88. 88. The cell of claim 87, wherein the cell is a eukaryotic cell.
89. 88. The cell of claim 87, wherein the cell is a mammalian cell.
90. 88. The cell of claim 87, wherein the cell is an immortalized cell.
91. 88. The cell of claim 87, wherein the cell is an insect cell.
92. 88. The cell of claim 87, wherein the cell is a yeast cell.
93. 88. The cell of claim 87, wherein the cell is a plant cell.
94. 88. The cell of claim 87, wherein the cell is a fungal cell.
95. 88. The cell of claim 87, wherein the cell is a prokaryotic cell.
96. 88. The cell of claim 87, wherein the cell is A549, HEK-293, HEK-293T, BHK, CHO, HeLa, MRC5, Sf9, Cos-1, Cos-7, Vero, BSC 1, BSC 40, BMT 10, WI38, HeLa, Saos, C2C12, L cells, HT1080, HepG2, Huh7, K562, primary cells, or derivatives thereof.
97. 88. The cell of claim 87, wherein the cell is an engineered cell.
98. 88. The cell of claim 87, wherein the cell is a stable cell.