Modified UTRS for retrotransposon systems

Modified RTE-1 retrotransposase with tailored UTRs addresses the inefficiencies of existing methods by enabling precise and efficient integration of nucleic acid sequences into genomes, enhancing site-specificity and integration frequency.

WO2025255548A1PCT designated stage Publication Date: 2025-12-11TESSERA THERAPEUTICS INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/032776
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-07
Filing Date
2025-06-06
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing methods for integrating nucleic acid sequences into a genome lack site specificity and efficiency, particularly for longer sequences, and require multiple steps or specialized proteins like CRISPR/Cas9 or Cre/loxP.

Method used

Utilizing modified RTE-1 retrotransposase with tailored 5' and 3' untranslated regions (UTRs) to direct precise insertion of sequences into a genome, enhancing site-specific integration of heterologous objects.

Benefits of technology

The modified RTE-1 retrotransposase system enables efficient and targeted integration of nucleic acid sequences into genomes, improving insertion frequency and specificity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025032776_11122025_PF_FP_ABST
    Figure US2025032776_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and compositions for modulating a target genome are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 2017469-0043 MODIFIED UTRS FOR RETROTRANSPOSON SYSTEMS CROSS REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims the benefit of U.S. Provisional Appln. No.63 / 657,255 filed June 7, 2024, the entire contents of which are hereby incorporated by reference herein in their entirety. BACKGROUND

[0002] Integration of a nucleic acid of interest into a genome occurs at low frequency and with little site specificity, in the absence of a specialized protein to promote the insertion event. Some existing approaches, like CRISPR / Cas9, are more suited for small edits and are less effective at integrating longer sequences. Other existing approaches, like Cre / loxP, require a first step of inserting a loxP site into the genome and then a second step of inserting a sequence of interest (e.g., a heterologous object sequence) into the loxP site. There is a need in the art for improved proteins for inserting sequences of interest into a genome. SUMMARY

[0003] An RTE-1 retrotransposase is capable of directing insertion of a desired sequence into a genome, when the desired sequence is provided on a template RNA comprising an RTE-15’ untranslated region (UTR) and an RTE-13’ UTR. The present disclosure provides, in part, modified UTRs for a template RNA compatible with an RTE-1 retrotransposase.

[0004] In one aspect, provided herein are nucleic acid molecules comprising a modified RTE-1 3’ UTR that comprises a nucleotide difference (e.g., substitution, insertion, or deletion) at one or more of: Region A, Region B, Region D, or Region E, as numbered according to SEQ ID NO: 2.

[0005] In some embodiments, the modified RTE-13’ UTR comprises 5, 6, 7, or 8 of, or all of: Region A, Region B, Region C, Region D, Region E, Region F, Region G, or Region H.

[0006] In another aspect, provided herein are nucleic acid molecules comprising a modified RTE-13’ UTR that comprises a nucleotide difference (e.g., substitution or deletion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all of positions A4, C6, A21, U29, C30, G31, A32, C47, C48, A49, C50, or U51, as numbered according to SEQ ID NO: 2. Page 1 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0007] In some embodiments, the nucleic acid molecule has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity to SEQ ID NO: 2. In some embodiments, the nucleic acid molecule has 1, 2, 3, or all of substitutions: A4U, C6A, A21U, or A32G, as numbered according to SEQ ID NO: 2.

[0008] In some embodiments, the nucleic acid molecule has a deletion of position 29, 30, 31, or 32 as numbered according to SEQ ID NO: 2.

[0009] In some embodiments, the modified RTE-13’ UTR has a length of 25-30, 30-35, 35-40, 40-45, 45-50, or 50-51 nucleotides.

[0010] In some embodiments, the nucleic acid molecule has an A21U substitution as numbered according to SEQ ID NO: 2.

[0011] In some embodiments, the nucleic acid molecule has a C6A substitution as numbered according to SEQ ID NO: 2.

[0012] In some embodiments, the nucleic acid molecule has an A32G substitution as numbered according to SEQ ID NO: 2.

[0013] In some embodiments, the nucleic acid molecule comprises a deletion of position 31, as numbered according to SEQ ID NO: 2.

[0014] In some embodiments, the nucleic acid molecule comprises a deletion of position 32, as numbered according to SEQ ID NO: 2.

[0015] In some embodiments, the nucleic acid molecule has an A32G substitution and a deletion of position 30 as numbered according to SEQ ID NO: 2.

[0016] In some embodiments, the nucleic acid molecule has an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0017] In some embodiments, the nucleic acid molecule has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0018] In some embodiments, the nucleic acid molecule has a C6A substitution, an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0019] In some embodiments, the nucleic acid molecule has a C6A substitution and an A32G substitution as numbered according to SEQ ID NO: 2. Page 2 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0020] In some embodiments, the nucleic acid molecule has an A21U substitution and an A32G substitution as numbered according to SEQ ID NO: 2.

[0021] In some embodiments, the nucleic acid molecule has a C6A substitution and an A21U substitution as numbered according to SEQ ID NO: 2.

[0022] In some embodiments, the nucleic acid molecule comprises the sequence of any one of SEQ ID NOs: 11-14 and 506-512.

[0023] In another aspect, provided herein are nucleic acid molecules comprising a modified RTE-15’ UTR that comprises a nucleotide difference (e.g., substitution, deletion, or insertion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO: 1.

[0024] In some embodiments, the nucleic acid molecule has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 98% identity to SEQ ID NO: 1.

[0025] In some embodiments, the nucleic acid molecule comprises a nucleotide substitution at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO: 1.

[0026] In some embodiments, the nucleic acid molecule has a modified RTE-15’ UTR has a length of 400-450, 450-500, 500-550, 550-600, 600-650, 650-700, or 700-710 nucleotides, e.g., 700-710 nucleotides.

[0027] In some embodiments, the nucleic acid molecule has an U138C substitution as numbered according to SEQ ID NO: 1.

[0028] In some embodiments, the nucleic acid molecule has an U180C substitution as numbered according to SEQ ID NO: 1.

[0029] In some embodiments, the nucleic acid molecule has an A208G substitution as numbered according to SEQ ID NO: 1.

[0030] In some embodiments, the nucleic acid molecule has an U264G substitution as numbered according to SEQ ID NO: 1. Page 3 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0031] In some embodiments, the nucleic acid molecule has an A271G substitution as numbered according to SEQ ID NO: 1.

[0032] In some embodiments, the nucleic acid molecule has a G307U substitution as numbered according to SEQ ID NO: 1.

[0033] In some embodiments, the nucleic acid molecule has a G358A substitution as numbered according to SEQ ID NO: 1.

[0034] In some embodiments, the nucleic acid molecule has an A398G substitution as numbered according to SEQ ID NO: 1.

[0035] In some embodiments, the nucleic acid molecule has an U399C substitution as numbered according to SEQ ID NO: 1.

[0036] In some embodiments, the nucleic acid molecule has an A405G substitution as numbered according to SEQ ID NO: 1.

[0037] In some embodiments, the nucleic acid molecule has an G469A substitution as numbered according to SEQ ID NO: 1.

[0038] In some embodiments, the nucleic acid molecule has an U488C substitution as numbered according to SEQ ID NO: 1.

[0039] In some embodiments, the nucleic acid molecule has an insertion of an A following U658 as numbered according to SEQ ID NO: 1.

[0040] In some embodiments, the nucleic acid molecule has a U180C substitution and a U399C substitution as numbered according to SEQ ID NO: 1.

[0041] In some embodiments, the nucleic acid molecule has a U180C substitution, an A398G substitution, and an A405G substitution as numbered according to SEQ ID NO: 1.

[0042] In some embodiments, the nucleic acid molecule has a U138C substitution, a U180C substitution, an A405G substitution, and a U488C substitution as numbered according to SEQ ID NO: 1.

[0043] In some embodiments, the nucleic acid molecule has a U180C substitution, an A398G substitution, an A405G substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1. Page 4 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0044] In some embodiments, the nucleic acid molecule has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1.

[0045] In some embodiments, the nucleic acid molecule has a U138C substitution, a U180C substitution, an A405G substitution, a U488C substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1.

[0046] In some embodiments, the nucleic acid molecule has a U138C substitution, a U180C substitution, a U264G substitution, a U399C substitution, an A405G substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1.

[0047] In some embodiments, the nucleic acid molecule has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G358A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1.

[0048] In some embodiments, the nucleic acid molecule comprises the sequence of any one of SEQ ID NOs: 3-10, and 501-505.

[0049] In some embodiments, the nucleic acid molecule binds an RTE-1 polypeptide of Table 1 or Table 22.

[0050] In some embodiments, the nucleic acid molecule comprises one or more chemically modified nucleotides.

[0051] In another aspect, provided herein are template RNA comprising, from 5’ to 3’: a) optionally, an RTE-15’ UTR; b) a heterologous object sequence; and c) a nucleic acid molecule comprising a modified RTE-13’ UTR described herein.

[0052] In some embodiments, the template RNA comprises an RTE-15’ UTR, and wherein the RTE-15’ UTR comprises a nucleic acid sequence according to SEQ ID NO: 1.

[0053] In some embodiments, the template RNA comprises an RTE-15’ UTR, and wherein the RTE-15’ UTR comprises a modified RTE-15’ UTR. Page 5 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0054] In some embodiments, the template RNA comprising, from 5’ to 3’: a) a nucleic acid molecule comprising a modified RTE-15’ UTR described herein; b) a heterologous object sequence; and c) optionally, an RTE-13’ UTR.

[0055] In some embodiments, the template RNA comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a nucleic acid sequence according to SEQ ID NO: 2.

[0056] In some embodiments, the template RNA comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a modified RTE-13’ UTR.

[0057] In some embodiments, the template RNA comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a modified RTE-13’ UTR described herein.

[0058] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence having an A21U substitution as numbered according to SEQ ID NO: 2.

[0059] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence having a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0060] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0061] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0062] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a Page 6 of 372 12815032v1Attorney Docket No.: 2017469-0043 heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution, an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0063] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and an A32G substitution as numbered according to SEQ ID NO: 2.

[0064] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution and an A32G substitution as numbered according to SEQ ID NO: 2.

[0065] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and an A21U substitution as numbered according to SEQ ID NO: 2.

[0066] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0067] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO: 2.

[0068] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as Page 7 of 372 12815032v1Attorney Docket No.: 2017469-0043 numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0069] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0070] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution and a U399C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0071] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution, an A398G substitution, and an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0072] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, an A405G substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0073] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution, an A398G substitution, an A405G substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2. Page 8 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0074] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0075] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0076] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0077] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO: 2.

[0078] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO: 2. Page 9 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0079] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, an A405G substitution, a U488C substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0080] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, a U264G substitution, a U399C substitution, an A405G substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0081] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 2.

[0082] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO: 2.

[0083] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified Page 10 of 372 12815032v1Attorney Docket No.: 2017469-0043 RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0084] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0085] In another aspect, the disclosure features a template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic sequence of SEQ ID NO:1 or a modified RTE-15’ UTR comprising the nucleic acid sequence of any one of SEQ ID NOs: 3-10 and 501-505; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2 or a modified RTE-13’ UTR comprising the nucleic acid sequence of any one of SEQ ID NOs: 11-14 and 506-512.

[0086] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:6; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0087] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:10; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0088] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:501; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0089] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:502; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2. Page 11 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0090] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:8; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0091] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:9; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0092] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:503; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0093] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:504; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0094] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:505; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:2.

[0095] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:13.

[0096] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:506.

[0097] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:507.

[0098] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:508. Page 12 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0099] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:509.

[0100] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:510.

[0101] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:511.

[0102] In some embodiments, the template comprises, from 5’ to 3’: a) an RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:512.

[0103] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:6; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:13.

[0104] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:6; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:506.

[0105] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:6; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:507.

[0106] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:10; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:13.

[0107] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:10; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:506. Page 13 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0108] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:10; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:507.

[0109] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:501; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:13.

[0110] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:501; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:506.

[0111] In some embodiments, the template comprises, from 5’ to 3’: a) a modified RTE-15’ UTR comprising the nucleic acid sequence of SEQ ID NO:501; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising the nucleic acid sequence of SEQ ID NO:507.

[0112] In another aspect, a template RNA described herein promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; and b) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

[0113] In another aspect, provided herein are gene modifying systems comprising: a template RNA of any aspect or embodiment described herein; and a gene modifying polypeptide, or a nucleic acid encoding the gene modifying polypeptide, the gene modifying polypeptide comprising an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392 , or an amino acid sequence having at least 80%, 90%, 95%, 97%, 98%, 99% identity to SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392. Page 14 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0114] In another aspect, provider herein are pharmaceutical compositions comprising a nucleic acid of any aspect or embodiment described herein or a gene modifying system of any aspect or embodiment described herein, and a pharmaceutically acceptable excipient or carrier.

[0115] In some embodiments, a pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle (LNP).

[0116] In some embodiments, a host cell (e.g., a mammalian cell, e.g., a human cell) comprises a gene modifying system, template RNA, or nucleic acid of any aspect or embodiment described herein.

[0117] In another aspect, provider herein are methods of making a cell of any aspect or embodiment described herein comprising providing the cell and contacting the cell with a gene modifying system of any aspect or embodiment described herein.

[0118] In another aspect, provider herein are methods of making a nucleic acid or template RNA of any aspect or embodiment comprising synthesizing the template RNA in vitro (e.g., by in vitro transcription or solid state synthesis).

[0119] In another aspect, provider herein are methods of modifying a genome of a mammalian cell, comprising contacting the cell with a gene modifying system of any aspect or embodiment described herein, thereby modifying the genome of the mammalian cell.

[0120] In another aspect, provider herein are cells made by a method of any aspect or embodiment described herein.

[0121] In another aspect, provider herein are methods for treating a subject having cancer comprising administering to the subject a gene modifying system any aspect or embodiment described herein or DNA encoding the same, or a pharmaceutical composition any aspect or embodiment described herein, thereby treating the subject having the cancer.

[0122] Features of the compositions or methods can include one or more of the following enumerated embodiments. Page 15 of 372 12815032v1Attorney Docket No.: 2017469-0043 Enumerated Embodiments

[0123] 1. A nucleic acid molecule comprising a nucleotide sequence having at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, but no more than 10, 15, or 20 nucleotide differences (e.g., insertions, substitutions, or deletions) relative to SEQ ID NO: 2.

[0124] 2. A nucleic acid molecule comprising a modified RTE-13’ UTR that comprises a nucleotide difference (e.g., substitution, insertion, or deletion) at one or more of: Region A, Region B, Region D, or Region E, as numbered according to SEQ ID NO: 2.

[0125] 3. The nucleic acid molecule of embodiment 2, wherein the modified RTE-13’ UTR comprises 5, 6, 7, or 8 of, or all of: Region A, Region B, Region C, Region D, Region E, Region F, Region G, or Region H.

[0126] 4. The nucleic acid molecule of any preceding embodiments, wherein Region A is located at nucleotides 1-20 of a sequence numbered according to SEQ ID NO: 2.

[0127] 5. The nucleic acid molecule of any preceding embodiments, wherein Region B is located at nucleotides 21-24 of a sequence numbered according to SEQ ID NO: 2.

[0128] 6. The nucleic acid molecule of any preceding embodiments, wherein Region C is located at nucleotides 25-26 of a sequence numbered according to SEQ ID NO: 2.

[0129] 7. The nucleic acid molecule of any preceding embodiments, wherein Region D is located at nucleotides 27-31 of a sequence numbered according to SEQ ID NO: 2.

[0130] 8. The nucleic acid molecule of any preceding embodiments, wherein Region E is located at nucleotides 32-36 of a sequence numbered according to SEQ ID NO: 2.

[0131] 9. The nucleic acid molecule of any preceding embodiments, wherein Region F is located at nucleotides 37-41 of a sequence numbered according to SEQ ID NO: 2.

[0132] 10. The nucleic acid molecule of any preceding embodiments, wherein Region G is located at nucleotides 42-45 of a sequence numbered according to SEQ ID NO: 2.

[0133] 11. The nucleic acid molecule of any preceding embodiments, wherein Region H is located at nucleotides 46-51 of a sequence numbered according to SEQ ID NO: 2. Page 16 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0134] 12. The nucleic acid molecule of any of the preceding embodiments, wherein Region A has a sequence according to nucleotides 1-20 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0135] 13. The nucleic acid molecule of any of the preceding embodiments, wherein Region B has a sequence according to nucleotides 21-24 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0136] 14. The nucleic acid molecule of any of the preceding embodiments, wherein Region C has a sequence according to nucleotides 25-26 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0137] 15. The nucleic acid molecule of any of the preceding embodiments, wherein Region D has a sequence according to nucleotides 27-31 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0138] 16. The nucleic acid molecule of any of the preceding embodiments, wherein Region E has a sequence according to nucleotides 32-36 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0139] 17. The nucleic acid molecule of any of the preceding embodiments, wherein Region F has a sequence according to nucleotides 37-41 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0140] 18. The nucleic acid molecule of any of the preceding embodiments, wherein Region G has a sequence according to nucleotides 42-45 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0141] 19. The nucleic acid molecule of any of the preceding embodiments, wherein Region H has a sequence according to nucleotides 46-51 of SEQ ID NO: 2, or a sequence with no more than 1, 2, or 3 nucleotide differences relative thereto.

[0142] 20. The nucleic acid molecule of any of the preceding embodiments, wherein Region B is part of a double-stranded region, e.g., Region B interacts with (e.g., pairs with no mismatches) Region G. Page 17 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0143] 21. The nucleic acid molecule of any of the preceding embodiments, wherein Region D is part of a double-stranded region, e.g., Region D interacts with (e.g., pairs with no mismatches) Region F.

[0144] 22. The nucleic acid molecule of any of the preceding embodiments, wherein Region G is part of a double-stranded region, e.g., Region G interacts with (e.g., pairs with no mismatches) Region B.

[0145] 23. The nucleic acid molecule of any of the preceding embodiments, wherein Region F is part of a double-stranded region, e.g., Region F interacts with (e.g., pairs with no mismatches) Region D.

[0146] 24. The nucleic acid molecule of any of the preceding embodiments, wherein regions D, E, and F form a stem-loop.

[0147] 25. The nucleic acid molecule of any of the preceding embodiments, wherein region E is a loop.

[0148] 26. The nucleic acid molecule of any of the preceding embodiments, wherein region C is a bulge.

[0149] 27. The nucleic acid molecule of any of the preceding embodiments, wherein region A is single-stranded.

[0150] 28. The nucleic acid molecule of any of the preceding embodiments, wherein region H is single-stranded.

[0151] 29. The nucleic acid molecule of any of the preceding embodiments, wherein the mutation in Region A, Region B, or Region E is a substitution.

[0152] 30. The nucleic acid molecule of any of the preceding embodiments, wherein the mutation in Region A is a substitution.

[0153] 31. The nucleic acid molecule of any of the preceding embodiments, wherein the mutation in Region B is a substitution.

[0154] 32. The nucleic acid molecule of any of the preceding embodiments, wherein the mutation in Region E is a substitution. Page 18 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0155] 33. The nucleic acid molecule of any of the preceding embodiments, wherein the mutation in Region D is a deletion.

[0156] 34. The nucleic acid molecule of any of the preceding embodiments, wherein the sequence difference comprises a deletion in a double stranded region.

[0157] 35. The nucleic acid molecule of any of the preceding embodiments, wherein the sequence difference comprises a substitution in a single stranded region (e.g., a loop).

[0158] 36. The nucleic acid molecule of any of embodiments 1-34, wherein the sequence difference comprises a substitution in a double stranded region.

[0159] 37. A nucleic acid molecule comprising a modified RTE-13’ UTR that comprises a nucleotide difference (e.g., substitution or deletion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all of positions A4, C6, A21, U29, C30, G31, A32, C47, C48, A49, C50, or U51, as numbered according to SEQ ID NO: 2.

[0160] 38. The nucleic acid molecule of any of the preceding embodiments, which has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity to SEQ ID NO: 2.

[0161] 39. The nucleic acid molecule of any of the preceding embodiments, which has 1, 2, 3, or all of substitutions: A4U, C6A, A21U, or A32G, as numbered according to SEQ ID NO: 2.

[0162] 40. The nucleic acid molecule of any of the preceding embodiments, which has a deletion of position 29 as numbered according to SEQ ID NO: 2.

[0163] 41. The nucleic acid molecule of any of the preceding embodiments, which has a deletion of position 30 as numbered according to SEQ ID NO: 2.

[0164] 42. The nucleic acid molecule of any of the preceding embodiments, which has a deletion of position 31 as numbered according to SEQ ID NO: 2.

[0165] 43. The nucleic acid molecule of any of the preceding embodiments, which has a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0166] 44. The nucleic acid molecule of any of the preceding embodiments, which has a deletion of one or more of positions 29-31 (e.g., at 1, 2, or all 3 of positions 29, 30, and / or 31) as numbered according to SEQ ID NO: 2. Page 19 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0167] 45. The nucleic acid molecule of any of the preceding embodiments, which has substitution at position A32 (e.g., an A32G substitution) and a deletion of position 30 as numbered according to SEQ ID NO: 2.

[0168] 46. The nucleic acid molecule of any of the preceding embodiments, which has a substitution at position A21 (e.g., an A21U substitution) and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0169] 47. The nucleic acid molecule of any of the preceding embodiments, which has a substitution at position C6 (e.g., a C6A substitution) and a deletion of position 32 as numbered according to SEQ ID NO: 2.

[0170] 48. The nucleic acid molecule of any of the preceding embodiments, which has substitution at position A21 (e.g., an A21U substitution) and a deletion of position 31 as numbered according to SEQ ID NO: 2.

[0171] 49. The nucleic acid molecule of any of the preceding embodiments, which does not include a deletion within Region H of SEQ ID NO: 2.

[0172] 50. The nucleic acid molecule of any of the preceding embodiments, which comprises the nucleotide sequence of Region H of SEQ ID NO: 2, or a contiguous sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity thereto.

[0173] 51. The nucleic acid molecule of any of the preceding embodiments, which comprises nucleotides 47-51 of SEQ ID NO: 2.

[0174] 52. The nucleic acid molecule of any of the preceding embodiments, which comprises a sequence of Table 19 or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0175] 53. The nucleic acid molecule of any of the preceding embodiments, which comprises one or more mutations listed in the column “Mutations” of Table 21.

[0176] 54. The nucleic acid molecule of any preceding embodiments, which comprises an A4U substitution relative to SEQ ID NO: 2.

[0177] 55. The nucleic acid molecule of any preceding embodiments, which comprises a C6A substitution relative to SEQ ID NO: 2. Page 20 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0178] 56. The nucleic acid molecule of any preceding embodiments, which comprises an A21U substitution relative to SEQ ID NO: 2.

[0179] 57. The nucleic acid molecule of any preceding embodiments, which comprises an A32G substitution relative to SEQ ID NO: 2.

[0180] 58. The nucleic acid of any of the preceding embodiments, which comprises a single stranded region at its 3’ end.

[0181] 59. The nucleic acid of embodiment 58, wherein the single stranded region has a length of at least 3, 4, 5, or 6 nucleotides.

[0182] 60. The nucleic acid of any of the preceding embodiments, wherein the modified RTE-13’ UTR has a length of 25-30, 30-35, 35-40, 40-45, 45-50, or 50-51 nucleotides.

[0183] 61. The nucleic acid molecule of any of the preceding embodiments, which binds an RTE-1 polypeptide of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392.

[0184] 62. The nucleic acid molecule of any of the preceding embodiments, which promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; b) providing a template RNA comprising, from 5’ to 3’: the nucleic acid molecule of any of embodiments 1-61, a heterologous object sequence, and an RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1; c) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

[0185] 63. The nucleic acid molecule of embodiment 62, which promotes a greater number of integration events compared to a control template RNA having a sequence comprising, from 5’ to 3’: an RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and an RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0186] 64. The nucleic acid molecule of embodiment 62, which promotes a greater expression of a protein (e.g., GFP or CAR) encoded by a heterologous object sequence, Page 21 of 372 12815032v1Attorney Docket No.: 2017469-0043 compared to a control template RNA having a sequence comprising, from 5’ to 3’: a RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and a RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0187] 65. The nucleic acid molecule of embodiment 64, wherein the expression of the protein encoded by the heterologous object sequence is at least 1.05, 1.06, 1.07, 1.08, 1.09, 1.1, 1.11, 1.12, 1.13, 1.14, 1.15, 1.16, 1.17, 1.18, 1.19, 1.2, 1.21, 1.22, 1.23, 1.24, 1.25, 1.26, 1.27, 1.28, 1.29, 1.3, 1.31, 1.32, 1.33, 1.34, 1.35, 1.36, 1.37, 1.38, 1.39, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, 2.1, 2.2, 2.3, 2.4, or 2.5 times expression of the protein in cells treated with the control template RNA.

[0188] 66. A nucleic acid molecule comprising a nucleotide sequence having at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, but no more than 10, 15, or 20 nucleotide differences (e.g., insertion, substitutions, or deletions) relative to SEQ ID NO: 1.

[0189] 67. A nucleic acid molecule comprising a modified RTE-15’ UTR that comprises a nucleotide difference (e.g., substitution or deletion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO: 1.

[0190] 68. The nucleic acid molecule of any of embodiments 66 or 67, which has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 98% identity to SEQ ID NO: 1.

[0191] 69. The nucleic acid molecule of any of embodiments 66-68, which comprises a nucleotide substitution at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO: 1.

[0192] 70. The nucleic acid molecule of any of embodiments 66-69, which comprises a nucleotide substitution at 1, 2, 3, 4, 5, 6, 7, 8, or all of positions U138, A208, U264, A271, G307, G358, U399, G469, or U488 as numbered according to SEQ ID NO: 1.

[0193] 71. The nucleic acid molecule of any of embodiments 66-70, which comprises a nucleotide substitution at 1, 2, 3, or all of positions U138, U180, A405, or U488 as numbered according to SEQ ID NO: 1. Page 22 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0194] 72. The nucleic acid molecule of any of embodiments 66-71, which comprises a nucleotide substitution at 1, 2, or all of positions U180, A398, or A405 as numbered according to SEQ ID NO: 1.

[0195] 73. The nucleic acid molecule of any of embodiments 66-72, which comprises a nucleotide substitution at 1, 2, 3, 4, or all of positions: U180, A208, A398, A405, or U407 as numbered according to SEQ ID NO: 1.

[0196] 74. The nucleic acid molecule of any of embodiments 66-73, which comprises a nucleotide substitution 1, 2, 3, 4, or all of positions: U180, A208, A398, A405, U407, or U658 as numbered according to SEQ ID NO: 1.

[0197] 75. The nucleic acid molecule of any of embodiments 66-74, which comprises a nucleotide substitution 1, 2, 3, 4, 5, or all of positions: U180, A208, A398, A405, U407, or G570 as numbered according to SEQ ID NO: 1.

[0198] 76. The nucleic acid molecule of any of embodiments 66-75, which comprises a nucleotide substitution 1, 2, 3, 4, 5, 6, or all of positions: U180, A208, A398, A405, U407, G570, or U658 as numbered according to SEQ ID NO: 1.

[0199] 77. The nucleic acid molecule of any of embodiments 66-76, which comprises a nucleotide substitution 1, 2, 3, 4, or all of positions: U180, A208, U399, or U658 as numbered according to SEQ ID NO: 1.

[0200] 78. The nucleic acid molecule of any of embodiments 66-77, which comprises 1, 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or all of the following nucleotide substitutions: U138C, U180C, A208G, U264G, A271G, G307U, G358A, A398G, U399C, A405G, U407C, G469A, U488C, G570A, or U658A, as numbered according to SEQ ID NO: 1.

[0201] 79. The nucleic acid molecule of any of embodiments 66-78, which has 1, 2, 3, 4, 5, 6, 7, 8, or 9 or all of substitutions: U138C, A208G, U264G, A271G, G307U, G358A, U399C, A405G, G469A, or U488C, as numbered according to SEQ ID NO: 1.

[0202] 80. The nucleic acid molecule of any of embodiments 66-79, which has 1, 2, 3, 4, 5, 6, 7, 8, or all of substitutions: U138C, A208G, U264G, A271G, G307U, G358A, U399C, G469A, or U488C, as numbered according to SEQ ID NO: 1. Page 23 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0203] 81. The nucleic acid molecule of any of embodiments 66-80, which has 1, 2, 3, or all of substitutions: U138C, U180C, A405G, or U488C as numbered according to SEQ ID NO: 1.

[0204] 82. The nucleic acid molecule of any of embodiments 66-81, which has 1, 2, or all of substitutions: U180C, A398G, or A405G as numbered according to SEQ ID NO: 1.

[0205] 83. The nucleic acid molecule of any of embodiments 66-82, which has 1, 2, 3, 4, or all of substitutions: U180C, A208G, A398G, A405G, or U407C as numbered according to SEQ ID NO: 1.

[0206] 84. The nucleic acid molecule of any of embodiments 66-83, which has 1, 2, 3, 4, or all of substitutions: U180C, A208G, A398G, A405G, U407C, or U658A as numbered according to SEQ ID NO: 1.

[0207] 85. The nucleic acid molecule of any of embodiments 66-84, which has 1, 2, 3, 4, 5, or all of substitutions: U180C, A208G, A398G, A405G, U407C, or G570A as numbered according to SEQ ID NO: 1.

[0208] 86. The nucleic acid molecule of any of embodiments 66-85, which has 1, 2, 3, 4, 5, 6, or all of substitutions: U180C, A208G, A398G, A405G, U407C, G570A, or U658A as numbered according to SEQ ID NO: 1.

[0209] 87. The nucleic acid molecule of any of embodiments 66-86, which has 1, 2, 3, 4, or all of substitutions: U180C, A208G, U399C, or U658A as numbered according to SEQ ID NO: 1.

[0210] 88. The nucleic acid molecule of any of embodiments 66-87, which comprises one or more of: a C at position 138, a G at position 208, a G at position 264, a G at position 271, a U at position 307, an A at position 358, a C at position 399, a G at position 405, an A at position 469, or a C at position 488.

[0211] 89. The nucleic acid molecule of any of embodiments 66-88, which comprises one or more of: a C at position 138, a C at position 180, a G at position 405, or a C at position 488.

[0212] 90. The nucleic acid molecule of any of embodiments 66-89, which comprises one or more of: a C at position 180, an A at position 398, or a G at position 405. Page 24 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0213] 91. The nucleic acid molecule of any of embodiments 66-90, which comprises a U138C substitution relative to SEQ ID NO: 1.

[0214] 92. The nucleic acid molecule of any of embodiments 66-91, which comprises a U180C substitution relative to SEQ ID NO: 1.

[0215] 93. The nucleic acid molecule of any of embodiments 66-92, which comprises an A208G substitution relative to SEQ ID NO: 1.

[0216] 94. The nucleic acid molecule of any of embodiments 66-93, which comprises an U264G substitution relative to SEQ ID NO: 1.

[0217] 95. The nucleic acid molecule of any of embodiments 66-94, which comprises an A271G substitution relative to SEQ ID NO: 1.

[0218] 96. The nucleic acid molecule of any of embodiments 66-95, which comprises a G307U substitution relative to SEQ ID NO: 1.

[0219] 97. The nucleic acid molecule of any of embodiments 66-96, which comprises a G358A substitution relative to SEQ ID NO: 1.

[0220] 98. The nucleic acid molecule of any of embodiments 66-97, which comprises an A398G substitution relative to SEQ ID NO: 1.

[0221] 99. The nucleic acid molecule of any of embodiments 66-98, which comprises a U399C substitution relative to SEQ ID NO: 1.

[0222] 100. The nucleic acid molecule of any of embodiments 66-99, which comprises an A405G substitution relative to SEQ ID NO: 1.

[0223] 101. The nucleic acid molecule of any of embodiments 66-100, which comprises an U407C substitution relative to SEQ ID NO: 1.

[0224] 102. The nucleic acid molecule of any of embodiments 66-98101 which comprises an G469A substitution relative to SEQ ID NO: 1.

[0225] 103. The nucleic acid molecule of any of embodiments 66-102, which comprises a U488C substitution relative to SEQ ID NO: 1. Page 25 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0226] 104. The nucleic acid molecule of any of embodiments 66-103, which comprises a G570A substitution relative to SEQ ID NO: 1.

[0227] 105. The nucleic acid molecule of any of embodiments 66-104, which comprises a U658A substitution relative to SEQ ID NO: 1.

[0228] 106. The nucleic acid molecule of any of embodiments 66-105, which comprises all of a U138C substitution, A208G substitution, U264G substitution, A271G substitution, G307U substitution, G358A substitution, U399C substitution, G469A substitution, and U488C substitution relative to SEQ ID NO: 1.

[0229] 107. The nucleic acid molecule of any of embodiments 66-106, which comprises all of a U138C substitution, U180C substitution, A405G substitution, and U488C substitution relative to SEQ ID NO: 1.

[0230] 108. The nucleic acid molecule of any of embodiments 66-107, which comprises all of a U180C substitution, A398G substitution, and A405G substitution relative to SEQ ID NO: 1.

[0231] 109. The nucleic acid of any of embodiments 66-108, wherein the modified RTE-15’ UTR has a length of 400-450, 450-500, 500-550, 550-600, 600-650, 650-700, or 700-710 nucleotides, e.g., 700-710 nucleotides.

[0232] 110. The nucleic acid molecule of any of embodiments 66-109, which comprises a sequence of Table 18 or a sequence having at least 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0233] 111. The nucleic acid molecule of any of embodiments 66-110, which comprises one or more mutations listed in the column “Mutations” of Table 20.

[0234] 112. The nucleic acid molecule of any of embodiments 66-111, which binds a RTE-1 polypeptide according to SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392.

[0235] 113. The nucleic acid molecule of any of embodiments 66-112, which promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; Page 26 of 372 12815032v1Attorney Docket No.: 2017469-0043 b) providing a template RNA comprising, from 5’ to 3’: the nucleic acid molecule of any of embodiments 66-112, a heterologous object sequence, and a RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2; c) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

[0236] 114. The nucleic acid molecule of any of embodiments 66-112, which promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; b) providing a template RNA comprising, from 5’ to 3’: the nucleic acid molecule of any of embodiments 66-112, a heterologous object sequence, and the nucleic acid molecule of any of embodiments 1-65; c) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

[0237] 115. The nucleic acid molecule of embodiment 114, which promotes a greater number of integration events compared to a control template RNA having a sequence comprising, from 5’ to 3’: an RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and a RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0238] 116. The nucleic acid molecule of embodiment 114, which promotes a greater expression of a protein (e.g., GFP or CAR) encoded by a heterologous object sequence, compared to a control template RNA having a sequence comprising, from 5’ to 3’: a RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and a RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0239] 117. The nucleic acid molecule of embodiment 116, wherein the expression of the protein encoded by the heterologous object sequence is at least 1.05, 1.06, 1.07, 1.08, 1.09, 1.1, 1.11, 1.12, 1.13, 1.14, 1.15, 1.16, 1.17, 1.18, 1.19, 1.2, 1.21, 1.22, 1.23, 1.24, 1.25, 1.26, 1.27, 1.28, 1.29, 1.3, 1.31, 1.32, 1.33, 1.34, 1.35, 1.36, 1.37, 1.38, 1.39, 1.4, 1.5, 1.6, 1.7, 1.8, 1.9, 2, Page 27 of 372 12815032v1Attorney Docket No.: 2017469-0043 2.1, 2.2, 2.3, 2.4, or 2.5 times expression of the protein in cells treated with the control template RNA.

[0240] 118. The nucleic acid molecule of any preceding embodiments, wherein the nucleic acid comprises one or more chemically modified nucleotides, optionally wherein the chemically modified nucleotide comprises N1-methylpseudouridine.

[0241] 119. The nucleic acid molecule any of embodiments 1-117, wherein the nucleic acid does not comprise any chemically modified nucleotides.

[0242] 120. The nucleic acid molecule of any of the preceding embodiments, which further comprises a heterologous object sequence.

[0243] 121. The nucleic acid molecule of embodiment 120, wherein the heterologous object sequence is directly adjacent to the modified RTE-15’ UTR.

[0244] 122. The nucleic acid molecule of embodiment 120 or 121, wherein the heterologous object sequence is directly adjacent to the modified RTE-13’ UTR.

[0245] 123. The nucleic acid of embodiment 120 or 122, wherein a nucleic acid linker sequence is situated between the heterologous object sequence and the modified RTE-15’ UTR.

[0246] 124. The nucleic acid of embodiment 120 or 121, wherein a nucleic acid linker sequence is situated between the heterologous object sequence and the modified RTE-13’ UTR.

[0247] 125. A template RNA comprising, from 5’ to 3’: a) a nucleic acid molecule of any of embodiments 66-112; b) a heterologous object sequence; and c) optionally, an RTE-13’ UTR.

[0248] 126. The template RNA of embodiment 125, wherein the RTE-13’ UTR has a nucleic acid sequence according to SEQ ID NO: 2.

[0249] 127. The template RNA of embodiment 125, wherein the RTE-13’ UTR is a modified RTE-13’ UTR.

[0250] 128. The template RNA of any of embodiments 125 or 127, wherein the RTE-13’ UTR is a modified RTE-13’ UTR according to any of embodiments 1-65. Page 28 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0251] 129. A template RNA comprising, from 5’ to 3’: a) optionally, a RTE-15’ UTR; b) a heterologous object sequence; and c) a nucleic acid molecule of any of embodiments 1-65 or 118-121.

[0252] 130. The template RNA of embodiment 129, wherein the RTE-15’ UTR has a nucleic acid sequence according to SEQ ID NO: 1.

[0253] 131. The template RNA of embodiment 129, wherein the RTE-15’ UTR is a modified RTE-15’ UTR.

[0254] 132. A template RNA comprising, from 5’ to 3’: a) a nucleic acid molecule of any of embodiments 66-112; b) a heterologous object sequence; and c) a nucleic acid molecule of any of embodiments 1-65.

[0255] 133. The template RNA of embodiment 132, wherein: the nucleic acid molecule of (a) comprises at least at least 1, 2, 3, 4, 5, 6, 7, 8, or 9 sequence differences relative to SEQ ID NO: 1, and the nucleic acid molecule of (c) comprises at least at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 sequence differences relative to SEQ ID NO: 2.

[0256] 134. The template RNA of any of embodiments 125-133, wherein the heterologous object sequence comprises the sequence of an EF1-alpha promoter.

[0257] 135. The template RNA of any of embodiments 125-134, wherein the heterologous object sequence comprises the sequence of an MND promoter (e.g., an MNDU3 promoter).

[0258] 136. The template RNA of any of embodiments 125-135, wherein the heterologous object sequence encodes a CAR.

[0259] 137. The template RNA of any of embodiments 125-136, which promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: Page 29 of 372 12815032v1Attorney Docket No.: 2017469-0043 a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; and b) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

[0260] 138. The template RNA of embodiment 137, which promotes a greater number of integration events compared to a control template RNA having a sequence comprising, from 5’ to 3’: an RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and an RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0261] 139. The template RNA of embodiment 137, which promotes a greater expression of a protein (e.g., GFP or CAR) encoded by a heterologous object sequence, compared to a control template RNA having a sequence comprising, from 5’ to 3’: an RTE-15’ UTR having a nucleic acid sequence of SEQ ID NO: 1, the heterologous object sequence, and an RTE-13’ UTR having a nucleic acid sequence of SEQ ID NO: 2.

[0262] 140. A gene modifying system comprising: a template RNA of any of embodiments 125-139; and a gene modifying polypeptide, or a nucleic acid encoding the gene modifying polypeptide, the gene modifying polypeptide comprising an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392, or an amino acid sequence having at least 80%, 90%, 95%, 97%, 98%, 99% identity to SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392.

[0263] 141. The system of embodiment 140, which is capable of introducing the sequence of the heterologous object sequence into the genome of a cell.

[0264] 142. The system of any of embodiments 140 or 141, wherein the nucleic acid encoding the gene modifying polypeptide is mRNA.

[0265] 143. A pharmaceutical composition comprising the nucleic acid of any one of embodiments 1-125, the template RNA of any of embodiments 125-139, or the system of any one of embodiments 140-142, and a pharmaceutically acceptable excipient or carrier. Page 30 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0266] 144. The pharmaceutical composition of embodiment 143, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle (LNP).

[0267] 145. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising the gene modifying system, template RNA, or nucleic acid of any one of the preceding embodiments.

[0268] 146. A method of making the cell of embodiment 145, the method comprising providing the cell and contacting the cell with the gene modifying system of any one of the preceding embodiments.

[0269] 147. The method of embodiment 146, wherein providing the cell comprises isolating a cell from a subject (e.g., a mammal, e.g., a human, e.g., a human patient)

[0270] 148. A method of making the nucleic acid or template RNA of any one of embodiments 1-139, the method comprising synthesizing the template RNA in vitro (e.g., by in vitro transcription or solid state synthesis).

[0271] 149. A method of modifying the genome of a mammalian cell, comprising contacting the cell with a system of any of embodiments 140-142, thereby modifying the genome of the mammalian cell.

[0272] 150. A cell made by the method of embodiment 149.

[0273] 151. A method for treating a subject having cancer, the method comprising administering to the subject the gene modifying system of any one of embodiments 140-142, or DNA encoding the same, or the pharmaceutical composition of any one of embodiments 143 or 144, thereby treating the subject having the cancer.

[0274] 152. The gene modifying system of any one of embodiments 140-142, or DNA encoding the same, or the pharmaceutical composition of any one of embodiments 143 or 144, for use in treating cancer.

[0275] 153. The gene modifying system of any one of embodiments 140-142, or DNA encoding the same, or the pharmaceutical composition of any one of embodiments 143 or 144, in the manufacture of a medicament for treating cancer. Page 31 of 372 12815032v1Attorney Docket No.: 2017469-0043 BRIEF DESCRIPTION OF THE DRAWINGS

[0276] FIG.1 is a diagram showing components of an exemplary gene modifying system and illustrates an exemplary system comprising two components: (1) a gene modifying polypeptide labeled as “Driver protein”, and (2) a template RNA. In this example, the gene modifying polypeptide is an RTE-1 retrotransposase. The exemplary template RNA comprises three components, from 5’ to 3’: (1) a 5’-UTR labeled as “UTRretro”, (2) a heterologous object sequence labeled as “Transgene”, and (3) a 3’-UTR labeled as “UTRretro”. In this example, the exemplary UTRs are RTE-1 UTRs.

[0277] FIG.2 is a diagram that illustrate the nomenclature of different positions and regions of an exemplary RTE-13’-UTR sequence of SEQ ID NO: 2.

[0278] FIGS.3A-3B are bar graphs showing the fold change of percentage of GFP+ cells (relative to percentage of GFP+ cells in wild-type controls) after treating HEK293T cells or primary T cells from one of two donors (D076 and D744) with a gene modifying system comprising a nucleic acid encoding a wild-type RTE-1 polypeptide and a template RNA having: a 5’-UTR, a heterologous object sequence encoding a GFP reporter, and a 3’-UTR. FIG.3A shows the performance of the indicated 5’ UTRs. FIG.3B shows the performance of the indicated 3’ UTRs.

[0279] FIG.4 is a schematic showing an experimental protocol as described herein, e.g. in Example 2.

[0280] FIG.5A is a table describing exemplary template RNAs comprising modified RTE-15’ and / or 3’ UTRs as described herein, including the total number of mutations in the 5’ UTRs and 3’ UTRs and the mutations (e.g., substitutions and / or deletions at specific residues). For example, exemplary template RNA RNAIVT8311 comprised a WT 5’ UTR comprising zero mutations and a mutant 3’ UTR comprising two mutations (A21T substitution and deletion of residue 32). Exemplary template RNAs further comprised a heterologous object (GFP reporter) sequence.

[0281] FIG.5B is a bar graph showing percentage of GFP+ cells following transfection with gene modifying systems comprising exemplary template RNAs as shown in FIG.5A, as assessed with flow cytometry relative to template RNAs comprising wild-type UTRs (dotted line). Page 32 of 372 12815032v1Attorney Docket No.: 2017469-0043 DETAILED DESCRIPTION

[0282] This disclosure relates to compositions, systems and methods for targeting, editing, modifying or manipulating a DNA sequence (e.g., inserting a heterologous object DNA sequence into a target site of a mammalian genome) at one or more locations in a DNA sequence in a cell, tissue or subject, e.g., in vivo, in vitro or ex vivo. The object DNA sequence may include, e.g., a coding sequence, a regulatory sequence, a gene expression unit.

[0283] More specifically, the disclosure provides retrotransposon-based systems for inserting a sequence of interest (e.g., a heterologous object sequence) into the genome. Examples of retrotransposon elements are listed, e.g., in Tables 10, 11, X, 3A, 3B, and Z1 of PCT Publication No. WO / 2021 / 178717, incorporated herein by reference in its entirety. Definitions

[0284] Antigen binding domain: The term “antigen binding domain” as used herein refers to that portion of antibody or a chimeric antigen receptor which binds an antigen. In some embodiments, an antigen binding domain binds to a cell surface antigen of a cell. In some embodiments an antigen binding domain binds an antigen characteristic of a cancer, e.g., a tumor associated antigen in a neoplastic cell. In some embodiments, an antigen binding domain binds an antigen characteristic of an infectious disease, e.g. a virus associated antigen in a virus infected cell. In some embodiments, an antigen binding domain binds an antigen characteristic of a cell targeted by a subject’s immune system in an autoimmune disease, e.g., a self-antigen. In some embodiments, an antigen binding domain is or comprises an antibody or antigen-binding portion thereof. In some embodiments, an antigen binding domain is or comprises an scFv or Fab.

[0285] Bulge sequence: As used herein, a “bulge,” as used with respect to a portion of a nucleic acid molecule (e.g., an RNA), refers to a single-stranded nucleic acid sequence (i.e., comprising one or more nucleotides not base-paired with a nucleotide of an opposite strand) within the nucleic acid molecule bounded on both the 5’ and 3’ ends by at least one base pair. In some instances, the opposite strand from the bulge may also include one or more nucleotides that are not base paired with nucleotides from the bulge sequence. In some instances, the opposite strand Page 33 of 372 12815032v1Attorney Docket No.: 2017469-0043 from the bulge may not include any nucleotides that are not base paired. In some instances, a bulge includes at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides that are not base-paired. In some instance, the bulge consists of 1, 2, 3, or 4 nucleotides that are not base-paired.

[0286] Domain: The term “domain” as used herein refers to a structure of a biomolecule that contributes to a specified function of the biomolecule. A domain may comprise a contiguous region (e.g., a contiguous sequence) or distinct, non-contiguous regions (e.g., non-contiguous sequences) of a biomolecule. Examples of protein domains include, but are not limited to, an endonuclease domain, a DNA binding domain, a reverse transcriptase domain; an example of a domain of a nucleic acid is a regulatory domain, such as a transcription factor binding domain.

[0287] Exogenous: As used herein, the term “exogenous,” when used with reference to a biomolecule (such as a nucleic acid sequence or polypeptide) means that the biomolecule was introduced into a host genome, cell, or organism by the hand of man. For example, a nucleic acid that is as added into an existing genome, cell, tissue, or subject using recombinant DNA techniques or other methods is exogenous to the existing nucleic acid sequence, cell, tissue or subject.

[0288] Integration: The term “integration”, as used herein, refers to the insertion of a first nucleic acid sequence (e.g., a heterologous object sequence) into a second nucleic acid sequence (e.g., in a host cell genome). “Integration” includes situations where the informational content of a heterologous object sequence of a template RNA being integrated into a genome. Typically, the RNA is not physically incorporated into the host genome but rather used as a template for synthesis of DNA, which DNA can be incorporated into the host genome.

[0289] Heterologous: The term “heterologous”, when used to describe a first element in reference to a second element means that the first element and second element do not exist in nature disposed as described. For example, a heterologous polypeptide, nucleic acid molecule, construct or sequence refers to (a) a polypeptide, nucleic acid molecule or portion of a polypeptide or nucleic acid molecule sequence that is not native to a cell in which it is expressed, (b) a polypeptide or nucleic acid molecule or portion of a polypeptide or nucleic acid molecule that has been altered or mutated relative to its native state, or (c) a polypeptide or nucleic acid molecule with an altered expression as compared to the native expression levels under similar conditions. For example, a heterologous regulatory sequence (e.g., promoter, enhancer) may be Page 34 of 372 12815032v1Attorney Docket No.: 2017469-0043 used to regulate expression of a gene or a nucleic acid molecule in a way that is different than the gene or a nucleic acid molecule is normally expressed in nature. In another example, a heterologous domain of a polypeptide or nucleic acid sequence (e.g., a DNA binding domain of a polypeptide or nucleic acid encoding a DNA binding domain of a polypeptide) may be disposed relative to other domains or may be a different sequence or from a different source, relative to other domains or portions of a polypeptide or its encoding nucleic acid. In certain embodiments, a heterologous nucleic acid molecule may exist in a native host cell genome, but may have an altered expression level or have a different sequence or both. In other embodiments, heterologous nucleic acid molecules may not be endogenous to a host cell or host genome but instead may have been introduced into a host cell by transformation (e.g., transfection, electroporation), wherein the added molecule may integrate into the host genome or can exist as extra- chromosomal genetic material either transiently (e.g., mRNA) or semi-stably for more than one generation (e.g., episomal viral vector, plasmid or other self-replicating vector). In some embodiments, a domain is heterologous relative to another domain, if the first domain is not naturally comprised in the same polypeptide as the other domain (e.g., a fusion between two domains of different proteins from the same organism).

[0290] Modified RTE-15’ UTR: The term “modified RTE-15’ UTR”, as used herein, refers to a nucleotide sequence having one or more sequence differences relative to SEQ ID NO: 1, and which is capable of directing insertion of the sequence of a heterologous object sequence into a genome, when present in a template RNA comprising, from 5’ to 3’, the modified RTE-15’ UTR, the heterologous object sequence, and a wild-type or modified RTE-13’ UTR, in the presence of an RTE-1 polypeptide. In some embodiments, the modified RTE-15’ UTR has a sequence at least which has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity to SEQ ID NO: 1.

[0291] Modified RTE-13’ UTR: The term “modified RTE-13’ UTR”, as used herein, refers to a nucleotide sequence having one or more sequence differences relative to SEQ ID NO: 2, and which is capable of directing insertion of the sequence of a heterologous object sequence into a genome, when present in a template RNA comprising, from 5’ to 3’, a wild-type or modified RTE-15’ UTR, the heterologous object sequence, and the modified RTE-13’ UTR, in the presence of an RTE-1 polypeptide. In some embodiments, the modified RTE-13’ UTR has a Page 35 of 372 12815032v1Attorney Docket No.: 2017469-0043 sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 98% identity to SEQ ID NO: 2. In some embodiments, the modified RTE-13’ UTR comprises 5, 6, 7, or all of: Region A, Region B, Region C, Region D, Region E, Region F, Region G, or Region H as annotated in Fig. 2.

[0292] Mutation or Mutated: The term “mutated” when applied to nucleic acid sequences means that nucleotides in a nucleic acid sequence may be inserted, deleted or changed compared to a reference (e.g., native) nucleic acid sequence. A single alteration may be made at a locus (a point mutation) or multiple nucleotides may be inserted, deleted, or changed at a single locus. In addition, one or more alterations may be made at any number of loci within a nucleic acid sequence. A nucleic acid sequence may be mutated by any method known in the art. In some embodiments a mutation occurs naturally. In some embodiments a desired mutation can be produced by a system described herein.

[0293] Nucleic acid molecule: “Nucleic acid molecule” refers to both RNA and DNA molecules including, without limitation, cDNA, genomic DNA and mRNA, and also includes synthetic nucleic acid molecules, such as those that are chemically synthesized or recombinantly produced, such as RNA templates, as described herein. The nucleic acid molecule can be double stranded or single stranded, circular or linear. If single stranded, the nucleic acid molecule can be the sense strand or the antisense strand. Unless otherwise indicated, and as an example for all sequences described herein under the general format “SEQ ID NO:,” “nucleic acid comprising SEQ ID NO:1” refers to a nucleic acid, at least a portion which has either (i) the sequence of SEQ ID NO:1, or (ii) a sequence complimentary to SEQ ID NO:1. The choice between the two is dictated by the context in which SEQ ID NO:1 is used. For instance, if the nucleic acid is used as a probe, the choice between the two is dictated by the requirement that the probe be complimentary to the desired target. Nucleic acid sequences of the present disclosure may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more naturally occurring nucleotides with an analog, inter- nucleotide modifications such as uncharged linkages (for example, methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), charged linkages (for example, phosphorothioates, phosphorodithioates, etc.), pendant moieties, (for example, polypeptides), Page 36 of 372 12815032v1Attorney Docket No.: 2017469-0043 intercalators (for example, acridine, psoralen, etc.), chelators, alkylators, and modified linkages (for example, alpha anomeric nucleic acids, etc.). Also included are synthetic molecules that mimic polynucleotides in their ability to bind to a designated sequence via hydrogen bonding and other chemical interactions. Such molecules are known in the art and include, for example, those in which peptide linkages substitute for phosphate linkages in the backbone of a molecule. Other modifications can include, for example, analogs in which the ribose ring contains a bridging moiety or other structure such as modifications found in “locked” nucleic acids. In some embodiments, a given nucleic acid molecule is a portion of a larger nucleic acid molecule.

[0294] Nucleotide difference: The term “nucleotide difference”, as used herein, refers to a difference between a first nucleotide sequence and a reference nucleotide sequence. In some embodiments, the nucleotide difference is a substitution, insertion, or deletion. In some embodiments, the nucleotide difference is an insertion of 1-10 nucleotides. In some embodiments, the nucleotide difference is a deletion of 1-3 nucleotides. In some embodiments, the nucleotide sequence comprises a first nucleotide difference that is a first substitution of a first nucleotide, and further comprises a second nucleotide difference that is a second substitution of a second nucleotide. The first and second substitutions may be adjacent or may be separated by one or more nucleotides.

[0295] Gene expression unit: a “gene expression unit” is a nucleic acid sequence comprising at least one regulatory nucleic acid sequence operably linked to at least one effector sequence. A first nucleic acid sequence is operably linked with a second nucleic acid sequence when the first nucleic acid sequence is placed in a functional relationship with the second nucleic acid sequence. For instance, a promoter or enhancer is operably linked to a coding sequence if the promoter or enhancer affects the transcription or expression of the coing sequence. Operably linked DNA sequences may be contiguous or non-contiguous. Where necessary to join two protein-coding regions, operably linked sequences may be in the same reading frame.

[0296] Gene modifying polypeptide: A “gene modifying polypeptide,” as used herein, refers to a polypeptide comprising a retrotransposase reverse transcriptase domain and a retrotransposase endonuclease domain, or a polypeptide comprising an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% amino acid sequence identity to said domains, which is capable of integrating a nucleic acid sequence (e.g., a sequence provided on a template Page 37 of 372 12815032v1Attorney Docket No.: 2017469-0043 nucleic acid) into a target DNA molecule (e.g., in a mammalian host cell, such as a genomic DNA molecule in the host cell). In some embodiments, the endonuclease domain is a catalytically inactive endonuclease domain. In some embodiments, the retrotransposase reverse transcriptase domain and a retrotransposase endonuclease domain are derived from the same retrotransposase. In some embodiments, the gene modifying polypeptide is capable of integrating the sequence substantially without relying on host machinery. In some embodiments, the gene modifying polypeptide integrates a sequence into a random position in a genome, and in some embodiments, the gene modifying polypeptide integrates a sequence into a specific target site. In some embodiments, a gene modifying polypeptide includes one or more domains that, collectively, facilitate 1) binding the template nucleic acid, 2) binding the target DNA molecule, and 3) facilitate integration of the at least a portion of the template nucleic acid into the target DNA. Gene modifying polypeptides include both naturally occurring polypeptides as well as engineered variants of the foregoing, e.g., having one or more amino acid substitutions to the naturally occurring sequence. Gene modifying polypeptides also include heterologous constructs, e.g., where one or more of the domains recited above are heterologous to each other, whether through a heterologous fusion (or other conjugate) of otherwise wild-type domains, as well as fusions of modified domains, e.g., by way of replacement or fusion of a heterologous sub- domain or other substituted domain. Exemplary gene modifying polypeptides, and systems comprising them and methods of using them, that can be used in the methods provided herein are described, e.g., in WO / 2021 / 178717, which is incorporated herein by reference, including Tables 10, 11, X, 3A, 3B, and Z1 therein. In some embodiments, a gene modifying polypeptide integrates a sequence into a gene. In some embodiments, a gene modifying polypeptide integrates a sequence into a sequence outside of a gene. A “gene modifying system,” as used herein, refers to a system comprising a gene modifying polypeptide and a template nucleic acid.

[0297] Host: The terms “host genome” or “host cell,” as used herein, refer to a cell and / or its genome into which protein and / or genetic material has been introduced. It should be understood that such terms are intended to refer not only to the particular subject cell and / or genome, but to the progeny of such a cell and / or the genome of the progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term “host cell” as used herein. A host genome or host cell may be an Page 38 of 372 12815032v1Attorney Docket No.: 2017469-0043 isolated cell or cell line grown in culture, or genomic material isolated from such a cell or cell line, or may be a host cell or host genome which composing living tissue or an organism. In some instances, a host cell may be an animal cell or a plant cell, e.g., as described herein. In certain instances, a host cell may be a bovine cell, horse cell, pig cell, goat cell, sheep cell, chicken cell, or turkey cell. In certain instances, a host cell may be a corn cell, soy cell, wheat cell, or rice cell.

[0298] Stem-loop sequence: As used herein, a “stem-loop sequence” refers to a nucleic acid sequence (e.g., RNA sequence) with sufficient self-complementarity to form a stem-loop, e.g., having a stem comprising at least two (e.g., 3, 4, 5, 6, 7, 8, 9, or 10) base pairs, and a loop with at least three (e.g., four) base pairs. The stem may comprise mismatches or bulges.

[0299] RTE-1 polypeptide: The term “RTE-1 polypeptide” refers to a polypeptide having retrotransposase activity and which is capable of directing insertion of the sequence of a heterologous object sequence into a genome, in the presence of a template RNA comprising, from 5’ to 3’, a wild-type or modified RTE-15’ UTR, the heterologous object sequence, and a wild-type or modified RTE-13’ UTR. In some embodiments, the RTE-1 polypeptide has an amino acid sequence having at least 80%, 90%, 95%, 97%, 98%, 99% identity to SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392. Gene modifying polypeptides

[0300] Non-long terminal repeat (LTR) retrotransposons are a type of mobile genetic elements that are widespread in eukaryotic genomes. They include, for example, the apurinic / apyrimidinic endonuclease (APE)-type, the restriction enzyme-like endonuclease (RLE)-type, and the Penelope-like element (PLE)-type.

[0301] The APE class retrotransposons are comprised of two functional domains: an endonuclease / DNA binding domain, and a reverse transcriptase domain. Examples of APE-class retrotransposons can be found, for example, in Table 1 of PCT Application No. PCT / US2019 / 048607, incorporated herein by reference in its entirety, including the sequence listing and sequences referred to in Table 1 therein. Page 39 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0302] The RLE class are comprised of three functional domains: a DNA binding domain, a reverse transcription domain, and an endonuclease domain. Examples of RLE-class retrotransposons can be found, for example, in Table 2 of PCT Application No. PCT / US2019 / 048607, incorporated herein by reference in its entirety, including the sequence listing and sequences referred to in Table 2 therein.

[0303] The reverse transcriptase domain of non-LTR retrotransposon functions by binding an RNA sequence template and reverse transcribing it into the host genome’s target DNA. The RNA sequence template has a 3’ untranslated region which is specifically bound to the retrotransposase, and a variable 5’ region generally having Open Reading Frame(s) (“ORF”) encoding retrotransposase proteins. The RNA sequence template may also comprise a 5’ untranslated region which specifically binds the retrotransposase.

[0304] As described herein, the elements of such retrotransposons can be functionally modularized and / or modified to target, edit, modify or manipulate a target DNA sequence, e.g., to insert an object (e.g., heterologous) nucleic acid sequence into a target genome, e.g., a mammalian genome, by reverse transcription. In some embodiments, a gene modifying system comprises: (A) a polypeptide or a nucleic acid encoding a polypeptide, wherein the polypeptide comprises (i) a retrotransposase reverse transcriptase domain, and (ii) a retrotransposase endonuclease domain that contains DNA binding functionality; and (B) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence. The RNA template element of a gene modifying system is typically heterologous to the polypeptide element and provides an object sequence to be inserted (reverse transcribed) into the host genome.

[0305] In some embodiments, an amino acid sequence encoded by an element of Table 1 is an amino acid sequence encoded by the full length sequence of an element listed in Table 1, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the full-length sequence of an element listed in Table 1 may comprise one or more (e.g., all of) of a 5’ UTR, polypeptide-encoding sequence, or 3’ UTR of a retrotransposon as described herein. In some embodiments, an amino acid sequence of Table 1 is an amino acid sequence encoded by the full length sequence of an element listed in Table 1, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity Page 40 of 372 12815032v1Attorney Docket No.: 2017469-0043 thereto. In some embodiments, the amino acid sequence of the reverse transcriptase domain of a gene modifying system is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of a reverse transcriptase domain of a retrotransposon whose DNA sequence is referenced in Table 1. In some embodiments, the amino acid sequence of the endonuclease domain of a gene modifying system is at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99% identical to the amino acid sequence of a endonuclease domain of a retrotransposon whose DNA sequence is referenced in Table 1. In certain embodiments, the heterologous DNA-binding domain is a DNA binding domain of a retrotransposon described in Table 1 herein. In some embodiments, a 5’ UTR of an element of Table 1 comprises a 5’ UTR of the full length sequence of an element listed in Table 1, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, a 3’ UTR of an element of Table 1 comprises a 3’ UTR of the full length sequence of an element listed in Table 1, or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.

[0306] Also indicated in Table 1 are the host organisms from which the nucleic acid sequences were obtained and a listing of domains present within the polypeptide encoded by the open reading frame of the nucleic acid sequence.

[0307] In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of SEQ ID NO: 402 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a functional fragment thereof (e.g., having one or both of reverse transcriptase activity and endonuclease activity), wherein the initial M of SEQ ID NO: 402 may be present or absent.

[0308] In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of SEQ ID NO: 403 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a functional fragment thereof (e.g., having one or both of reverse transcriptase activity and endonuclease activity), wherein the initial M of SEQ ID NO: 403 may be present or absent. Page 41 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0309] In some embodiments, a gene modifying polypeptide described herein comprises an NLS with the amino acid sequence PKKKRKV (SEQ ID NO: 378) and a GGS linker. In some embodiments, a gene modifying polypeptide comprises the amino acid sequence (bold: NLS; italics: linker): MPKKKRKVGGSDSTAHPNQGRGLEKVSQTLPALQTPGQHTAAGGSSPLSGRNQRKNTK KLLLGAWNIRTLLDRENTPRPERRTALIGKELARYNIDIAALSETRLPEEGSLSEPTTGYTF FWKGRASNEDRIHGVGLAIKTSLLKQLPDLPVGISERLMKIRLPLSKDRYATIISAYAPTLT STEETIEQFYSDLSAVLHSVPTNDKLILLGDFNARVGQDHERWKGVLGKHGVGKMNNN GLLLLSKCSEFELTITNTVFRMANKYKTTWMHPRSKQWHLIDYIIVRRRDIQDVKITRAM RGAECWTDHRLVRATLQMRIAPRHPKRAQTVRAFYNVSRLRDPSYLQTFQSCLDDKLS AKGPLTGSSTEKWNQFRDAVKETSKAVLGPKQRNHQDWFDENNTAIEDLLSKKNKAFM EWQNNPNSAPKKDRFKSLQATAQREIRKMQDRWWEKKAEEIQRFADMKNYKQFFSAL KTVYGPLKPTTTPLLSSDGDTLIKDKKGISNRWKEHFSQLLNRPSSVDQSALDQIPQNRTI EQLDVPPSIEEVQKAIKQMSAGKAPGKDGIPTEVYKALNGKALQAFHIVLTSIWEEEDMP PELRDASIVALYKNKGSRAACDNYRGISLLSTAGKILARVILNRLLSSVSEQNLPESQCGF RPDRSTIDMVFTVRQMQEKCLEQNLSLYIVFIDLTKAFDTVNRDALWVILSKLGCPAKFV KLIQLFHVDMTGEVLSGGETSDRFNISNGVKQGCVLAPVLFNLFFTQVLRHAVMDLDLG VYIKYRLDGSLFDLRRLTAKTKTTERLILEALFADDCALMAHQENHLQTIVDRFSTATKL FGLTISLSKTEVLFQPAPGRPTNQPCITIDGTQLSNVNTFKYLGSTIANDGSLDHEINARIQ KASQALGRLRCKVLQHRGVSTATKLKVYNAVVLSSLLYGCETWTLYRKHMKQLEQFHQ RSLRSIMRIRWQDRITNQEVLDRANSTSIEVMVLKTQLRWSGHVIRMDPQRIPRQVFYGE LSAGLRKQGRPKKRFKDQLKSNLKWAGITPKQLELAASDRSSWRTHINHAATTFEDERR RRLAAARERRHQATTAPPVTTGVPCPMCHKLCASAFGLQSHMRVHRR (SEQ ID NO:392).

[0310] In some embodiments, the gene modifying polypeptide comprises an amino acid sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical to SEQ ID NO: 392, or a functional fragment thereof (e.g., having one or both of reverse transcriptase activity and endonuclease activity), wherein the initial M of SEQ ID NO: 392 may be present or absent. Page 42 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0311] In some embodiments, the gene modifying polypeptide is encoded by a nucleic acid sequence of (underlining = 5’ UTR; bold italics underlining = start codon; bold = NLS; italics = linker; normal text = RTE-1 protein coding sequence; Courier bold = XTEN linker; Courier italics = HitBit tag; Courier underline = 3’ UTR): AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGCCAA AGAAGAAAAGAAAAGUAGGAGGUUCCGACUCCACAGCUCAUCCCAACCAGGGAC GUGGCCUGGAGAAGGUGUCCCAGACUCUGCCUGCGCUGCAGACCCCAGGGCAGCA CACUGCCGCUGGGGGAUCCUCUCCCCUGUCCGGCAGGAAUCAGCGGAAGAACACC AAAAAGCUGUUGCUGGGCGCCUGGAACAUCCGCACCCUGCUCGACCGCGAAAAUA CUCCCCGUCCUGAAAGACGCACGGCACUGAUUGGGAAAGAGCUCGCCAGAUAUAA CAUCGACAUCGCUGCCCUGUCCGAAACCCGCCUGCCAGAGGAAGGAUCUCUGAGC GAGCCUACCACUGGCUACACGUUCUUUUGGAAGGGCAGAGCAAGCAACGAGGACC GUAUCCAUGGUGUAGGCCUGGCUAUCAAGACCUCUCUCCUGAAACAACUCCCAGA UCUGCCAGUCGGCAUCUCAGAGAGACUGAUGAAGAUCAGGCUUCCUCUCUCAAAA GACCGCUACGCAACCAUCAUUUCCGCAUACGCGCCGACCCUGACUAGCACCGAGG AAACCAUUGAGCAGUUCUACUCUGAUCUUUCAGCUGUUUUGCACUCCGUGCCUAC CAACGACAAGCUGAUCCUCCUGGGCGACUUCAACGCUCGCGUGGGUCAGGACCAU GAGCGCUGGAAAGGCGUGCUGGGUAAGCAUGGCGUGGGUAAGAUGAACAAUAAC GGCCUCCUGCUCCUGUCUAAGUGUAGCGAAUUCGAGCUGACUAUUACCAACACUG UUUUUCGUAUGGCUAACAAGUACAAAACUACAUGGAUGCACCCCAGGAGCAAGC AGUGGCACCUGAUUGACUACAUUAUCGUGCGUAGGCGGGAUAUCCAGGACGUGA AGAUUACCCGUGCAAUGCGCGGCGCUGAGUGCUGGACAGACCAUCGCCUGGUGAG AGCUACUCUCCAGAUGAGGAUCGCACCGCGUCACCCUAAGCGCGCCCAGACCGUG CGGGCCUUUUACAACGUAUCACGGCUUCGGGAUCCUUCCUACCUUCAGACAUUUC AAUCUUGCCUCGAUGACAAGCUCUCUGCUAAGGGUCCCCUCACCGGAUCUUCAAC AGAGAAGUGGAACCAAUUCCGCGAUGCGGUGAAGGAAACCAGCAAAGCAGUCCU GGGACCGAAGCAGCGCAACCACCAAGACUGGUUUGACGAAAAUAACACUGCCAUC GAAGACCUCCUGUCUAAAAAGAACAAAGCAUUUAUGGAGUGGCAGAAUAACCCA AACAGUGCACCCAAGAAAGACCGUUUCAAAAGCCUGCAGGCCACCGCGCAGCGCG AGAUCCGGAAGAUGCAGGACCGCUGGUGGGAAAAAAAGGCCGAAGAGAUCCAGA GGUUCGCGGAUAUGAAAAACUACAAGCAGUUCUUUAGCGCACUGAAGACCGUGU Page 43 of 372 12815032v1Attorney Docket No.: 2017469-0043 ACGGACCUCUCAAACCCACCACGACUCCCCUGCUCUCUUCCGAUGGUGACACUCU CAUCAAGGACAAAAAGGGAAUCAGCAACCGCUGGAAAGAACACUUUUCUCAACU UCUCAACCGCCCAAGCUCCGUGGACCAGAGCGCUCUGGACCAGAUCCCGCAGAAC CGUACCAUCGAGCAGCUGGACGUGCCUCCCUCCAUCGAAGAGGUCCAGAAGGCUA UCAAGCAGAUGUCUGCUGGGAAGGCUCCUGGCAAGGAUGGGAUCCCAACAGAGG UCUAUAAAGCUCUCAACGGAAAGGCGCUGCAGGCCUUUCACAUUGUCCUGACCUC CAUCUGGGAGGAAGAGGACAUGCCACCUGAACUUAGGGACGCCUCUAUUGUGGCC CUGUACAAGAAUAAGGGCAGCCGCGCAGCUUGUGAUAACUAUCGUGGCAUCUCAC UCCUGAGUACAGCUGGCAAGAUUCUGGCUCGUGUGAUCCUGAACCGGUUGCUGA GUUCCGUUUCCGAGCAGAACCUGCCCGAAUCACAGUGCGGCUUCCGCCCCGAUCG GUCCACCAUCGAUAUGGUGUUCACCGUGCGUCAGAUGCAGGAGAAGUGCCUGGA ACAGAAUCUGAGCCUGUACAUCGUUUUCAUCGAUCUGACUAAGGCCUUCGACACG GUGAACCGCGAUGCCCUGUGGGUGAUCCUGAGCAAGCUGGGGUGCCCCGCUAAGU UCGUAAAGCUCAUCCAGUUGUUCCACGUUGAUAUGACGGGAGAGGUCCUGUCCG GUGGAGAAACCUCCGACCGGUUUAACAUCUCCAAUGGCGUGAAACAGGGCUGCGU CCUGGCCCCAGUGCUGUUCAACCUGUUUUUCACACAGGUGCUGAGACAUGCGGUU AUGGAUCUUGACCUGGGUGUGUAUAUCAAGUAUCGCUUGGACGGCAGUCUGUUC GAUCUCAGACGCCUGACAGCUAAGACUAAGACUACCGAGCGCCUCAUUCUCGAAG CCCUGUUUGCCGACGAUUGUGCCCUUAUGGCCCACCAGGAGAACCACCUGCAGAC GAUUGUCGACCGCUUCUCUACAGCGACCAAGCUGUUCGGCCUCACCAUCUCCCUC UCAAAGACCGAAGUCCUGUUUCAGCCCGCACCCGGGAGACCCACGAAUCAGCCGU GCAUUACCAUUGACGGGACCCAACUUUCCAACGUCAACACCUUCAAGUAUCUGGG AUCUACCAUCGCCAAUGAUGGCAGCCUGGAUCAUGAAAUCAACGCAAGGAUCCAA AAGGCUAGCCAGGCACUGGGCCGCCUCCGGUGCAAGGUCCUGCAGCACCGGGGCG UUUCCACUGCCACUAAACUGAAGGUGUACAACGCUGUCGUGCUCUCUUCCCUGCU UUAUGGCUGCGAAACAUGGACUCUGUACCGGAAGCAUAUGAAACAGCUUGAGCA GUUCCACCAGAGGAGCUUGCGUAGCAUCAUGAGGAUCCGCUGGCAGGACCGCAUC ACCAACCAGGAGGUGCUGGACAGGGCUAAUUCCACCAGCAUCGAAGUUAUGGUGC UGAAGACCCAGCUCCGUUGGUCUGGCCAUGUCAUUCGGAUGGACCCUCAGCGUAU CCCCCGCCAGGUAUUCUACGGGGAACUGUCUGCGGGGCUCAGAAAGCAGGGCCGC CCAAAGAAACGCUUUAAGGAUCAGCUGAAAAGCAAUCUGAAGUGGGCCGGCAUC Page 44 of 372 12815032v1Attorney Docket No.: 2017469-0043 ACCCCCAAGCAGCUGGAGCUGGCGGCAUCUGACCGCUCCAGCUGGCGCACCCACA UUAACCACGCAGCUACUACCUUUGAGGACGAACGCCGGCGCAGACUGGCCGCUGC ACGCGAGAGGCGCCACCAAGCCACUACCGCCCCUCCCGUGACCACUGGGGUCCCU UGCCCUAUGUGCCACAAGCUGUGUGCGUCUGCCUUCGGUUUGCAGUCCCACAUGC GUGUGCAUCGCCGCAGCGGCAGCGAGACUCCCGGGACCUCAGAGUCCGCCACACCCGAAAGU GUGAGCGGCUGGCGGCUGUUCAAGAAGAUUAGCUGAGCUGGAGCCUCGGUGGCCAUGCUUCUUG CCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAU AAAGUCUGAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 900).In some embodiments, the gene modifying polypeptide comprises an amino acid sequence encoded by the nucleic acid sequence of nucleotides 48-3672 of SEQ ID NO: 900. In some embodiments, the gene modifying polypeptide comprises an amino acid sequence encoded by the nucleic acid sequence of nucleotides 81-3672 of SEQ ID NO: 900. In some embodiments, the gene modifying polypeptide comprises an amino acid sequence encoded by the nucleic acid sequence of nucleotides 48-3407 of SEQ ID NO: 900. In some embodiments, the gene modifying polypeptide comprises an amino acid sequence encoded by the nucleic acid sequence of nucleotides 81-3407 of SEQ ID NO: 900. In certain embodiments, the gene modifying polypeptide comprises an N- terminal methionine (e.g., encoded by an AUG codon).

[0312] In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises the full length sequence of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 48-3672 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 81-3672 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 48-3407 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 81-3407 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 48- 3407 and 3492-3672 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 81-3407 and 3492-3672 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene Page 45 of 372 12815032v1Attorney Docket No.: 2017469-0043 modifying polypeptide as described herein comprises nucleotides 48-3407 and 3492-3592 of SEQ ID NO: 900. In some embodiments, a nucleic acid molecule encoding a gene modifying polypeptide as described herein comprises nucleotides 81-3407 and 3492-3592 of SEQ ID NO: 900. In certain embodiments, the nucleic acid molecule does not comprise nucleotides 3456- 3488 of SEQ ID NO: 900.

[0313] In some embodiments, the gene modifying polypeptide is encoded by a nucleic acid sequence of: AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGGACU CCACAGCUCAUCCCAACCAGGGACGUGGCCUGGAGAAGGUGUCCCAGACUCUGCC UGCGCUGCAGACCCCAGGGCAGCACACUGCCGCUGGGGGAUCCUCUCCCCUGUCC GGCAGGAAUCAGCGGAAGAACACCAAAAAGCUGUUGCUGGGCGCCUGGAACAUCC GCACCCUGCUCGACCGCGAAAAUACUCCCCGUCCUGAAAGACGCACGGCACUGAU UGGGAAAGAGCUCGCCAGAUAUAACAUCGACAUCGCUGCCCUGUCCGAAACCCGC CUGCCAGAGGAAGGAUCUCUGAGCGAGCCUACCACUGGCUACACGUUCUUUUGGA AGGGCAGAGCAAGCAACGAGGACCGUAUCCAUGGUGUAGGCCUGGCUAUCAAGA CCUCUCUCCUGAAACAACUCCCAGAUCUGCCAGUCGGCAUCUCAGAGAGACUGAU GAAGAUCAGGCUUCCUCUCUCAAAAGACCGCUACGCAACCAUCAUUUCCGCAUAC GCGCCGACCCUGACUAGCACCGAGGAAACCAUUGAGCAGUUCUACUCUGAUCUUU CAGCUGUUUUGCACUCCGUGCCUACCAACGACAAGCUGAUCCUCCUGGGCGACUU CAACGCUCGCGUGGGUCAGGACCAUGAGCGCUGGAAAGGCGUGCUGGGUAAGCA UGGCGUGGGUAAGAUGAACAAUAACGGCCUCCUGCUCCUGUCUAAGUGUAGCGA AUUCGAGCUGACUAUUACCAACACUGUUUUUCGUAUGGCUAACAAGUACAAAAC UACAUGGAUGCACCCCAGGAGCAAGCAGUGGCACCUGAUUGACUACAUUAUCGUG CGUAGGCGGGAUAUCCAGGACGUGAAGAUUACCCGUGCAAUGCGCGGCGCUGAG UGCUGGACAGACCAUCGCCUGGUGAGAGCUACUCUCCAGAUGAGGAUCGCACCGC GUCACCCUAAGCGCGCCCAGACCGUGCGGGCCUUUUACAACGUAUCACGGCUUCG GGAUCCUUCCUACCUUCAGACAUUUCAAUCUUGCCUCGAUGACAAGCUCUCUGCU AAGGGUCCCCUCACCGGAUCUUCAACAGAGAAGUGGAACCAAUUCCGCGAUGCGG UGAAGGAAACCAGCAAAGCAGUCCUGGGACCGAAGCAGCGCAACCACCAAGACUG GUUUGACGAAAAUAACACUGCCAUCGAAGACCUCCUGUCUAAAAAGAACAAAGC AUUUAUGGAGUGGCAGAAUAACCCAAACAGUGCACCCAAGAAAGACCGUUUCAA Page 46 of 372 12815032v1Attorney Docket No.: 2017469-0043 AAGCCUGCAGGCCACCGCGCAGCGCGAGAUCCGGAAGAUGCAGGACCGCUGGUGG GAAAAAAAGGCCGAAGAGAUCCAGAGGUUCGCGGAUAUGAAAAACUACAAGCAG UUCUUUAGCGCACUGAAGACCGUGUACGGACCUCUCAAACCCACCACGACUCCCC UGCUCUCUUCCGAUGGUGACACUCUCAUCAAGGACAAAAAGGGAAUCAGCAACCG CUGGAAAGAACACUUUUCUCAACUUCUCAACCGCCCAAGCUCCGUGGACCAGAGC GCUCUGGACCAGAUCCCGCAGAACCGUACCAUCGAGCAGCUGGACGUGCCUCCCU CCAUCGAAGAGGUCCAGAAGGCUAUCAAGCAGAUGUCUGCUGGGAAGGCUCCUG GCAAGGAUGGGAUCCCAACAGAGGUCUAUAAAGCUCUCAACGGAAAGGCGCUGC AGGCCUUUCACAUUGUCCUGACCUCCAUCUGGGAGGAAGAGGACAUGCCACCUGA ACUUAGGGACGCCUCUAUUGUGGCCCUGUACAAGAAUAAGGGCAGCCGCGCAGCU UGUGAUAACUAUCGUGGCAUCUCACUCCUGAGUACAGCUGGCAAGAUUCUGGCUC GUGUGAUCCUGAACCGGUUGCUGAGUUCCGUUUCCGAGCAGAACCUGCCCGAAUC ACAGUGCGGCUUCCGCCCCGAUCGGUCCACCAUCGAUAUGGUGUUCACCGUGCGU CAGAUGCAGGAGAAGUGCCUGGAACAGAAUCUGAGCCUGUACAUCGUUUUCAUC GAUCUGACUAAGGCCUUCGACACGGUGAACCGCGAUGCCCUGUGGGUGAUCCUGA GCAAGCUGGGGUGCCCCGCUAAGUUCGUAAAGCUCAUCCAGUUGUUCCACGUUGA UAUGACGGGAGAGGUCCUGUCCGGUGGAGAAACCUCCGACCGGUUUAACAUCUCC AAUGGCGUGAAACAGGGCUGCGUCCUGGCCCCAGUGCUGUUCAACCUGUUUUUCA CACAGGUGCUGAGACAUGCGGUUAUGGAUCUUGACCUGGGUGUGUAUAUCAAGU AUCGCUUGGACGGCAGUCUGUUCGAUCUCAGACGCCUGACAGCUAAGACUAAGAC UACCGAGCGCCUCAUUCUCGAAGCCCUGUUUGCCGACGAUUGUGCCCUUAUGGCC CACCAGGAGAACCACCUGCAGACGAUUGUCGACCGCUUCUCUACAGCGACCAAGC UGUUCGGCCUCACCAUCUCCCUCUCAAAGACCGAAGUCCUGUUUCAGCCCGCACC CGGGAGACCCACGAAUCAGCCGUGCAUUACCAUUGACGGGACCCAACUUUCCAAC GUCAACACCUUCAAGUAUCUGGGAUCUACCAUCGCCAAUGAUGGCAGCCUGGAUC AUGAAAUCAACGCAAGGAUCCAAAAGGCUAGCCAGGCACUGGGCCGCCUCCGGUG CAAGGUCCUGCAGCACCGGGGCGUCUCCACUGCCACUAAACUGAAGGUGUACAAC GCUGUCGUGCUCUCUUCCCUGCUUUAUGGCUGCGAAACAUGGACUCUGUACCGGA AGCAUAUGAAACAGCUUGAGCAGUUCCACCAGAGGAGCUUGCGUAGCAUCAUGA GGAUCCGCUGGCAGGACCGCAUCACCAACCAGGAGGUGCUGGACAGGGCUAAUUC CACCAGCAUCGAAGUUAUGGUGCUGAAGACCCAGCUCCGUUGGUCUGGCCAUGUC Page 47 of 372 12815032v1Attorney Docket No.: 2017469-0043 AUUCGGAUGGACCCUCAGCGUAUCCCCCGCCAGGUAUUCUACGGGGAACUGUCUG CGGGGCUCAGAAAGCAGGGCCGCCCAAAGAAACGCUUUAAGGAUCAGCUGAAAA GCAAUCUGAAGUGGGCCGGCAUCACCCCCAAGCAGCUGGAGCUGGCGGCAUCUGA CCGCUCCAGCUGGCGCACCCACAUUAACCACGCAGCUACUACCUUUGAGGACGAA CGCCGGCGCAGACUGGCCGCUGCACGCGAGAGGCGCCACCAAGCCACUACCGCCC CUCCCGUGACCACUGGGGUCCCUUGCCCUAUGUGCCACAAGCUGUGUGCGUCUGC CUUCGGUUUGCAGUCCCACAUGCGUGUGCAUCGCCGCAGCGGCAGCGAGACUCCC GGAACCUCAGAGUCCGCCACACCCGAAAGUGUGAGCGGCUGGCGGCUGUUCAAGA AGAUUAGCUGAAUGGUUUAUAUUGCGGCCGCUUAAUUAAGCUGCCUUCUGCGGG GCUUGCCUUCUGGCCAUGCCCUUCUUCUCUCCCUUGCACCUGUACCUCUUGGUCU UUGAAUAAAGCCUGAGUAGGAAGUCUAGAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA (SEQ ID NO: 901).

[0314] In some embodiments, the gene modifying polypeptide is encoded by a nucleic acid sequence of: AGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACCAUGCCCA AGAAGAAGAGAAAAGUAGGAGGUUCCGAUUCCACAGCUCAUCCCAACCAGGGAC GGGGACUGGAAAAGGUUUCCCAGACUCUGCCUGCACUGCAAACCCCAGGUCAGCA UACUGCAGCUGGUGGAUCCUCUCCUCUGAGCGGAAGAAACCAGCGGAAGAACACC AAGAAGCUGCUGCUGGGAGCCUGGAACAUACGCACACUGCUGGAUCGCGAAAAU ACUCCCAGGCCUGAAAGACGCACCGCACUGAUUGGGAAAGAGCUGGCCAGAUACA ACAUCGACAUAGCCGCCCUGUCCGAAACCAGACUGCCUGAAGAAGGAUCCUUAUC CGAGCCUACCACCGGAUACACCUUCUUUUGGAAGGGGAGAGCAAGCAACGAGGAC AGGAUCCAUGGAGUGGGACUGGCUAUCAAGACCUCCCUGCUGAAGCAGCUGCCAG AUCUGCCAGUGGGAAUCUCCGAGAGACUGAUGAAGAUCAGGCUGCCUCUGUCCAA GGACCGCUACGCAACAAUCAUUUCCGCAUACGCCCCCACCCUGACUUCCACCGAG GAAACCAUUGAGCAGUUCUACUCCGAUCUGUCCGCCGUGUUGCACUCCGUGCCUA CCAACGACAAGCUGAUCCUGCUUGGAGACUUCAACGCUCGCGUGGGACAAGACCA UGAGCGCUGGAAAGGAGUGCUGGGUAAACACGGGGUGGGGAAGAUGAACAACAA CGGGCUGCUGCUGCUGUCCAAGUGCAGCGAGUUCGAGCUGACCAUUACCAACACC Page 48 of 372 12815032v1Attorney Docket No.: 2017469-0043 GUGUUUAGGAUGGCCAACAAGUACAAGACCACCUGGAUGCACCCCAGGAGCAAGC AGUGGCACCUGAUCGAUUACAUCAUCGUGCGGAGGAGGGAUAUCCAGGACGUGA AAAUCACCCGUGCAAUGCGCGGAGCUGAAUGCUGGACAGACCAUAGGCUGGUGA GAGCAACACUGCAGAUGAGGAUUGCACCUAGGCACCCUAAGAGGGCUCAAACCGU GCGGGCCUUUUACAAUGUGUCCCGGCUGAGAGAUCCCUCCUACCUGCAGACCUUU CAGUCUUGCCUGGAUGACAAGCUGUCUGCCAAGGGACCCCUGACCGGAUCUUCCA CAGAGAAGUGGAACCAGUUUCGCGACGCAGUGAAGGAAACCAGCAAAGCAGUGC UGGGACCCAAACAGCGCAACCACCAGGACUGGUUCGACGAGAACAACACCGCUAU CGAGGACCUGCUGUCCAAGAAGAACAAGGCCUUUAUGGAGUGGCAGAACAACCCC AACUCCGCCCCCAAGAAAGACAGGUUCAAAAGCCUGCAAGCCACCGCCCAGAGAG AGAUCCGGAAAAUGCAGGACAGAUGGUGGGAAAAGAAGGCCGAGGAGAUCCAGA GGUUCGCCGAUAUGAAGAACUACAAGCAGUUCUUUAGCGCACUGAAAACCGUGU ACGGACCUCUGAAACCCACCACCACUCCCCUGCUCUCUUCCGAUGGAGACACCCU GAUCAAGGACAAGAAGGGAAUCAGCAACCGCUGGAAAGAGCACUUUUCCCAGCU GCUGAACAGGCCAUCCUCCGUGGACCAAAGCGCUUUAGACCAGAUCCCCCAGAAC CGUACUAUCGAGCAGCUGGAUGUGCCUCCCUCCAUCGAAGAGGUGCAGAAGGCCA UCAAGCAGAUGUCUGCUGGGAAGGCUCCUGGGAAGGAUGGGAUCCCAACAGAGG UGUACAAAGCCCUGAACGGAAAGGCCCUGCAGGCCUUUCACAUUGUGCUGACCUC CAUCUGGGAGGAAGAGGACAUGCCUCCUGAGCUUAGGGACGCCUCUAUUGUGGCC CUGUACAAGAAUAAGGGAAGCCGCGCCGCUUGUGAUAACUACCGGGGAAUCUCAC UGCUGAGCACAGCUGGGAAGAUUCUGGCCCGUGUGAUCCUGAACCGGUUGCUGA GUUCCGUGUCCGAGCAGAACUUACCAGAAUCACAGUGCGGGUUUCGCCCCGAUCG GUCCACCAUCGAUAUGGUGUUCACAGUGAGGCAGAUGCAGGAGAAGUGCCUGGA GCAGAACCUGAGCCUGUACAUCGUGUUUAUCGACCUGACCAAGGCCUUUGAUACC GUGAACCGCGACGCACUGUGGGUUAUCCUUAGCAAGCUGGGAUGUCCCGCCAAGU UCGUGAAGCUGAUCCAGCUGUUCCACGUUGAUAUGACCGGAGAGGUGCUGUCCG GAGGAGAAACCUCCGACCGGUUUAACAUCUCCAAUGGGGUGAAACAGGGGUGUG UGCUGGCCCCAGUGCUGUUCAACCUGUUUUUCACCCAGGUGCUGAGACACGCCGU GAUGGAUCUGGACCUGGGUGUGUACAUCAAGUACCGCCUGGACGGAAGUCUGUU CGAUCUCAGACGCCUGACCGCCAAGACUAAGACUACCGAGCGCCUGAUUCUGGAA GCCCUGUUUGCCGACGAUUGUGCCCUUAUGGCCCACCAGGAGAACCACCUCCAGA Page 49 of 372 12815032v1Attorney Docket No.: 2017469-0043 CCAUUGUGGACAGAUUCUCCACCGCCACCAAGCUGUUCGGACUGACCAUCUCCCU GUCAAAGACCGAAGUGCUGUUUCAACCCGCACCUGGAAGGCCCACCAAUCAGCCU UGCAUCACUAUUGACGGGACCCAGCUGUCCAACGUGAACACCUUCAAGUACCUGG GAUCCACCAUAGCCAACGAUGGAAGCCUGGAUCACGAGAUCAACGCCAGGAUCCA GAAGGCCAGCCAAGCACUUGGGAGGCUAAGAUGUAAGGUGCUGCAGCACAGAGG GGUGAGCACUGCCACCAAGCUGAAGGUGUACAACGCUGUGGUGCUGUCCUCCCUG CUGUAUGGAUGCGAAACCUGGACCCUGUACCGGAAGCACAUGAAGCAGCUGGAGC AGUUCCACCAGAGGAGCUUGAGGAGCAUCAUGAGGAUCCGCUGGCAGGAUAGAA UCACCAACCAGGAGGUGCUGGACAGGGCUAACUCCACCAGCAUCGAGGUUAUGGU GCUGAAAACCCAGCUGCGUUGGUCUGGACACGUGAUUCGGAUGGAUCCUCAGCGU AUUCCCCGCCAGGUAUUCUAUGGGGAGCUGUCUGCCGGACUGAGAAAACAGGGCA GACCCAAGAAGCGCUUUAAGGAUCAGCUGAAGUCCAACCUGAAGUGGGCUGGGA UCACCCCUAAACAGCUGGAGUUGGCAGCAUCUGAUCGCUCCAGCUGGCGCACACA UAUCAACCACGCAGCCACCACCUUUGAGGAUGAGAGGAGGCGCAGAUUAGCUGCA GCAAGAGAGCGUCGCCAUCAAGCAACUACAGCCCCCCCAGUUACCACAGGAGUAC CUUGCCCUAUGUGCCACAAGCUAUGUGCCUCCGCCUUCGGGUUGCAAUCCCACAU GAGAGUGCAUAGGCGAAGUGGGAGCGAGACUCCUGGAACAUCAGAAUCCGCUAC ACCUGAAUCCGUGAGCGGGUGGCGGCUGUUCAAGAAGAUCAGCUGAGCUGGAGC CUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCCAGCCCCUCCUCCCCUUCC UGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGAAAAAAAAAAAAAAAAUU AAAAAAAAAAAAAAAAAAAAAAAAAUUUAAAAAAAAAAAAAAAAAAAAAAAAAA AAAAAAAAUUUUUAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA AAAA (SEQ ID NO: 391).

[0315] In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of SEQ ID NO: 403, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein one or both of: i) amino acid position 345 is other than D, e.g., is N, and ii) amino acid position 523 is other than T, e.g., is S; or a functional fragment thereof (e.g., having one or both of reverse transcriptase activity and endonuclease activity). In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of SEQ ID NO: 403, or a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto, wherein one or both of: i) amino acid position 345 is Page 50 of 372 12815032v1Attorney Docket No.: 2017469-0043 N, and ii) amino acid position 523 is S; or a functional fragment thereof (e.g., having one or both of reverse transcriptase activity and endonuclease activity).

[0316] In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of (in an N-terminal to C-terminal direction), SEQ ID NO: 400, SEQ ID NO: 402, and SEQ ID NO: 404. In some embodiments, the gene modifying polypeptide comprises an amino acid sequence of (in an N-terminal to C-terminal direction), SEQ ID NO: 403 and SEQ ID NO: 404. Page 51 of 372 12815032v1)actaagctgcadR eDt TI20gt cc agccacU 1 c ai ’3QE : aOag ttaaac aaccatg atccadn Si(Ngt c ataga teba g c agy gtt c aggag a g ctacaa ata gaggggcgcc gtt gtgcac gtgcggctaaccgat cgactt ct ga g gactt acggcg aacagagagg cgtctgttaatacctguq (M AE EKI LN W G A G W K QSQ AIR C V ne3 iysss4f0 iAni,0-d9oNm6mDeetasNFZy moE-4e1,T7nht1e r seDL R0g2so seaspa:. eco di nomsileciv epsn do tsNtoerukpqnes a artgrneomc 1AoOsMihodoeNrtDlbyaReRteTen -D1vnr]7ht:1 e E )2e me TM_ 1( 30ott1g3nilsb l Ra E1 518A0[ u T21L CLML TH K D R WEKS(M AE EKSI LN W G A G W3s4n0i,0aNF-9 m6 oE-Z4D1,7 LTR102m:. sp ile acon ditsNatgoneer omkcOsMihodoDty 1ene- DnrmET M) v2230otel R_tE1 ( 518A21O AQGS TRLE FT IQIVQVA Hs TKKRPILIMVGFL T EKGL E SusAnQYKNGEKD PGIT T LeTDNR N RQLTPHLHKFM VASER LAQLsLNDQD nSKKPKI IKPA QGMT SFM SRADFD KVL LDH NR GLSTQ AEEITL SLTEQAGGKAFo FMLPDL DTHVF E PGSQS P LAC RD DATDDLAEEL PSRLDFPLLASQ HAPD KYTNV NF IRSGRA KFRG K QDASSK GE I FIQ AWIGG R CFQ VILG LV CDAQ L MFAGMAV LITRHRQAR H K D R WRCELK3s4n0i0a-9 m6 o4D7102m:. sionNatgerkcOoDtyene1v2nrme30ot l 5tE18A21Attorney Docket No.: 2017469-0043 Polypeptide component of gene modifying system RT domain

[0318] In certain aspects, the reverse transcriptase domain of the gene modifying system is based on a reverse transcriptase domain of an APE-type or RLE-type non-LTR retrotransposon, or of a PLE-type retrotransposon. A wild-type reverse transcriptase domain of an APE-type, RLE-type, or PLE-type retrotransposon can be used in a gene modifying system or can be modified (e.g., by insertion, deletion, or substitution of one or more residues) to alter the reverse transcriptase activity for target DNA sequences. In some embodiments, the reverse transcriptase is altered from its natural sequence to have altered codon usage, e.g. improved for human cells. In some embodiments, the reverse transcriptase domain is a heterologous reverse transcriptase from a different LTR-retrotransposon, non-LTR retrotransposon, or other source.

[0319] In some embodiments, a polypeptide (e.g., RT domain) comprises an RNA-binding domain, e.g., that specifically binds to an RNA sequence. In some embodiments, a template RNA comprises an RNA sequence that is specifically bound by the RNA-binding domain. Endonuclease domain

[0320] In some embodiments, the polypeptide comprises an endonuclease domain (e.g., a heterologous endonuclease domain). In certain embodiments, the endonuclease / DNA binding domain of an APE-type retrotransposon, the endonuclease domain of an RLE-type retrotransposon, or the endonuclease domain of a PLE-type retrotransposon can be used or can be modified (e.g., by insertion, deletion, or substitution of one or more residues) in a gene modifying system described herein. In some embodiments, the endonuclease domain or endonuclease / DNA binding domain is altered from its natural sequence to have altered codon usage, e.g. improved for human cells. In some embodiments, the endonuclease element is a heterologous endonuclease element.

[0321] In some embodiments, a gene modifying polypeptide possesses the function of DNA target site cleavage via an endonuclease domain. In some embodiments, the endonuclease domain is also a DNA-binding domain. In some embodiments, the endonuclease domain is also a template nucleic acid (e.g., template RNA) binding domain. In certain embodiments, the endonuclease / DNA binding domain of an APE-type retrotransposon or the endonuclease domain Page 55 of 372 12815032v1Attorney Docket No.: 2017469-0043 of an RLE-type retrotransposon can be used or can be modified (e.g., by insertion, deletion, or substitution of one or more residues) in a gene modifying system described herein. Template nucleic acid binding domain

[0322] A gene modifying polypeptide typically contains regions capable of associating with the template nucleic acid (e.g., template RNA). In some embodiments, the template nucleic acid binding domain is an RNA binding domain. In some embodiments, the RNA binding domain is a modular domain that can associate with RNA molecules containing specific signatures, e.g., structural motifs, e.g., secondary structures present in the 3’ UTR in non-LTR retrotransposons. In other embodiments, the template nucleic acid binding domain (e.g., RNA binding domain) RNA binding domain is contained within the reverse transcription domain, e.g., the reverse transcriptase-derived component has a known signature for RNA preference, e.g., secondary structures present in the 3’ UTR in non-LTR retrotransposons. DNA binding domain

[0323] In certain aspects, the DNA-binding domain of a gene modifying polypeptide described herein is selected, designed, or constructed for binding to a desired host DNA target sequence. In certain embodiments, the DNA-binding domain of the engineered retrotransposon is a heterologous DNA-binding protein or domain relative to a native retrotransposon sequence. In still other embodiments, DNA-binding domains are modified, for example by site-specific mutation. In some embodiments, the DNA binding domain is altered from its natural sequence to have altered codon usage, e.g. improved for human cells.

[0324] In certain aspects, the host DNA-binding site integrated into by the gene modifying system can be in a gene, in an intron, in an exon, an ORF, outside of a coding region of any gene, in a regulatory region of a gene, or outside of a regulatory region of a gene. In other aspects, the retrotransposon may bind to one or more than one host DNA sequence. In other aspects, the retrotransposon may have low sequence specificity, e.g., bind to multiple sequences or lack sequence preference. Localization sequences for gene modifying systems

[0325] In certain embodiments, a gene modifying system RNA further comprises an intracellular localization sequence, e.g., a nuclear localization sequence. Page 56 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0326] The nuclear localization sequence may be an RNA sequence that promotes the import of the RNA into the nucleus. In certain embodiments, the nuclear localization signal is located on the template RNA. In certain embodiments, the retrotransposase polypeptide is encoded on a first RNA, and the template RNA is a second, separate, RNA, and the nuclear localization signal is located on the template RNA and not on an RNA encoding the retrotransposase polypeptide. While not wishing to be bound by theory, in some embodiments, the RNA encoding the retrotransposase is targeted primarily to the cytoplasm to promote its translation, while the template RNA is targeted primarily to the nucleus to promote its retrotransposition into the genome. In some embodiments, the nuclear localization signal is at the 3’ end, 5’ end, or in an internal region of the template RNA. In some embodiments the nuclear localization signal is 3’ of the heterologous sequence (e.g., is directly 3’ of the heterologous sequence) or is 5’ of the heterologous sequence (e.g., is directly 5’ of the heterologous sequence). In some embodiments, the nuclear localization signal is placed outside of the 5’ UTR or outside of the 3’ UTR of the template RNA. In some embodiments the nuclear localization signal is placed between the 5’ UTR and the 3’ UTR, wherein optionally the nuclear localization signal is not transcribed with the transgene (e.g., the nuclear localization signal is an anti-sense orientation or is downstream of a transcriptional termination signal or polyadenylation signal). In some embodiments, the nuclear localization sequence is situated inside of an intron. In some embodiments a plurality of the same or different nuclear localization signals are in the RNA, e.g., in the template RNA. In some embodiments, the nuclear localization signal is less than 5, 10, 25, 50, 75, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, or 1000 bp in length. Various RNA nuclear localization sequences can be used. For example, Lubelsky and Ulitsky, Nature 555 (107-111), 2018 describe RNA sequences, which drive RNA localization into the nucleus. In some embodiments, the nuclear localization signal is a SINE-derived nuclear RNA localization (SIRLOIN) signal. In some embodiments, the nuclear localization signal binds a nuclear- enriched protein. In some embodiments, the nuclear localization signal binds the HNRNPK protein. In some embodiments the nuclear localization signal is rich in pyrimidines, e.g., is a C / T rich, C / U rich, C rich, T rich, or U rich region. In some embodiments, the nuclear localization signal is derived from a long non-coding RNA. In some embodiments, the nuclear localization signal is derived from MALAT1 long non-coding RNA or is the 600 nucleotide M region of MALAT1 (described in Miyagawa et al., RNA 18, (738-751), 2012). In some embodiments, the Page 57 of 372 12815032v1Attorney Docket No.: 2017469-0043 nuclear localization signal is derived from BORG long non-coding RNA or is a AGCCC motif (described in Zhang et al., Molecular and Cellular Biology 34, 2318-2329 (2014). In some embodiments, the nuclear localization sequence is described in Shukla et al., The EMBO Journal e98452 (2018). In some embodiments, the nuclear localization signal is derived from a non-LTR retrotransposon, an LTR retrotransposon, retrovirus, or an endogenous retrovirus.

[0327] In some embodiments, a polypeptide described herein comprises one or more (e.g., 2, 3, 4, 5) nuclear targeting sequences, for example, a nuclear localization sequence (NLS), e.g., as described above. In some embodiments, the NLS is a bipartite NLS. In some embodiments, an NLS facilitates the import of a protein comprising an NLS into the cell nucleus. In some embodiments, the NLS is fused to the N-terminus of a gene modifying polypeptide described herein. In some embodiments, the NLS is fused to the C-terminus of the gene modifying polypeptide. In some embodiments, a linker sequence is disposed between the NLS and the neighboring domain of the gene modifying polypeptide.

[0328] In some embodiments, the NLS comprises the amino acid sequence MPKKKRKVGGS (SEQ ID NO: 400), or a functional fragment or variant thereof, wherein optionally the NLS is situated N-terminal of the RT and EN domains. In some embodiments, the NLS comprises the amino acid sequence PKKKRKV (SEQ ID NO: 378), or a functional fragment or variant thereof, wherein optionally the NLS is situated N-terminal of the RT and EN domains. In some embodiments, a gene modifying polypeptide described herein comprises an NLS with the amino acid sequence PKKKRKV (SEQ ID NO: 378) and a GGS linker, wherein optionally the NLS is situated N-terminal of RT and EN domains. In some embodiments, the NLS comprises an amino acid sequence encoded by the nucleic acid sequence CCAAAGAAGAAAAGAAAAGUA, or a functional fragment or variant thereof, wherein optionally the NLS is situated N-terminal of the RT and EN domains.

[0329] In some embodiments, a gene modifying polypeptide described herein comprises, from N terminus to C terminus, the first residue of a wild-type gene modifying polypeptide (e.g., M, the first residue of SEQ ID NO: 402 herein), an NLS sequence, a linker, and a wild-type gene modifying polypeptide (e.g., residues 2-1110 of SEQ ID NO: 402).

[0330] In some embodiments, a gene modifying polypeptide described herein comprises an XTEN linker, e.g., an XTEN linker comprising the amino acid sequence SGSETPGTSESATPES Page 58 of 372 12815032v1Attorney Docket No.: 2017469-0043 (SEQ ID NO: 2200), or a functional fragment or variant thereof. In some embodiments, the gene modifying polypeptide comprises (e.g., C-terminal of the RT and EN domains) the amino acid sequence SGSETPGTSESATPESVSGWRLFKKIS (SEQ ID NO: 404), which comprises an XTEN linker and a HiBiT tag. In some embodiments, an NLS comprises the amino acid sequence MDSLLMNRRKFLYQFKNVRWAKGRRETYLC (SEQ ID NO: 366), PKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 367), PKKKRKVPKKRKVEGADKRTADGSEFESPKKKRKV (SEQ ID NO: 373), RKSGKIAAIWKRPRKPKKKRKV (SEQ ID NO: 368) KRTADGSEFESPKKKRKV(SEQ ID NO: 369), KKTELQTTNAENKTKKL (SEQ ID NO: 370), or KRGINDRNFWRGENGRKTR (SEQ ID NO: 371), KRPAATKKAGQAKKKK (SEQ ID NO: 372), PAAKRVKLD (SEQ ID NO: 374), KRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 375), KRTADGSEFE (SEQ ID NO: 376), KRTADGSEFESPKKKAKVE (SEQ ID NO: 377), AGKRTADGSEFEKRTADGSEFESPKKKAKVE (SEQ ID NO: 441), or a functional fragment or variant thereof.

[0331] In some aspects, the present disclosure comprises a polypeptide comprising an NLS comprises the amino acid sequence MPKKKRKVGGS (SEQ ID NO: 400), or a functional fragment or variant thereof. In some embodiments, the functional fragment or variant comprises 1, 2, 3, or 4 sequence differences relative to SEQ ID NO: 400. In some embodiments, the polypeptide is an RTE-1 polypeptide. In some embodiments, the RTE-1 polypeptide comprises an amino acid sequence of Table 1 or a sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto. In some embodiments, the NLS is N-terminal of the amino acid sequence of Table 1 or the sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity thereto.

[0332] In some embodiments, a gene modifying polypeptide comprises an NLS as comprised in SEQ ID NO: 400 and / or SEQ ID NO: 441, or an NLS having an amino acid sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto.

[0333] In some embodiments, a gene modifying polypeptide comprises an NLS as listed in Table 8 of PCT Publication No. WO / 2021 / 178717, incorporated herein by reference in its entirety.

[0334] In some embodiments, the NLS is a bipartite NLS. A bipartite NLS typically comprises two basic amino acid clusters separated by a spacer sequence (which may be, e.g., about 10 Page 59 of 372 12815032v1Attorney Docket No.: 2017469-0043 amino acids in length). A monopartite NLS typically lacks a spacer. In certain embodiments, a gene modifying system polypeptide further comprises an intracellular localization sequence, e.g., a nuclear localization sequence and / or a nucleolar localization sequence. The nuclear localization sequence and / or nucleolar localization sequence may be amino acid sequences that promote the import of the protein into the nucleus and / or nucleolus, where it can promote integration of heterologous sequence into the genome. In certain embodiments, a gene modifying system polypeptide (e.g., a retrotransposase, e.g., a polypeptide according to Table 1 herein) further comprises a nucleolar localization sequence. In some embodiments, a nucleic acid described herein (e.g., an RNA encoding a gene modifying polypeptide, or a DNA encoding the RNA) comprises a microRNA binding site. In some embodiments, the microRNA binding site is used to increase the target-cell specificity of a gene modifying system. For instance, the microRNA binding site can be chosen on the basis that is recognized by a miRNA that is present in a non- target cell type, but that is not present (or is present at a reduced level relative to the non-target cell) in a target cell type. Thus, when the RNA encoding the gene modifying polypeptide is present in a non-target cell, it would be bound by the miRNA, and when the RNA encoding the gene modifying polypeptide is present in a target cell, it would not be bound by the miRNA (or bound but at reduced levels relative to the non-target cell). While not wishing to be bound by theory, binding of the miRNA to the RNA encoding the gene modifying polypeptide may reduce production of the gene modifying polypeptide, e.g., by degrading the mRNA encoding the polypeptide or by interfering with translation. Accordingly, the heterologous object sequence would be inserted into the genome of target cells more efficiently than into the genome of non- target cells. A system having a microRNA binding site in the RNA encoding the gene modifying polypeptide (or encoded in the DNA encoding the RNA) may also be used in combination with a template RNA that is regulated by a second microRNA binding site, e.g., as described herein in the section entitled “Template RNA component of gene modifying system.” Linkers

[0335] In some embodiments, domains of the compositions and systems described herein (e.g., the endonuclease and reverse transcriptase domains of a polypeptide or the DNA binding domain and reverse transcriptase domains of a polypeptide) may be joined by a linker. A composition described herein comprising a linker element has the general form S1-L-S2, wherein S1 and S2 Page 60 of 372 12815032v1Attorney Docket No.: 2017469-0043 may be the same or different and represent two domain moieties (e.g., each a polypeptide or nucleic acid domain) associated with one another by the linker. In some embodiments, a linker may connect two polypeptides. In some embodiments, a linker may connect two nucleic acid molecules. In some embodiments, a linker may connect a polypeptide and a nucleic acid molecule. A linker may be a chemical bond, e.g., one or more covalent bonds or non-covalent bonds. A linker may be flexible, rigid, and / or cleavable. In some embodiments, the linker is a peptide linker. Generally, a peptide linker is at least 2, 3, 4, 5, 6, 7, 8, 9, 10 or more amino acids in length, e.g., 2-50 amino acids in length, 2-30 amino acids in length.

[0336] Some commonly used flexible linkers have sequences consisting primarily of stretches of Gly and Ser residues (“GS” linker). Flexible linkers may be useful for joining domains that require a certain degree of movement or interaction and may include small, non-polar (e.g. Gly) or polar (e.g. Ser or Thr) amino acids. Incorporation of Ser or Thr can also maintain the stability of the linker in aqueous solutions by forming hydrogen bonds with the water molecules, and therefore reduce unfavorable interactions between the linker and the other moieties. Examples of such linkers include those having the structure [GGS]>1or [GGGS]>1. Rigid linkers are useful to keep a fixed distance between domains and to maintain their independent functions. Rigid linkers may also be useful when a spatial separation of the domains is critical to preserve the stability or bioactivity of one or more components in the agent. Rigid linkers may have an alpha helix-structure or Pro-rich sequence, (XP)n, with X designating any amino acid, preferably Ala, Lys, or Glu. Cleavable linkers may release free functional domains in vivo. In some embodiments, linkers may be cleaved under specific conditions, such as the presence of reducing reagents or proteases. In vivo cleavable linkers may utilize the reversible nature of a disulfide bond. One example includes a thrombin-sensitive sequence (e.g., PRS) between the two Cys residues. In vitro thrombin treatment of CPRSC results in the cleavage of the thrombin- sensitive sequence, while the reversible disulfide linkage remains intact. Such linkers are known and described, e.g., in Chen et al.2013. Fusion Protein Linkers: Property, Design and Functionality. Adv Drug Deliv Rev.65(10): 1357–1369. In vivo cleavage of linkers in compositions described herein may also be carried out by proteases that are expressed in vivo under pathological conditions (e.g. cancer or inflammation), in specific cells or tissues, or constrained within certain cellular compartments. The specificity of many proteases offers slower cleavage of the linker in constrained compartments. Page 61 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0337] In some embodiments the amino acid linkers are (or are homologous to) the endogenous amino acids that exist between such domains in a native polypeptide. In some embodiments, the endogenous amino acids that exist between such domains are substituted but the length is unchanged from the natural length. In some embodiments, additional amino acid residues are added to the naturally existing amino acid residues between domains.

[0338] In some embodiments, the amino acid linkers are designed computationally or screened to maximize protein function (Anad et al., FEBS Letters, 587:19, 2013).

[0339] In addition to being fully encoded on a single transcript, a polypeptide can be generated by separately expressing two or more polypeptide fragments that reconstitute the holoenzyme. In some embodiments, the gene modifying polypeptide is generated by expressing as separate subunits that reassemble the holoenzyme through engineered protein-protein interactions. In some embodiments, reconstitution of the holoenzyme does not involve covalent binding between subunits. Peptides may also fuse together through trans-splicing of inteins (Tornabene et al. Sci Transl Med 11, eaav4523 (2019)). In some embodiments, the gene modifying holoenzyme is expressed as separate subunits that are designed to create a fusion protein through the presence of split inteins (e.g., as described herein) in the subunits. In some embodiments, the gene modifying holoenzyme is reconstituted through the formation of covalent linkages between subunits. In some embodiments, the breaking up of a gene modifying polypeptide into subunits may aid in delivery of the protein by keeping the nucleic acid encoding each part within optimal packaging limits of a viral delivery vector, e.g., AAV (Tornabene et al. Sci Transl Med 11, eaav4523 (2019)). In some embodiments, the gene modifying polypeptide is designed to be dimerized through the use of covalent or non-covalent interactions as described above. Exemplary Linkers are shown in Table L1 below. Table L1: Exemplary linker sequences SEQ IDPage 62 of 372 12815032v1Attorney Docket No.: 2017469-0043 Amino Acid SequenceSEQ IDNOPage 63 of 372 12815032v1Attorney Docket No.: 2017469-0043 Amino Acid SequenceSEQ IDNOPage 64 of 372 12815032v1Attorney Docket No.: 2017469-0043 Amino Acid SequenceSEQ IDNOPage 65 of 372 12815032v1Attorney Docket No.: 2017469-0043 Amino Acid SequenceSEQ IDNOchosen from: (SGGS)n(SEQ ID NO: 1234), (GGGS)n(SEQ ID NO: 1149), (GGGGS)n(SEQ ID NO: 1111), (G)n, (EAAAK)n(SEQ ID NO: 1113), (GGS)n, or (XP)n. Promoters

[0341] In some embodiments, one or more promoter or enhancer elements are operably linked to a nucleic acid encoding a gene modifying protein or a template nucleic acid, e.g., that controls expression of the heterologous object sequence. In certain embodiments, the one or more promoter or enhancer elements comprise cell-type or tissue specific elements. In some embodiments, the promoter or enhancer is the same or derived from the promoter or enhancer that naturally controls expression of the heterologous object sequence.

[0342] Exemplary tissue specific promoters that are commercially available can be found, for example, at a uniform resource locator (e.g., invivogen.com / tissue-specific-promoters). In some embodiments, a promoter is a native promoter or a minimal promoter, e.g., which consists of a single fragment from the 5’ region of a given gene. In some embodiments, a native promoter comprises a core promoter and its natural 5’ UTR. In some embodiments, the 5’ UTR comprises an intron. In other embodiments, these include composite promoters, which combine promoter elements of different origins or were generated by assembling a distal enhancer with a minimal promoter of the same origin.

[0343] Exemplary cell or tissue specific promoters are provided, for example, in PCT Publication No. WO / 2021 / 178717, incorporated herein by reference in its entirety. Exemplary nucleic acid sequences encoding such promoters are known in the art and can be readily accessed using a variety of resources, such as the NCBI database, including RefSeq, as well as the Eukaryotic Promoter Database (epd.epfl.ch / / index.php). Page 66 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0344] Depending on the host / vector system utilized, any of a number of suitable transcription and translation control elements, including constitutive and inducible promoters, transcription enhancer elements, transcription terminators, etc. may be used in the expression vector (see e.g., Bitter et al. (1987) Methods in Enzymology, 153:516-544; incorporated herein by reference in its entirety).

[0345] In some embodiments, a nucleic acid encoding a gene modifying polypeptide or template nucleic acid is operably linked to a control element, e.g., a transcriptional control element, such as a promoter. The transcriptional control element may, in some embodiments, be functional in either a eukaryotic cell, e.g., a mammalian cell; or a prokaryotic cell (e.g., bacterial or archaeal cell). In some embodiments, a nucleotide sequence encoding a polypeptide is operably linked to multiple control elements, e.g., that allow expression of the nucleotide sequence encoding the polypeptide in both prokaryotic and eukaryotic cells. Insertions produced by systems described herein

[0346] In some embodiments, a gene modifying system is capable of producing an insertion into the target site of at least 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides). In some embodiments, a gene modifying system is capable of producing an insertion into the target site of at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 nucleotides (and optionally no more than 500, 400, 300, 200, or 100 nucleotides). In some embodiments, a gene modifying system is capable of producing an insertion into the target site of at least 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 3.5, 4, 4.5, 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, 9.5 or 10 kilobases (and optionally no more than 1, 5, 10, or 20 kilobases).

[0347] In some embodiments, the gene modifying polypeptide results in insertion of the heterologous object sequence (e.g., the GFP gene) at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome. In some embodiments, a cell described herein (e.g., a cell comprising a heterologous sequence) comprises the heterologous object sequence at an average copy number of at least 0.01, 0.025, 0.05, 0.075, 0.1, 0.15, 0.2, 0.25, 0.3, 0.4, 0.5, 0.75, 1, 1.25, 1.5, 1.75, 2, 2.5, 3, 4, or 5 copies per genome. Page 67 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0348] In some embodiments, a gene modifying system causes integration of a sequence in a target DNA with relatively few truncation events at the terminus. For instance, in some embodiments, a gene modifying protein results in about 25-100%, 50-100%, 60-100%, 70-100%, 75-95%, 80%-90%, or 86.17% of integrants into the target site being non-truncated, as measured by an assay described herein, e.g., an assay of Example 6 and Figure 8 of PCT Application No. PCT / US2019 / 048607, incorporated herein by reference in its entirety. In some embodiments, a gene modifying protein results in at least about 30%, 40%, 50%, 60%, 70%, 80%, or 90% of integrants into the target site being non-truncated, as measured by an assay described herein. In some embodiments, the number of full-length integrants in the target insertion site is greater than the number of truncated integrants, e.g., the number of full-length integrants is at least 1.1x, 1.2x, 1.5x, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, or 10x the number of the truncated integrants, or the number of full-length integrants is at least 1.1x-10x, 2x-10x, 3x-10x, or 5x-10x the number of the truncated integrants.

[0349] In some embodiments, a system or method described herein results in “scarless” insertion of the heterologous object sequence, while in some embodiments, the target site can show deletions or duplications of endogenous DNA as a result of insertion of the heterologous sequence. The mechanisms of different retrotransposons could result in different patterns of duplications or deletions in the host genome occurring during retrotransposition at the target site. In some embodiments, the system results in a scarless insertion, with no duplications or deletions in the surrounding genomic DNA. In some embodiments, the system results in a deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system results in a deletion of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion. In some embodiments, the system results in a duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA upstream of the insertion. In some embodiments, the system results in a duplication of less than 1, 2, 3, 4, 5, 10, 50, or 100 bp of genomic DNA downstream of the insertion. Template RNA component of gene modifying system

[0350] The gene modifying systems described herein can transcribe an RNA sequence template into host target DNA sites by target-primed reverse transcription. In some embodiments, the Page 68 of 372 12815032v1Attorney Docket No.: 2017469-0043 template comprises a modified RTE-15’ UTR and / or a modified RTE-13’ UTR as described herein.

[0351] In some embodiments, a template RNA comprises at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, but no more than 10, 15, or 20 nucleotide differences (e.g., insertions, substitutions, or deletions) relative to a 5’ UTR and / or 3’ UTR of a template sequence as described herein (e.g., comprising the nucleic acid sequence of a RTE-15’ UTR, e.g., comprising the nucleic acid sequence of SEQ ID NO: 1).

[0352] The present disclosure provides, in some aspects, nucleic acid molecule comprising a nucleotide sequence having at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, but no more than 10, 15, or 20 sequence modifications (e.g., insertion, substitutions, or deletions) relative to SEQ ID NO: 1.

[0353] The present disclosure provides, in some aspects, nucleic acid molecule comprising a nucleotide sequence having at least 1, 2, 3, 4, 5, 6, 7, 8, or 9, but no more than 10, 15, or 20 sequence modifications (e.g., insertion, substitutions, or deletions) relative to SEQ ID NO: 2.

[0354] It is understood that, when a template RNA is described as comprising an open reading frame or the reverse complement thereof, in some embodiments the template RNA must be converted into double stranded DNA (e.g., through reverse transcription) before the open reading frame can be transcribed and translated.

[0355] The template nucleic acid (e.g., template RNA) component of a gene modifying system described herein typically is able to bind the gene modifying protein of the system. In some embodiments, the template RNA has a 3’ region that is capable of binding a gene modifying protein. The binding region, e.g., 3’ region, may be a structured RNA region, e.g., having at least 1, 2 or 3 hairpin loops, capable of binding the gene modifying protein of the system. The binding region may associate the template nucleic acid (e.g., template RNA) with any of the polypeptide modules. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with an RNA-binding domain in the polypeptide. In some embodiments, the binding region of the template nucleic acid (e.g., template RNA) may associate with the reverse transcription domain of the polypeptide (e.g., specifically bind to the RT domain). For example, where the reverse transcription domain is derived from a non-LTR retrotransposon, the template nucleic acid (e.g., template RNA) may contain a binding region derived from a non- LTR retrotransposon, e.g., a 3’ UTR from a non-LTR retrotransposon. In some embodiments a Page 69 of 372 12815032v1Attorney Docket No.: 2017469-0043 system or method described herein comprises a single template nucleic acid (e.g., template RNA). In some embodiments a system or method described herein comprises a plurality of template nucleic acids (e.g., template RNAs). In some embodiments, when the system comprises a plurality of nucleic acids, each nucleic acid comprises a conjugating domain. In some embodiments, a conjugating domain enables association of nucleic acid molecules, e.g., by hybridization of complementary sequences.

[0356] In some embodiments, the template nucleic acid may comprise one or more UTRs (e.g., a 5’ UTR or a 3’ UTR). In some embodiments, the UTR facilitates interaction of the template with the reverse transcriptase domain of the polypeptide. In some embodiments, the template possesses one or more sequences aiding in association of the template with the gene modifying polypeptide. In some embodiments, these sequences may be derived from retrotransposon UTRs. In some embodiments, the UTRs may be located flanking the desired insertion sequence. In some embodiments, a sequence with target site homology may be located outside of one or both UTRs. In some embodiments, the sequence with target site homology can anneal to the target sequence to prime reverse transcription. In some embodiments, the 5’ and / or 3’ UTR may be located terminal to the target site homology sequence. In some embodiments, the gene modifying system may result in the insertion of a desired payload without any additional sequence (e.g., a gene expression unit without UTRs used to bind the gene modifying protein).

[0357] The template RNA component of a gene modifying system described herein typically is able to bind the gene modifying protein of the system. In some embodiments, the template RNA has a 5’ region that is capable of binding a gene modifying protein. The binding region, e.g., 5’ region, may be a structured RNA region, e.g., having at least 1, 2 or 3 hairpin loops, capable of binding the gene modifying protein of the system. In some embodiments, the 5’ untranslated region comprises a pseudoknot, e.g., a pseudoknot that is capable of binding to the gene modifying protein.

[0358] In some embodiments, the template RNA (e.g., an untranslated region of the hairpin RNA, e.g., a 5’ untranslated region) comprises a stem-loop sequence. In some embodiments, the template RNA (e.g., an untranslated region of the hairpin RNA, e.g., a 5’ untranslated region) comprises a hairpin. In some embodiments, the template RNA (e.g., an untranslated region of the hairpin RNA, e.g., a 5’ untranslated region) comprises a helix. In some embodiments, the Page 70 of 372 12815032v1Attorney Docket No.: 2017469-0043 template RNA (e.g., an untranslated region of the hairpin RNA, e.g., a 5’ untranslated region) comprises a pseudoknot. In some embodiments, the template RNA comprises a ribozyme. In some embodiments the ribozyme is similar to a hepatitis delta virus (HDV) ribozyme, e.g., has a secondary structure like that of the HDV ribozyme and / or has one or more activities of the HDV ribozyme, e.g., a self-cleavage activity. See, e.g., Eickbush et al., Molecular and Cellular Biology, 2010, 3142-3150.

[0359] In some embodiments, the template RNA (e.g., an untranslated region of the hairpin RNA, e.g., a 3’ untranslated region) comprises one or more stem-loops or helices. Exemplary structures of R23’ UTRs are shown, for example, in Ruschak et al. “Secondary structure models of the 3′ untranslated regions of diverse R2 RNAs” RNA.2004 Jun; 10(6): 978–987, e.g., at Figure 3, therein, and in Eikbush and Eikbush, “R2 and R2 / R1 hybrid non-autonomous retrotransposons derived by internal deletions of full-length elements” Mobile DNA (2012) 3:10; e.g., at Figure 3 therein, which articles are hereby incorporated by reference in their entirety.

[0360] In some embodiments, a template RNA described herein comprises a sequence that is capable of binding to a gene modifying protein described herein. For instance, in some embodiments, the template RNA comprises an MS2 RNA sequence capable of binding to an MS2 coat protein sequence in the gene modifying protein. In some embodiments, the template RNA comprises an RNA sequence capable of binding to a B-box sequence. In some embodiments, in addition to or in place of a UTR, the template RNA is linked (e.g., covalently) to a non-RNA UTR, e.g., a protein or small molecule.

[0361] In some embodiments, the template RNA has a tail at the 3’ end. In some embodiments, the template RNA has a poly-A tail at the 3’ end. In some embodiments, the template RNA does not have a poly-A tail at the 3’ end.

[0362] In some embodiments the template RNA has a 5’ region of at least 10, 15, 20, 25, 30, 40, 50, 60, 80, 100, 120, 140, 160, 180, 200 or more bases of at least 40%, 50%, 60%, 70%, 80%, 90%, 95% or greater homology to the 5’ sequence of a retrotransposon, e.g., a retrotransposon described herein.

[0363] The template RNA of the system typically comprises an object sequence for insertion into a target DNA. The object sequence may be coding or non-coding. Page 71 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0364] In some embodiments, a system or method described herein comprises a single template RNA. In some embodiments, a system or method described herein comprises a plurality of template RNAs.

[0365] In some embodiments, the object sequence may contain an open reading frame. In some embodiments, the template RNA has a Kozak sequence. In some embodiments, the template RNA has an internal ribosome entry site. In some embodiments, the template RNA has a self- cleaving peptide such as a T2A or P2A site. In some embodiments, the template RNA has a start codon. In some embodiments, the template RNA has a splice acceptor site. In some embodiments, the template RNA has a splice donor site. Exemplary splice acceptor and splice donor sites are described in WO2016044416, incorporated herein by reference in its entirety. Exemplary splice acceptor site sequences are known to those of skill in the art and include, by way of example only, CTGACCCTTCTCTCTCTCCCCCAGAG (from human HBB gene) and TTTCTCTCCCACAAG (from human immunoglobulin-gamma gene). In some embodiments the template RNA, has a microRNA binding site downstream of the stop codon. In some embodiments, the template RNA has a polyA tail downstream of the stop codon of an open reading frame. In some embodiments, the template RNA comprises one or more exons. In some embodiments, the template RNA comprises one or more introns. In some embodiments, the template RNA comprises a eukaryotic transcriptional terminator. In some embodiments, the template RNA comprises an enhanced translation element or a translation enhancing element. In some embodiments, the RNA comprises the human T-cell leukemia virus (HTLV-1) R region. In some embodiments, the RNA comprises a posttranscriptional regulatory element that enhances nuclear export, such as that of Hepatitis B Virus (HPRE) or Woodchuck Hepatitis Virus (WPRE). In some embodiments, in the template RNA, the heterologous object sequence encodes a polypeptide and is coded in an antisense direction with respect to the 5’ and 3’ UTR. In some embodiments, in the template RNA, the heterologous object sequence encodes a polypeptide and is coded in a sense direction with respect to the 5’ and 3’ UTR.

[0366] In some embodiments, the object sequence may contain a non-coding sequence. For example, the template RNA may comprise a promoter or enhancer sequence. In some embodiments, the template RNA comprises a tissue specific promoter or enhancer, each of which may be unidirectional or bidirectional. A system having a tissue-specific promoter sequence in Page 72 of 372 12815032v1Attorney Docket No.: 2017469-0043 the template RNA may also be used in combination with a DNA encoding a gene modifying polypeptide, driven by a tissue-specific promoter, e.g., to achieve higher levels of gene modifying protein in target cells than in non-target cells.

[0367] In some embodiments, the non-coding sequence is transcribed in an antisense-direction with respect to the 5’ and 3’ UTR. In some embodiments, the non-coding sequence is transcribed in a sense direction with respect to the 5’ and 3’ UTR.

[0368] In some embodiments, the template RNA comprises a non-coding heterologous object sequence, e.g., a regulatory sequence. In some embodiments, integration of the heterologous object sequence thus alters the expression of an endogenous gene. In some embodiments, integration of the heterologous object sequence upregulates expression of an endogenous gene. In some embodiments, integration of the heterologous object sequence downregulated expression of an endogenous gene.

[0369] In some embodiments, the template RNA comprises a site that coordinates epigenetic modification. In some embodiments, the template RNA comprises an element that inhibits, e.g., prevents, epigenetic silencing. In some embodiments, the template RNA comprises a chromatin insulator. For example, the template RNA comprises a CTCF site or a site targeted for DNA methylation.

[0370] In order to promote higher level or more stable gene expression, the template RNA may include features that prevent or inhibit gene silencing. In some embodiments, these features prevent or inhibit DNA methylation. In some embodiments, these features promote DNA demethylation. In some embodiments, these features prevent or inhibit histone deacetylation. In some embodiments, these features prevent or inhibit histone methylation. In some embodiments, these features promote histone acetylation. In some embodiments, these features promote histone demethylation. In some embodiments, multiple features may be incorporated into the template RNA to promote one or more of these modifications. CpG dinculeotides are subject to methylation by host methyl transferases. In some embodiments, the template RNA is depleted of CpG dinucleotides, e.g., does not comprise CpG nucleotides or comprises a reduced number of CpG dinucleotides compared to a corresponding unaltered sequence. In some embodiments, the promoter driving transgene expression from integrated DNA is depleted of CpG dinucleotides. Page 73 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0371] In some embodiments, the template RNA comprises a gene expression unit composed of at least one regulatory region operably linked to an effector sequence. The effector sequence may be a sequence that is transcribed into RNA (e.g., a coding sequence or a non-coding sequence such as a sequence encoding a micro RNA).

[0372] In some embodiments, the object sequence of the template RNA is inserted into a target genome in an endogenous intron. In some embodiments, the object sequence of the template RNA is inserted into a target genome and thereby acts as a new exon. In some embodiments, the insertion of the object sequence into the target genome results in replacement of a natural exon or the skipping of a natural exon. Methods and Compositions for Modified RNA (e.g., template RNA)

[0373] In some embodiments, an RNA component of the system (e.g., a template RNA, as described herein) comprises one or more nucleotide modifications. In some embodiments, the modification pattern of the template RNA can significantly affect in vivo activity compared to unmodified or end-modified guides. Without wishing to be bound by theory, this process may be due, at least in part, to a stabilization of the RNA conferred by the modifications. Non-limiting examples of such modifications may include 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxyethyl) (2'- O-MOE), 2'- fluoro (2'-F), phosphorothioate (PS) bond between nucleotides, G-C substitutions, and inverted abasic linkages between nucleotides and equivalents thereof.

[0374] In some embodiments, the template RNA (e.g., at the portion thereof that binds a target site) comprises a 5' terminus region. In some embodiments, the template RNA does not comprise a 5' terminus region. In some embodiments, the 5' terminus region comprises a 5' end modification. In some embodiments, the template RNA comprises a 2'-O-methyl (2'-O-Me) modified nucleotide. In some embodiments, the template RNA comprises a 2'-O-(2-methoxy ethyl) (2'-O-moe) modified nucleotide. In some embodiments, the template RNA comprises a 2'- fluoro (2'- F) modified nucleotide. In some embodiments, the template RNA comprises a phosphorothioate (PS) bond between nucleotides. In some embodiments, the template RNA comprises a 5' end modification, a 3' end modification, or 5' and 3' end modifications. In some embodiments, the 5' end modification comprises a phosphorothioate (PS) bond between nucleotides. In some embodiments, the 5' end modification comprises a 2'-O-methyl (2'-O-Me), 2'-O-(2-methoxy ethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modified nucleotide. In some Page 74 of 372 12815032v1Attorney Docket No.: 2017469-0043 embodiments, the 5' end modification comprises at least one phosphorothioate (PS) bond and one or more of a 2'-O-methyl (2'-O- Me), 2'-O-(2-methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modified nucleotide. The end modification may comprise a phosphorothioate (PS), 2'-O- methyl (2'-O-Me) , 2'-O-(2- methoxyethyl) (2'-O-MOE), and / or 2'-fluoro (2'-F) modification. Equivalent end modifications are also encompassed by embodiments described herein. In some embodiments, the template RNA comprises an end modification in combination with a modification of one or more regions of the template RNA. In some embodiments, structure- guided and systematic approaches are used to introduce modifications (e.g., 2′-OMe-RNA, 2′-F- RNA, and PS modifications) to a template RNA, for example, as described in Mir et al. Nat Commun 9:2641 (2018) (incorporated by reference herein in its entirety). In some embodiments, the incorporation of 2′-F-RNAs increases thermal and nuclease stability of RNA:RNA or RNA:DNA duplexes, e.g., while minimally interfering with C3′-endo sugar puckering. In some embodiments, 2′-F may be better tolerated than 2′-OMe at positions where the 2′-OH is important for RNA:DNA duplex stability. In some embodiments, structure-guided and systematic approaches (e.g., as described in Mir et al. Nat Commun 9:2641 (2018); incorporated herein by reference in its entirety) are employed to find modifications for the template RNA. In some embodiments, a structure of polypeptide bound to template RNA is used to determine non- protein-contacted nucleotides of the RNA that may then be selected for modifications, e.g., with lower risk of disrupting the association of the RNA with the polypeptide. Secondary structures in a template RNA can also be predicted in silico by software tools, e.g., the RNAstructure tool available at rna.urmc.rochester.edu / RNAstructureWeb (Bellaousov et al. Nucleic Acids Res 41:W471-W474 (2013); incorporated by reference herein in its entirety), e.g., to determine secondary structures for selecting modifications, e.g., hairpins, stems, and / or bulges.

[0375] Further included here are compositions and methods for the assembly of full or partial template RNA molecules. In some embodiments, RNA molecules may be assembled by the connection of two or more (e.g., two, three, four, five, six, seven, eight, nine, ten, or more) RNA segments with each other. In an aspect, the disclosure provides methods for producing nucleic acid molecules, the methods comprising contacting two or more linear RNA segments with each other under conditions that allow for the 5′ terminus of a first RNA segment to be covalently linked with the 3′ terminus of a second RNA segment. In some embodiments, the joined molecule may be contacted with a third RNA segment under conditions that allow for the 5’ Page 75 of 372 12815032v1Attorney Docket No.: 2017469-0043 terminus of the joined molecule to be covalently linked with the 3’ terminus of the third RNA segment. In embodiments, the method further comprises joining a fourth, fifth, or additional RNA segments to the elongated molecule. This form of assembly may, in some instances, allow for rapid and efficient assembly of RNA molecules.

[0376] In some embodiments, RNA segments may be produced by chemical synthesis. In some embodiments, RNA segments may be produced by in vitro transcription of a nucleic acid template, e.g., by providing an RNA polymerase to act on a cognate promoter of a DNA template to produce an RNA transcript. In some embodiments, in vitro transcription is performed using, e.g., a T7, T3, or SP6 RNA polymerase, or a derivative thereof, acting on a DNA, e.g., dsDNA, ssDNA, linear DNA, plasmid DNA, linear DNA amplicon, linearized plasmid DNA, e.g., encoding the RNA segment, e.g., under transcriptional control of a cognate promoter, e.g., a T7, T3, or SP6 promoter. In some embodiments, a combination of chemical synthesis and in vitro transcription is used to generate the RNA segments for assembly. In embodiments, the gene modifying polypeptide binding segments are produced by chemical synthesis and the heterologous object sequence segment is produced by in vitro transcription. Without wishing to be bound by theory, in vitro transcription may be better suited for the production of longer RNA molecules. In some embodiments, reaction temperature for in vitro transcription may be lowered, e.g., be less than 37°C (e.g., between 0-10C, 10-20C, or 20-30C), to result in a higher proportion of full-length transcripts (Krieg Nucleic Acids Res 18:6463 (1990)). In some embodiments, a protocol for improved synthesis of long transcripts is employed to synthesize a long template RNA, e.g., a template RNA greater than 5 kb, such as the use of e.g., T7 RiboMAX Express, which can generate 27 kb transcripts in vitro (Thiel et al. J Gen Virol 82(6):1273-1281 (2001)). In some embodiments, modifications to RNA molecules as described herein may be incorporated during synthesis of RNA segments (e.g., through the inclusion of modified nucleotides or alternative binding chemistries), following synthesis of RNA segments through chemical or enzymatic processes, following assembly of one or more RNA segments, or a combination thereof. Additional Template Features

[0377] In some embodiments, the template (e.g., template RNA) comprises certain structural features, e.g., determined in silico. In embodiments, the template RNA is predicted to have Page 76 of 372 12815032v1Attorney Docket No.: 2017469-0043 minimal energy structures between -280 and -480 kcal / mol (e.g., between -280 to -300, -300 to - 350, -350 to -400, -400 to -450, or -450 to -480 kcal / mol), e.g., as measured by RNA structure, e.g., as described in Turner and Mathews Nucleic Acids Res 38:D280-282 (2009) (incorporated herein by reference in its entirety).

[0378] In some embodiments, the template (e.g., template RNA) comprises certain structural features, e.g., determined in vitro. In embodiments, the template RNA is sequence optimized, e.g., to reduce secondary structure as determined in vitro, for example, by SHAPE-MaP (e.g., as described in Siegfried et al. Nat Methods 11:959-965 (2014); incorporated herein by reference in its entirety). In some embodiments, the template (e.g., template RNA) comprises certain structural features, e.g., determined in cells. In embodiments, the template RNA is sequence optimized, e.g., to reduce secondary structure as measured in cells, for example, by DMS- MaPseq (e.g., as described in Zubradt et al. Nat Methods 14:75-82 (2017); incorporated by reference herein in its entirety). Chemically modified nucleic acids and nucleic acid end features

[0379] A nucleic acid described herein (e.g., a template nucleic acid, e.g., a template RNA; or a nucleic acid (e.g., mRNA) encoding a gene modifying polypeptide) can comprise unmodified or modified nucleobases. Naturally occurring RNAs are synthesized from four basic ribonucleotides: ATP, CTP, UTP and GTP, but may contain post-transcriptionally modified nucleotides. Further, approximately one hundred different nucleoside modifications have been identified in RNA (Rozenski, J, Crain, P, and McCloskey, J. (1999). The RNA Modification Database: 1999 update. Nucl Acids Res 27: 196-197). An RNA can also comprise wholly synthetic nucleotides that do not occur in nature.

[0380] In some embodiments, the chemically modification is one provided in PCT / US2016 / 032454, US Pat. Pub. No.20090286852, of International Application No. WO / 2012 / 019168, WO / 2012 / 045075, WO / 2012 / 135805, WO / 2012 / 158736, WO / 2013 / 039857, WO / 2013 / 039861, WO / 2013 / 052523, WO / 2013 / 090648, WO / 2013 / 096709, WO / 2013 / 101690, WO / 2013 / 106496, WO / 2013 / 130161, WO / 2013 / 151669, WO / 2013 / 151736, WO / 2013 / 151672, WO / 2013 / 151664, WO / 2013 / 151665, WO / 2013 / 151668, WO / 2013 / 151671, WO / 2013 / 151667, WO / 2013 / 151670, WO / 2013 / 151666, WO / 2013 / 151663, WO / 2014 / 028429, WO / 2014 / 081507, WO / 2014 / 093924, WO / 2014 / 093574, WO / 2014 / 113089, WO / 2014 / 144711, WO / 2014 / 144767, Page 77 of 372 12815032v1Attorney Docket No.: 2017469-0043 WO / 2014 / 144039, WO / 2014 / 152540, WO / 2014 / 152030, WO / 2014 / 152031, WO / 2014 / 152027, WO / 2014 / 152211, WO / 2014 / 158795, WO / 2014 / 159813, WO / 2014 / 164253, WO / 2015 / 006747, WO / 2015 / 034928, WO / 2015 / 034925, WO / 2015 / 038892, WO / 2015 / 048744, WO / 2015 / 051214, WO / 2015 / 051173, WO / 2015 / 051169, WO / 2015 / 058069, WO / 2015 / 085318, WO / 2015 / 089511, WO / 2015 / 105926, WO / 2015 / 164674, WO / 2015 / 196130, WO / 2015 / 196128, WO / 2015 / 196118, WO / 2016 / 011226, WO / 2016 / 011222, WO / 2016 / 011306, WO / 2016 / 014846, WO / 2016 / 022914, WO / 2016 / 036902, WO / 2016 / 077125, or WO / 2016 / 077123, each of which is herein incorporated by reference in its entirety. It is understood that incorporation of a chemically modified nucleotide into a polynucleotide can result in the modification being incorporated into a nucleobase, the backbone, or both, depending on the location of the modification in the nucleotide. In some embodiments, the backbone modification is one provided in EP 2813570, which is herein incorporated by reference in its entirety. In some embodiments, the modified cap is one provided in US Pat. Pub. No.20050287539, which is herein incorporated by reference in its entirety.

[0381] In some embodiments, the chemically modified nucleic acid (e.g., RNA, e.g., mRNA) comprises one or more of ARCA: anti-reverse cap analog (m27.3'-OGP3G), GP3G (Unmethylated Cap Analog), m7GP3G (Monomethylated Cap Analog), m32.2.7GP3G (Trimethylated Cap Analog), m5CTP (5'-methyl-cytidine triphosphate), m6ATP (N6-methyl- adenosine-5'-triphosphate), s2UTP (2-thio-uridine triphosphate), and Ѱ (pseudouridine triphosphate).

[0382] In some embodiments, the chemically modified nucleic acid comprises a 5’ cap, e.g.: a 7- methylguanosine cap (e.g., a O-Me-m7G cap); a hypermethylated cap analog; an NAD+-derived cap analog (e.g., as described in Kiledjian, Trends in Cell Biology 28, 454-464 (2018)); or a modified, e.g., biotinylated, cap analog (e.g., as described in Bednarek et al., Phil Trans R Soc B 373, 20180167 (2018)).

[0383] In some embodiments, the chemically modified nucleic acid comprises a 3’ feature selected from one or more of: a polyA tail; a 16-nucleotide long stem-loop structure flanked by unpaired 5 nucleotides (e.g., as described by Mannironi et al., Nucleic Acid Research 17, 9113- 9126 (1989)); a triple-helical structure (e.g., as described by Brown et al., PNAS 109, 19202- 19207 (2012)); a tRNA, Y RNA, or vault RNA structure (e.g., as described by Labno et al., Page 78 of 372 12815032v1Attorney Docket No.: 2017469-0043 Biochemica et Biophysica Acta 1863, 3125-3147 (2016)); incorporation of one or more deoxyribonucleotide triphosphates (dNTPs), 2’O-Methylated NTPs, or phosphorothioate-NTPs; a single nucleotide chemical modification (e.g., oxidation of the 3’ terminal ribose to a reactive aldehyde followed by conjugation of the aldehyde-reactive modified nucleotide); or chemical ligation to another nucleic acid molecule.

[0384] In some embodiments, the nucleic acid (e.g., template nucleic acid or nucleic acid encoding the gene modifying polypeptide) comprises one or more modified nucleotides, e.g., selected from dihydrouridine, inosine, 7-methylguanosine, 5-methylcytidine (5mC), 5′ Phosphate ribothymidine, 2′-O-methyl ribothymidine, 2′-O-ethyl ribothymidine, 2′-fluoro ribothymidine, C- 5 propynyl-deoxycytidine (pdC), C-5 propynyl-deoxyuridine (pdU), C-5 propynyl-cytidine (pC), C-5 propynyl-uridine (pU), 5-methyl cytidine, 5-methyl uridine, 5-methyl deoxycytidine, 5- methyl deoxyuridine methoxy, 2,6-diaminopurine, 5′-Dimethoxytrityl-N4-ethyl-2′- deoxycytidine, C-5 propynyl-f-cytidine (pfC), C-5 propynyl-f-uridine (pfU), 5-methyl f-cytidine, 5-methyl f-uridine, C-5 propynyl-m-cytidine (pmC), C-5 propynyl-f-uridine (pmU), 5-methyl m- cytidine, 5-methyl m-uridine, LNA (locked nucleic acid), MGB (minor groove binder) pseudouridine (Ψ), 1-N-methylpseudouridine (1-Me-Ψ), or 5-methoxyuridine (5-MO-U).

[0385] In some embodiments, the nucleic acid comprises a backbone modification, e.g., a modification to a sugar or phosphate group in the backbone. In some embodiments, the nucleic acid comprises a nucleobase modification.

[0386] In some embodiments, the nucleic acid comprises one or more chemically modified nucleotides of Table 2, one or more chemical backbone modifications of Table 3, one or more chemically modified caps of Table 4. For instance, in some embodiments, the nucleic acid comprises two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of chemical modifications. As an example, the nucleic acid may comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of modified nucleobases, e.g., as described herein, e.g., in Table 2. Alternatively or in combination, the nucleic acid may comprise two or more (e.g., 3, 4, 5, 6, 7, 8, 9, or 10 or more) different types of backbone modifications, e.g., as described herein, e.g., in Table 3. Alternatively or in combination, the nucleic acid may comprise one or more modified cap, e.g., as described herein, e.g., in Table 4. For instance, in some embodiments, the nucleic acid comprises one or more type of modified nucleobase and one or more type of backbone Page 79 of 372 12815032v1Attorney Docket No.: 2017469-0043 modification; one or more type of modified nucleobase and one or more modified cap; one or more type of modified cap and one or more type of backbone modification; or one or more type of modified nucleobase, one or more type of backbone modification, and one or more type of modified cap.

[0387] In some embodiments, the nucleic acid comprises one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) modified nucleobases. In some embodiments, all nucleobases of the nucleic acid are modified. In some embodiments, the nucleic acid is modified at one or more (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, or more) positions in the backbone. In some embodiments, all backbone positions of the nucleic acid are modified. Table 2: Modified nucleotides 5-aza-uridine N2-methyl-6-thio-guanosinePage 80 of 372 12815032v1Attorney Docket No.: 2017469-0043 4-methoxy-pseudouridine 2-methylthio-N6-(cis- 4-methoxy-2-thio-pseudouridine hydroxyisopentenyl)adenosine rPage 81 of 372 12815032v1Attorney Docket No.: 2017469-0043 2-methylthio-adenine 5-carboxymethylaminomethyluridine 2-methoxy-adenine 5-carboxymethylaminomethyl-2'-O-Table 3: Backbone modifications 2’-O-Methyl backbonePage 82 of 372 12815032v1Attorney Docket No.: 2017469-0043 Table 4: Modified caps m7GpppA m7G CProduction of Compositions and Systems

[0388] Methods of designing and constructing nucleic acid constructs and proteins or polypeptides (such as the systems, constructs, and polypeptides described herein) are known. Generally, recombinant methods may be used.

[0389] The disclosure provides, in part, a nucleic acid, e.g., vector, encoding a gene modifying polypeptide described herein, a template nucleic acid described herein, or both. In some embodiments, a vector comprises a selective marker, e.g., an antibiotic resistance marker. In some embodiments, a vector encoding a gene modifying polypeptide is integrated into a target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). In some embodiments, a vector encoding a gene modifying polypeptide is not integrated into a target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). In some embodiments, a vector encoding a template nucleic acid (e.g., template RNA) is not integrated into a target cell genome (e.g., upon administration to a target cell, tissue, organ, or subject). Page 83 of 372 12815032v1Attorney Docket No.: 2017469-0043 Exemplary methods for producing a therapeutic pharmaceutical protein or polypeptide described herein involve expression in mammalian cells, although recombinant proteins can also be produced using insect cells, yeast, bacteria, or other cells under control of appropriate promoters. Mammalian expression vectors may comprise non-transcribed elements such as an origin of replication, a suitable promoter, and other 5' or 3' flanking non-transcribed sequences, and 5' or 3' non-translated sequences such as necessary ribosome binding sites, a polyadenylation site, splice donor and acceptor sites, and termination sequences. DNA sequences derived from the SV40 viral genome, for example, SV40 origin, early promoter, splice, and polyadenylation sites may be used to provide other genetic elements required for expression of a heterologous DNA sequence. Appropriate cloning and expression vectors for use with bacterial, fungal, yeast, and mammalian cellular hosts are described in Green & Sambrook, Molecular Cloning: A Laboratory Manual (Fourth Edition), Cold Spring Harbor Laboratory Press (2012).

[0390] Various mammalian cell culture systems can be employed to express and manufacture recombinant protein. Examples of mammalian expression systems include CHO, COS, HEK293, HeLA, and BHK cell lines. Processes of host cell culture for production of protein therapeutics are described in Zhou and Kantardjieff (Eds.), Mammalian Cell Cultures for Biologics Manufacturing (Advances in Biochemical Engineering / Biotechnology), Springer (2014). Compositions described herein may include a vector, such as a viral vector, e.g., a lentiviral vector, encoding a recombinant protein. In some embodiments, a vector, e.g., a viral vector, may comprise a nucleic acid encoding a recombinant protein.

[0391] In some embodiments, quality standards include, but are not limited to: (i) the length of mRNA encoding the gene modifying polypeptide, e.g., whether the mRNA has a length that is above a reference length or within a reference length range, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA present is greater than 3000, 4000, or 5000 nucleotides long; (ii) the presence, absence, and / or length of a polyA tail on the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA present contains a polyA tail (e.g., a polyA tail that is at least 5, 10, 20, 30, 50, 70, 100 nucleotides in length); Page 84 of 372 12815032v1Attorney Docket No.: 2017469-0043 (iii) the presence, absence, and / or type of a 5’ cap on the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA present contains a 5’ cap, e.g., whether that cap is a 7-methylguanosine cap, e.g., a O-Me-m7G cap; (iv) the presence, absence, and / or type of one or more modified nucleotides (e.g., selected from pseudouridine, dihydrouridine, inosine, 7-methylguanosine, 1-N-methylpseudouridine (1- Me-Ψ), 5-methoxyuridine (5-MO-U), 5-methylcytidine (5mC), or a locked nucleotide) in the mRNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA present contains one or more modified nucleotides; (v) the stability of the mRNA (e.g., over time and / or under a pre-selected condition), e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the mRNA remains intact (e.g., greater than 100, 125, 150, 175, or 200 nucleotides long) after a stability test; or (vi) the potency of the mRNA in a system for modifying DNA, e.g., whether at least 1% of target sites are modified after a system comprising the mRNA is assayed for potency. Kits, Articles of Manufacture, and Pharmaceutical Compositions

[0392] In an aspect the disclosure provides a kit comprising a gene modifying polypeptide or a gene modifying system, e.g., as described herein. In some embodiments, the kit comprises a gene modifying polypeptide (or a nucleic acid encoding the polypeptide) and a template RNA (or DNA encoding the template RNA). In some embodiments, the kit further comprises a reagent for introducing the system into a cell, e.g., transfection reagent, LNP, and the like. In some embodiments, the kit is suitable for any of the methods described herein. In some embodiments, the kit comprises one or more elements, compositions (e.g., pharmaceutical compositions), gene modifying polypeptides, and / or gene modifying systems, or a functional fragment or component thereof, e.g., disposed in an article of manufacture. In some embodiments, the kit comprises instructions for use thereof.

[0393] In an aspect, the disclosure provides an article of manufacture, e.g., in which a kit as described herein, or a component thereof, is disposed.

[0394] In an aspect, the disclosure provides a pharmaceutical composition comprising a gene modifying polypeptide or a gene modifying system, e.g., as described herein. In some Page 85 of 372 12815032v1Attorney Docket No.: 2017469-0043 embodiments, the pharmaceutical composition further comprises a pharmaceutically acceptable carrier or excipient. In some embodiments, the pharmaceutical composition comprises a template RNA and / or an RNA encoding the polypeptide. In embodiments, the pharmaceutical composition has one or more (e.g., 1, 2, 3, or 4) of the following characteristics: (a) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) DNA template relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (b) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) uncapped RNA relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (c) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) partial length RNAs relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (d) substantially lacks unreacted cap dinucleotides. Chemistry, Manufacturing, and Controls (CMC)

[0395] Purification of protein therapeutics is described, for example, in Franks, Protein Biotechnology: Isolation, Characterization, and Stabilization, Humana Press (2013); and in Cutler, Protein Purification Protocols (Methods in Molecular Biology), Humana Press (2010).

[0396] In some embodiments, a gene modifying system, polypeptide, and / or template nucleic acid (e.g., template RNA) conforms to certain quality standards. In some embodiments, a gene modifying system, polypeptide, and / or template nucleic acid (e.g., template RNA) produced by a method described herein conforms to certain quality standards. Accordingly, the disclosure is directed, in some aspects, to methods of manufacturing a gene modifying system, polypeptide, and / or template nucleic acid (e.g., template RNA) that conforms to certain quality standards, e.g., in which said quality standards are assayed. The disclosure is also directed, in some aspects, to methods of assaying said quality standards in a gene modifying system, polypeptide, and / or template nucleic acid (e.g., template RNA). In some embodiments, quality standards include, but are not limited to, one or more (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12) of the following: (i) the length of the template RNA, e.g., whether the template RNA has a length that is above a reference length or within a reference length range, e.g., whether at least 80, 85, 90, 95, Page 86 of 372 12815032v1Attorney Docket No.: 2017469-0043 96, 97, 98, or 99% of the template RNA present is greater than 100, 125, 150, 175, or 200 nucleotides long; (ii) the presence, absence, and / or length of a polyA tail on the template RNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the template RNA present contains a polyA tail (e.g., a polyA tail that is at least 5, 10, 20, 30, 50, 70, 100 nucleotides in length); (iii) the presence, absence, and / or type of a 5’ cap on the template RNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the template RNA present contains a 5’ cap, e.g., whether that cap is a 7-methylguanosine cap, e.g., a O-Me-m7G cap; (iv) the presence, absence, and / or type of one or more modified nucleotides (e.g., selected from pseudouridine, dihydrouridine, inosine, 7-methylguanosine, 1-N-methylpseudouridine (1- Me-Ψ), 5-methoxyuridine (5-MO-U), 5-methylcytidine (5mC), or a locked nucleotide) in the template RNA, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the template RNA present contains one or more modified nucleotides; (v) the stability of the template RNA (e.g., over time and / or under a pre-selected condition), e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the template RNA remains intact (e.g., greater than 100, 125, 150, 175, or 200 nucleotides long) after a stability test; (vi) the potency of the template RNA in a system for modifying DNA, e.g., whether at least 1% of target sites are modified after a system comprising the template RNA is assayed for potency; (vii) the length of the polypeptide, first polypeptide, or second polypeptide, e.g., whether the polypeptide, first polypeptide, or second polypeptide has a length that is above a reference length or within a reference length range, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide present is greater than 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1600, 1700, 1800, 1900, or 2000 amino acids long (and optionally, no larger than 2500, 2000, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, or 600 amino acids long); (viii) the presence, absence, and / or type of post-translational modification on the polypeptide, first polypeptide, or second polypeptide, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide contains phosphorylation, Page 87 of 372 12815032v1Attorney Docket No.: 2017469-0043 methylation, acetylation, myristoylation, palmitoylation, isoprenylation, glipyatyon, or lipoylation, or any combination thereof; (ix) the presence, absence, and / or type of one or more artificial, synthetic, or non- canonical amino acids (e.g., selected from ornithine, β-alanine, GABA, δ-Aminolevulinic acid, PABA, a D-amino acid (e.g., D-alanine or D-glutamate), aminoisobutyric acid, dehydroalanine, cystathionine, lanthionine, Djenkolic acid, Diaminopimelic acid, Homoalanine, Norvaline, Norleucine, Homonorleucine, homoserine, O-methyl-homoserine and O-ethyl-homoserine, ethionine, selenocysteine, selenohomocysteine, selenomethionine, selenoethionine, tellurocysteine, or telluromethionine) in the polypeptide, first polypeptide, or second polypeptide, e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide present contains one or more artificial, synthetic, or non- canonical amino acids; (x) the stability of the polypeptide, first polypeptide, or second polypeptide (e.g., over time and / or under a pre-selected condition), e.g., whether at least 80, 85, 90, 95, 96, 97, 98, or 99% of the polypeptide, first polypeptide, or second polypeptide remains intact (e.g., greater than 600, 650, 700, 750, 800, 850, 900, 950, 1000, 1050, 1100, 1150, 1200, 1250, 1300, 1350, 1400, 1450, 1500, 1600, 1700, 1800, 1900, or 2000 amino acids long (and optionally, no larger than 2500, 2000, 1500, 1400, 1300, 1200, 1100, 1000, 900, 800, 700, or 600 amino acids long)) after a stability test; (xi) the potency of the polypeptide, first polypeptide, or second polypeptide in a system for modifying DNA, e.g., whether at least 1 % of target sites are modified after a system comprising the polypeptide, first polypeptide, or second polypeptide is assayed for potency; or (xii) the presence, absence, and / or level of one or more of a pyrogen, virus, fungus, bacterial pathogen, or host cell protein, e.g., whether the system is free or substantially free of pyrogen, virus, fungus, bacterial pathogen, or host cell protein contamination.

[0397] In some embodiments, a system or pharmaceutical composition described herein is endotoxin free.

[0398] In some embodiments, the presence, absence, and / or level of one or more of a pyrogen, virus, fungus, bacterial pathogen, and / or host cell protein is determined. In embodiments, Page 88 of 372 12815032v1Attorney Docket No.: 2017469-0043 whether the system is free or substantially free of pyrogen, virus, fungus, bacterial pathogen, and / or host cell protein contamination is determined.

[0399] In some embodiments, a pharmaceutical composition or system as described herein has one or more (e.g., 1, 2, 3, or 4) of the following characteristics: (a) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) DNA template relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (b) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) uncapped RNA relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (c) less than 1% (e.g., less than 0.5%, 0.4%, 0.3%, 0.2%, or 0.1%) partial length RNAs relative to the template RNA and / or the RNA encoding the polypeptide, e.g., on a molar basis; (d) substantially lacks unreacted cap dinucleotides. Applications

[0400] By integrating coding genes into a RNA sequence template, the gene modifying system can address therapeutic needs, for example, by providing expression of a therapeutic transgene in individuals with loss-of-function mutations, by replacing gain-of-function mutations with normal transgenes, by providing regulatory sequences to eliminate gain-of-function mutation expression, and / or by controlling the expression of operably linked genes, transgenes and systems thereof. In certain embodiments, the RNA sequence template encodes a promotor region specific to the therapeutic needs of the host cell, for example a tissue specific promotor or enhancer. In still other embodiments, a promotor can be operably linked to a coding sequence.

[0401] In embodiments, the gene modifying system can provide therapeutic transgenes expressing, e.g., replacement blood factors or replacement enzymes, e.g., lysosomal enzymes. For example, the compositions, systems and methods described herein are useful to express, in a target human genome, agalsidase alpha or beta for treatment of Fabry Disease; imiglucerase, taliglucerase alfa, velaglucerase alfa, or alglucerase for Gaucher Disease; sebelipase alpha for lysosomal acid lipase deficiency (Wolman disease / CESD); laronidase, idursulfase, elosulfase alpha, or galsulfase for mucopolysaccharidoses; alglucosidase alpha for Pompe disease. For Page 89 of 372 12815032v1Attorney Docket No.: 2017469-0043 example, the compositions, systems, and methods described herein are useful to express, in a target human genome factor I, II, V, VII, X, XI, XII or XIII for blood factor deficiencies.

[0402] In some embodiments, the heterologous object sequence encodes an intracellular protein (e.g., a cytoplasmic protein, a nuclear protein, an organellar protein such as a mitochondrial protein or lysosomal protein, or a membrane protein). In some embodiments, the heterologous object sequence encodes a membrane protein, e.g., a CAR or a membrane protein other than a CAR, and / or an endogenous human membrane protein. In some embodiments, the heterologous object sequence encodes an extracellular protein. In some embodiments, the heterologous object sequence encodes an enzyme, a structural protein, a signaling protein, a regulatory protein, a transport protein, a sensory protein, a motor protein, a defense protein, or a storage protein. Other exemplary proteins that may be encoded by a heterologous object sequence include, without limitation, an immune receptor protein, e.g. a synthetic immune receptor protein such as a chimeric antigen receptor protein (CAR), a T cell receptor, a B cell receptor, or an antibody.

[0403] In some embodiments, the systems, reaction mixtures, and cell populations modified using the methods disclosed herein can be used to treat a subject in need thereof. In some embodiments, the subject has a cancer, e.g., a hematological cancer or a solid tumor. In some embodiments, the subject has an infectious disease. In other embodiments, the subject has an autoimmune or an inflammatory disease. Chimeric Antigen Receptors

[0404] In some embodiments, the heterologous object sequence encodes a chimeric antigen receptor (CAR) comprising an antigen binding domain. In some embodiments, the CAR is or comprises a first generation CAR comprising an antigen binding domain, a transmembrane domain, and a single intracellular signaling domain. In some embodiments, the CAR is or comprises a second generation CAR comprising an antigen binding domain, a transmembrane domain, and two intracellular signaling domains (e.g., a first intracellular signaling domain and a second intracellular signaling domain). In some embodiments, the CAR comprises a third generation CAR comprising an antigen binding domain, a transmembrane domain, and at least three intracellular signaling domains. In some embodiments, a fourth generation CAR comprising an antigen binding domain, a transmembrane domain, three or four intracellular signaling domains, and a domain which upon successful signaling of the CAR induces Page 90 of 372 12815032v1Attorney Docket No.: 2017469-0043 expression of a cytokine gene. In some embodiments, the antigen binding domain is or comprises an scFv, Fab, a diabody, a D domain binder, centyrins (e.g., antibody-like scaffolds, e.g., a CARTyrin), one or more single domain antibodies such as VHH domains (e.g., comprises two VHH binding domains). In some embodiments, the CAR antigen binding domain binds to two epitopes of the target antigen (e.g., is a biepitopic binding domain). In some embodiments, the CAR comprises two antigen binding domains, such that each antigen binding domain binds to a different target antigen on a cell, e.g., a neoplastic cell. Antigen Binding Domains

[0405] In some embodiments, a CAR antigen binding domain is or comprises an antibody or antigen-binding portion thereof. In some embodiments, a CAR antigen binding domain is or comprises an scFv, Fab, a diabody, a D domain binder, centyrins (e.g., antibody-like scaffolds, e.g., a CARTyrin), one or more single domain antibodies such as VHH domains (e.g., comprises two VHH binding domains). In some embodiments, the CAR antigen binding domain binds to two epitopes of the target antigen (e.g., is a biepitopic binding domain). In some embodiments, the CAR comprises two antigen binding domains, such that each antigen binding domain binds to a different target antigen on a cell, e.g., a neoplastic cell.

[0406] In some embodiments, the CAR comprises a camelid antigen-binding domain. In some embodiments, the CAR comprises a murine binding domain. In some embodiments, the CAR comprises a humanized binding domain. In some embodiments, the CAR comprises a human binding domain.

[0407] In some embodiments, an antigen binding domain binds to a cell surface antigen of a cell. In some embodiments, a cell surface antigen is characteristic of one type of cell. In some embodiments, a cell surface antigen is characteristic of more than one type of cell.

[0408] In some embodiments, the antigen binding domain targets an antigen characteristic of a neoplastic cell. In some embodiments, the antigen characteristic of a neoplastic cell is selected from a receptor listed in Table 5, or an antigenic fragment or antigenic portion thereof. In some embodiments, the antigen binding domain binds one or more antigens of a blood cancer (e.g., a leukemia, a lymphoma, or a multiple myeloma). In some embodiments, the blood cancer antigen is a B cell antigen. In some embodiments, the antigen is BCMA. In some embodiments, the Page 91 of 372 12815032v1Attorney Docket No.: 2017469-0043 antigen is GPRC5D. In some embodiments, the antigen is CD20. In some embodiments, the antigen binding domain binds an antigen of a solid tumor. Table 5: Exemplary Neoplastic Cell Antigens prostate ifi 1Page 92 of 372 12815032v1Attorney Docket No.: 2017469-0043 histidine kinase T-cell alpha VEGFR SPage 93 of 372 12815032v1Attorney Docket No.: 2017469-0043 VEGF-D CD25 ABL IGF-1receptorRAGE-1Page 94 of 372 12815032v1Attorney Docket No.: 2017469-0043 CXCR1CD137 (4-1BB) ANPEP GPRC5D LY75Page 95 of 372 12815032v1Attorney Docket No.: 2017469-0043 CIC-Kb Th22 CLL-1 LAGE-la CD151 CD340 [040, g g g g c of a T-cell. In some embodiments, the antigen characteristic of a T-cell is selected from an exemplary T cell antigen listed in Table 6, or an antigenic fragment thereof. Table 6: Exemplary T-cell Antigens a cell surface receptor CALM1 HLA-DRB1 MAP2K6 (MKK6) NFKBIA RAF1 APage 96 of 372 12815032v1Attorney Docket No.: 2017469-0043 an ion channel protein CD3G(CD3 )HLA-DRB5 MAP3K3 PAK2 SHP26 1A6 0Page 97 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0410] In some embodiments, the antigen binding domain targets an antigen characteristic of an autoimmune or inflammatory disorder. In some embodiments, the autoimmune or inflammatory disorder is selected from chronic graft-vs-host disease (GVHD), lupus, arthritis, immune complex glomerulonephritis, goodpasture, uveitis, hepatitis, systemic sclerosis or scleroderma, type I diabetes, multiple sclerosis, cold agglutinin disease, Pemphigus vulgaris, Grave's disease, autoimmune hemolytic anemia, Hemophilia A, Primary Sjogren's Syndrome, thrombotic thrombocytopenia purrpura, neuromyelits optica, Evan's syndrome, IgM mediated neuropathy, cyroglobulinemia, dermatomyositis, idiopathic thrombocytopenia, ankylosing spondylitis, bullous pemphigoid, acquired angioedema, chronic urticarial, antiphospholipid demyelinating polyneuropathy, and autoimmune thrombocytopenia or neutropenia or pure red cell aplasias, while exemplary non-limiting examples of alloimmune diseases include allosensitization (see, for example, Blazar et al., 2015, Am. J. Transplant, 15(4):931-41) or xenosensitization from hematopoietic or solid organ transplantation, blood transfusions, pregnancy with fetal allosensitization, neonatal alloimmune thrombocytopenia, hemolytic disease of the newborn, sensitization to foreign antigens such as can occur with replacement of inherited or acquired deficiency disorders treated with enzyme or protein replacement therapy, blood products, and gene therapy. In some embodiments, the antigen characteristic of an autoimmune or inflammatory disorder is selected from an exemplary antigen listed in Table 7, or an antigenic fragment thereof. In some embodiments, the antigen binding domain targets citrullinated vimentin (e.g., associated with rheumatoid arthritis). In some embodiments, the antigen binding domain targets a human leukocyte antigen (HLA) (e.g., to induce transplant tolerance). Table 7: Exemplary Autoimmune or Inflammatory Disorder Antigens a cell surface receptor a G protein-coupled receptor-like tyrosine histidine kinase nPage 98 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0411] In some embodiments, a CAR antigen binding domain binds to a ligand expressed on B cells, plasma cells, plasmablasts. In some embodiments, the antigen expressed on B cells, plasma cells, or plasmablasts is selected from an exemplary antigen listed in Table 8, or an antigenic fragment thereof. In some embodiments, the B cell antigen is BCMA. In some embodiments, the B cell antigen is GPRC5D. In some embodiments, the B cell antigen is CD20. In some embodiments, a CAR that binds to an antigen listed in Table 8 is utilized to deplete B cells (e.g., autoreactive B cells producing autoantibodies) to induce immune tolerance. Table 8: Exemplary B cell, Plasma Cell, and Plasmablast Antigens CD10 CD24 CD138 TNF LFA-1

[0412] In so, acteristic of an infectious disease. In some embodiments, wherein the infectious disease is selected from HIV, hepatitis B virus, hepatitis C virus, Human herpes virus, Human herpes virus 8 (HHV-8, Kaposi sarcoma-associated herpes virus (KSHV)), Human T-lymphotrophic virus-1 (HTLV-1), Merkel cell polyomavirus (MCV), Simian virus 40 (SV40), Eptstein-Barr virus, CMV, human papillomavirus. In some embodiments, the antigen characteristic of an infectious disease is selected from an exemplary antigen listed in Table 9, or an antigenic fragment thereof. Table 9: Exemplary Infectious Disease Antigens a cell surface receptor receptor serine / threonine kinase rPage 99 of 372 12815032v1Attorney Docket No.: 2017469-0043 tyrosine kinase associated receptor CD4-induced epitope on HIV-1 Env.

[0413] Oses a signal peptide. An amino acid sequence of an exemplary signal peptide is MALPVTALLLPLALLLHAARP (SEQ ID NO: 15542), which may be encoded by an exemplary nucleic acid sequence of ATGGCTCTGCCGGTGACCGCCCTGCTTCTGCCTCTTGCCCTGCTCTTGCATGCCGCTC GCCCG (SEQ ID NO: 15543) or ATGGCTCTGCCCGTCACCGCTCTGCTGCTGCCTCTGGCTCTGCTGCTGCACGCTGCTC GCCCT (SEQ ID NO: 15544). Transmembrane Domain

[0414] In some embodiments, the CAR transmembrane domain comprises at least a transmembrane region of an exemplary transmembrane domain listed in Table 10, or a functional fragment thereof. Table 10: Exemplary Transmembrane Domains alpha chain of the T cellCD224-TCR CD804Page 100 of 372 12815032v1Attorney Docket No.: 2017469-0043 CD9 CD8α CD3γCD22[041ain. In some embodiments, the CD8 transmembrane domain has an amino acid sequence of a CD8 transmembrane domain listed in Table 11, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the CD8 transmembrane domain is encoded by a nucleic acid sequence of a CD8 transmembrane domain listed in Table 11, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Table 11: Sequences of Exemplary Transmembrane Domains Name Sequence SEQ IDIntracellular Signaling Domains

[0416] In some embodiments, the CAR comprises at least one signaling domain selected from one or more intracellular signaling domains listed in Table 12, or a functional fragment thereof. In some embodiments, the CAR comprises a first intracellular signaling domain and a second Page 101 of 372 12815032v1Attorney Docket No.: 2017469-0043 intracellular signaling domain. In some embodiments, the first intracellular signaling domain mediates downstream signaling during T-cell activation. In some embodiments, the second intracellular signaling domain is a costimulatory domain. Table 12: Exemplary Intracellular Signaling Domains B7-1 / CD80 4-1BB Ligand / TNFSF9 RELT / TNFRSF19L CD7 aaPage 102 of 372 12815032v1Attorney Docket No.: 2017469-0043 CRTAM CD30TIM-1 / KIM-1 / HAVCR CD2immunoreceptor tyrosine-based activation motif (ITAM), or functional variant thereof; (ii) a CD28 domain or functional variant thereof; and (iii) a 4-1BB domain, or a CD134 domain, or functional variant thereof. In some embodiments, the CAR comprises a CD3 zeta domain or an immunoreceptor tyrosine-based activation motif (ITAM), or functional variant thereof. In some embodiments, the CAR comprises (i) a CD3 zeta domain, or an immunoreceptor tyrosine-based activation motif (ITAM), or functional variant thereof; (ii) a CD28 domain, or a 4-1BB domain, or functional variant thereof, and / or (iii) a 4-1BB domain, or a CD134 domain, or functional variant thereof. In some embodiments, the CAR comprises a (i) a CD3 zeta domain, or an immunoreceptor tyrosine-based activation motif (ITAM), or functional variant thereof; (ii) a CD28 domain or functional variant thereof; (iii) a 4-1BB domain, or a CD134 domain, or functional variant thereof; and (iv) a cytokine or costimulatory ligand transgene.

[0418] In some embodiments, the CAR comprises a CD28 co-stimulatory domain.

[0419] In some embodiments, the CAR comprises a CD3z signaling domain.

[0420] In some embodiments, intracellular signaling domain comprises an intracellular signaling domain listed in Table 13, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, intracellular signaling domain is encoded Page 103 of 372 12815032v1Attorney Docket No.: 2017469-0043 by a nucleic acid sequence of an intracellular signaling domain listed in Table 13, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Table 13: Sequences of Exemplary Intracellular Signaling Domains Name Sequence SEQ ID NOPage 104 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0421] In some embodiments, the CAR further comprises one or more spacers, e.g., wherein the spacer is a first spacer between the antigen binding domain and the transmembrane domain. In some embodiments, the first spacer includes at least a portion of an immunoglobulin constant region or variant or modified version thereof. In some embodiments, the spacer is a second spacer between the transmembrane domain and a signaling domain. In some embodiments, the second spacer is an oligopeptide, e.g., wherein the oligopeptide comprises glycine-serine doublets. In some embodiments, the CAR further comprises a hinge domain. In some embodiments, the hinge domain is a CD8 hinge domain. In some embodiments, the CD8 hinge domain has an amino acid sequence of a CD8 hinge domain in Table 14, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the CD8 hinge domain is encoded by a nucleic sequence listed in Table 14, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Table 14: Sequences of Exemplary Hinge Domains Name Sequence SEQ ID

[0422] In some embodiments, the CAR comprises a sequence of a CAR listed in Table 15, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. In some embodiments, the CAR is encoded by a nucleic acid sequence listed in Table 15, or sequence having at least 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity thereto. Page 105 of 372 12815032v1Attorney Docket No.: 2017469-0043 Table 15: Sequences of Exemplary CAR Molecules BCMA CAR 1 Descri tion Se uence SEQ OPage 106 of 372 12815032v1Attorney Docket No.: 2017469-0043 GATTTTGCAACTTACTACTGTCAGCAAAAATACGACCTCCT CACTTTTGGCGGAGGGACCAAGGTTGAGATCAAAGGCAGCPage 107 of 372 12815032v1Attorney Docket No.: 2017469-0043 TTGGTATCAGCAGAAACCAGGGAAAGCCCCTAAGCTCCTG ATCTATGCTGCATCCAGTTTGCAAAGTGGGGTCCCATCAAGPage 108 of 372 12815032v1Attorney Docket No.: 2017469-0043 TACTATTGCGCAGCAAGGAGAATCGACGCAGCAGACTTTG ATTCCTGGGGCCAGGGCACCCAGGTGACAGTGTCTAGCPage 109 of 372 12815032v1Attorney Docket No.: 2017469-0043 AGGCGTGCCGGCCAGCGGCGGGGGGCGCAGTGCACACGA GGGGGCTGGACTTCGCCTGTGATATCTACATCTGGGCGCCCPage 110 of 372 12815032v1Attorney Docket No.: 2017469-0043 GGGCAGGGAACACAGGTGACCGTGAGCAGCACCACGACG CCAGCGCCGCGACCACCAACACCGGCGCCCACCATCGCGTPage 111 of 372 12815032v1Attorney Docket No.: 2017469-0043 ACTTCACCCTGACCATCGACCCCGTGGAAGAGGACGACGT GGCCGTGTACTACTGCCTGCAGAGCAGAACCATCCCCCGGPage 112 of 372 12815032v1Attorney Docket No.: 2017469-0043 ATGAGCTGGGTCCGCCAGGCTCCAGGGAAGGGGCTGGAG TGGGTCTCATCTATTAGTGGTAGTGGTGATTACATATACTACPage 113 of 372 12815032v1Attorney Docket No.: 2017469-0043 GGCACCCTGGTCACCGTCTCCTCATTCGTGCCCGTGTTCCT GCCTGCCAAGCCTACAACAACCCCTGCTCCTAGACCTCCTAPage 114 of 372 12815032v1Attorney Docket No.: 2017469-0043 BCMA CAR-8 IYIWAPLAGTCGVLLLSLVITYC 15554 CD8αPage 115 of 372 12815032v1Attorney Docket No.: 2017469-0043 gccctgtgcagacaacccaggaggaggatggctgctcctgtaggttcccagaagaggaggag ggaggatgtgagctgcgcgtgaagttttctcggagcgccgacgcacctgcataccagcagggaPage 116 of 372 12815032v1Attorney Docket No.: 2017469-0043 GTGAAACCACATACTATAATTCAGCTCTCAAATCCAGACTG ACCATCATCAAGGACAACTCCAAGAGCCAAGTTTTCTTAAPage 117 of 372 12815032v1Attorney Docket No.: 2017469-0043 TGGTACTTGCGGGGTCCTGCTGCTTTCACTCGTGATCACTC TTTACTGTAGGAGTAAGAGGAGCAGGCTCCTGCACAGTGAPage 118 of 372 12815032v1Attorney Docket No.: 2017469-0043 GPRC5D Gaggtgcagctggtggagtctgggggaggcttggtcaagcctggagggtccctgagactctcc 15528 CAR2 VH tgtgcagcctctggattcaccttcagtgactactacatgagctggatccgccaggctccagggaaPage 119 of 372 12815032v1Attorney Docket No.: 2017469-0043 LYLQMNSLRAEDTAVYYCARSGGQWKYYDYWGQGTLVTVS SPage 120 of 372 12815032v1Attorney Docket No.: 2017469-0043 TATCCAGGAAATGGTGATACTTCCTACAATCAGAAGTTCAA AGGCAAGGCCACATTGACTGCAGACAAATCCTCCAGCACAPage 121 of 372 12815032v1Attorney Docket No.: 2017469-0043 GCCTCGGCGGAAGAACCCCCAGGAAGGCCTGTATAACGAA CTGCAGAAAGACAAGATGGCCGAGGCCTACAGCGAGATCGfrom Table 15), a CD8 hinge domain (e.g., from Table 14), a CD8 transmembrane domain (e.g., from Table 11), and a 4-1BB costimulatory domain (e.g., from Table 13). In some embodiments, the anti-BCMA CAR additionally comprises a CD3z signaling domain (e.g., from Table 13). In some embodiments, the anti-BCMA CAR comprises a BCMA binding domain (e.g., from Table 15), a CD8 hinge domain (e.g., from Table 14), a CD8 transmembrane domain (e.g., from Table 11), a 4-1BB costimulatory domain (e.g., from Table 13), and a CD3z signaling domain (e.g., from Table 13). In some embodiments, the anti-BCMA CAR comprises a BCMA binding domain (e.g., from Table 15), a CD20 epitope, a CD8 hinge domain (e.g., from Table 14), a CD8 transmembrane domain (e.g., from Table 11), a 4-1BB costimulatory domain (e.g., from Table 13), and a CD3z signaling domain (e.g., from Table 13). In some embodiments, the BCMA binding domain is murine. In some embodiments, the BCMA binding domain is humanized. In some embodiments, the BCMA binding domain is human. In some embodiments, the anti- BCMA CAR comprises a BCMA binding domain that comprises an scFv. In some embodiments, the anti-BCMA CAR comprises a BCMA binding domain that comprises two VHH domains (e.g., two linked camelid VHH antigen binding domains, e.g., VHH1 and VHH2 from Table 15). In some embodiments, the anti-BCMA CAR comprises a BCMA binding domain that comprises a D domain.

[0424] In some embodiments, the anti-BCMA CAR comprises a BCMA binding domain of Ciltacabtagene autoleucel (Carvykti), CT103A, CART-ddBCMA, NXC-201, idecabtagene vicleucel, ALLO-715, MCARH171, MCM998, P-BCMA-101, CTX120, or PBCAR269A. In some embodiments, the anti-BCMA CAR comprises Ciltacabtagene autoleucel (Carvykti), CT103A, CART-ddBCMA, NXC-201, idecabtagene vicleucel, ALLO-715, MCARH171, MCM998, P-BCMA-101, CTX120, or PBCAR269A.

[0425] In some embodiments, the anti-GPRC5D CAR comprises a VL and a VH domain from Table 15. In some embodiments, the anti-GPRC5D CAR comprises a VL domain, a VH domain, or an scFv of GPRC5D CAR1, GPRC5D CAR2, GPRC5D CAR3, or GPRC5D CAR4 Page 122 of 372 12815032v1Attorney Docket No.: 2017469-0043 in Table 15. In some embodiments, the anti-GPRC5D CAR comprises a VL domain, a VH domain, or an scFv of GPRC5D CAR1, GPRC5D CAR2, GPRC5D CAR3, or GPRC5D CAR4 in Table 15, a CD28 transmembrane domain (e.g., from Table 11), and a 4-1BB costimulatory domain (e.g., from Table 13). In some embodiments, the anti-GPRC5D CAR additionally comprises a CD3z signaling domain (e.g., from Table 13). In some embodiments, the GPRC5D binding domain is murine. In some embodiments, the GPRC5D binding domain is humanized. In some embodiments, the GPRC5D binding domain is human. In some embodiments, the anti- GPRC5D CAR comprises a GPRC5D binding domain that comprises an scFv. In some embodiments, the anti- GPRC5D CAR comprises a BCMA binding domain that comprises two VHH domains (e.g., two linked camelid VHH antigen binding domains, e.g., VHH1 and VHH2 from Table 15). In some embodiments, the anti- GPRC5D CAR comprises a GPRC5D binding domain that comprises a D domain.

[0426] In some embodiments, the anti-GPRC5D CAR comprises a GPRC5D binding domain of MCARH109, BMS-986393, or RD138. In some embodiments, the anti-GPRC5D CAR comprises MCARH109, BMS-986393, or RD138.

[0427] In some embodiments, the anti-CD19 CAR comprises a CD19 binding domain of FMC63. In some embodiments, the anti-CD19 CAR comprises a VL domain and a VH domain from Table 15. In some embodiments, the anti-CD19 CAR comprises a CD19 binding domain of FMC63, a CD28 hinge domain, a CD28 transmembrane domain, a CD28 costimulatory domain, and a CD3z signaling domain. In some embodiments, the anti-CD19 CAR comprises a CD19 binding domain of FMC63, a CD8α hinge domain, a CD8α transmembrane domain, a 4- 1BB costimulatory domain, and a CD3z signaling domain. In some embodiments, the anti-CD19 CAR comprises a CD19 binding domain of FMC63, an IgG4 hinge domain, a CD28 transmembrane domain, a 4-1BB costimulatory domain, and a CD3z signaling domain. In some embodiments, the anti-CD19 CAR comprises FMC63. In some embodiments, the anti-CD19 CAR comprises Axicabtagene ciloleucel, Brexucabtagene autoleucel, Tisagenlecleucel, or Lisocabtagene maraleucel. In some embodiments, the CD19 binding domain is murine. In some embodiments, the CD19 binding domain is humanized. In some embodiments, the CD19 binding domain is human. In some embodiments, the anti-CD19 CAR comprises a CD19 binding domain that comprises an scFv. Page 123 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0428] In some embodiments, the anti-CD20 CAR comprises a VL domain and a VH domain from Table 15. In some embodiments, the anti-CD20 CAR comprises a CD20 binding domain from Table 15. In some embodiments, the anti-CD20 CAR comprises a CD20 CAR from Table 15. In some embodiments, the CD20 binding domain is murine. In some embodiments, the CD20 binding domain is humanized. In some embodiments, the CD20 binding domain is human. In some embodiments, the anti- CD20 CAR comprises a CD20 binding domain that comprises a scFv.

[0429] In some embodiments, the anti-GPRC5D CAR comprises a VL domain and a VH domain from Table 15. In some embodiments, the anti-GPRC5D CAR comprises a GPRC5D binding domain from Table 15. In some embodiments, the anti-GPRC5D CAR comprises a GPRC5D CAR from Table 15. In some embodiments, the GPRC5D binding domain is murine. In some embodiments, the GPRC5D binding domain is humanized. In some embodiments, the anti- GPRC5D CAR comprises a GPRC5D binding domain that comprises a scFv.

[0430] In some embodiments, the CAR comprises two antigen binding domains that target different antigens on the surface of a cell, e.g., a neoplastic cell, e.g., a blood cancer cell such as one associated with a leukemia, a lymphoma, or a multiple myeloma. In some embodiments, a CAR-T cell is engineered to comprise two CARs with antigen binding domains that target a different antigen on the surface of a cell, e.g., a neoplastic cell, e.g., a blood cancer cell such as one associated with a leukemia, a lymphoma, or a multiple myeloma. In some embodiments, the antigen binding domains of the CAR target CD20 and CD22. In some embodiments, the antigen binding domains of the CAR target CD19 and CD20. In some embodiments, the antigen binding domains of the CAR target CD19 and CD22. In some embodiments, the antigen binding domains of the CAR target GPRC5D and BCMA. CAR Compositions, Methods of Manufacture, and Uses

[0431] Additionally provided herein is a system for modifying DNA of a mammalian cell (e.g., a T cell, e.g., a cytotoxic, helper, or regulatory T cell, e.g., a primary T cell) to express a CAR, the system comprising: (a) a gene modifying polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, as disclosed herein, and Page 124 of 372 12815032v1Attorney Docket No.: 2017469-0043 (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence encoding a chimeric antigen receptor (CAR), wherein the CAR comprises an antigen-binding domain, a transmembrane domain, a first intracellular signaling domain, and a second intracellular signaling domain, as disclosed herein.

[0432] In some embodiments, the system comprises: (a) a gene modifying polypeptide or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence encoding a chimeric antigen receptor (CAR), wherein one or more of: (i) the CAR comprises an antigen binding domain that binds one or more antigens of a blood cancer (e.g., a leukemia or lymphoma), wherein optionally the antigen is a B cell antigen; (ii) the CAR comprises an antigen binding domain that binds one or more antigens of a solid tumor; (iii) the CAR comprise an antigen binding domain of any one of Tables 5-9 or 15; (iv) the CAR comprise a linker domain of Table 16; (v) the CAR comprises a transmembrane domain of Table 10 or 11; (vi) the CAR comprises a hinge domain (e.g., a hinge domain of Table 14); (vii) the CAR comprises an intracellular signaling domain of Table 12 or 13; (viii) the CAR comprises a costimulatory domain of Table 12 or 13; (ix) the CAR comprises an antigen binding domain which comprises an scFv, a Fab, a diabody, a D domain binder, a centryin, or one or more single domain antibodies (e.g., VHH domains); or (x) the CAR comprises an amino acid sequence of Table 15 or an amino acid sequence according to any one of SEQ ID NOs: 1100, 15490, 15492, 15498, 15500, 15502, 15503, 15505, Page 125 of 372 12815032v1Attorney Docket No.: 2017469-0043 15507, 15509 and 15510, 15555, 15557 and 15558, 15559, 15560, 15561, 15515, 15526, 15531, 15536, 15541, or 15548; (xi) wherein the CAR comprises a first intracellular signaling domain and a second intracellular signaling domain.

[0433] In some embodiments, provided herein are a population of cells comprising immune effector cells (e.g., T cells, e.g., primary T cells) or regulatory T cell (e.g., primary T reg cells) comprising a plurality of copies of a gene encoding a CAR (“a CAR gene”). In some embodiments, less than less than or equal to 70%, 65%, 60%, or 55% of copies of the CAR gene in the population are situated within a gene endogenous to a cell of the population. In some embodiments, less than 10%, 9%, 8%, 7%, 6%, or 5% of copies of the CAR gene in the population are situated within an exon of a gene endogenous to a cell of the population. In some embodiments, less than 70%, 65%, 60%, 55%, or 50% of copies of the CAR gene in the population are situated within an intron of a gene endogenous to a cell of the population. In some embodiments, less than 10%, 9%, 8%, or 7% of copies of the CAR gene in the population are situated upstream of a gene (e.g., within 2 kb of a transcriptional start site (TSS)) endogenous to a cell of the population. In some embodiments, at least 20%, 25%, 30%, 35%, or 40% of copies of the CAR gene in the population are situated within an intergenic region endogenous to a cell of the population.

[0434] In some embodiments, at least 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, 99%, or 100% of cells in the plurality comprises a single copy of the CAR gene. In some embodiments, each cell in the plurality comprises a single copy of the CAR gene. In some embodiments, at least 0.1% of cells in the population comprise the CAR gene.

[0435] In some embodiments, the cell population comprises one or more of cancer cells, regulatory T cells, monocytes, and NK cells. In some embodiments, the immune effector cells and / or regulatory immune cells comprise T cells, e.g., primary T cells. In some embodiments, the immune effector cells or regulatory T cells comprise a leukapheresis sample or an apheresis sample. In some embodiments, the population of cells is substantially free of lentivirus proteins. In some embodiments, the population of cells is substantially free of lentivirus nucleic acids.

[0436] Additionally provided herein are methods of modifying a mammalian cell to express a CAR. In some embodiments, the method involves contacting the cell (e.g., an immune effector Page 126 of 372 12815032v1Attorney Docket No.: 2017469-0043 cell or a regulatory T cell) with a system disclosed herein. In some embodiments, the immune effector cell is a cell that expresses one or more Fc receptors and mediates one or more effector functions. In some embodiments, the immune effector cell may include, but may not be limited to, one or more of a monocyte, macrophage, neutrophil, dendritic cell, eosinophil, mast cell, platelet, large granular lymphocyte, Langerhans' cell, natural killer (NK) cell, T-lymphocyte (e.g., T-cell), a Gamma delta T cell, B-lymphocyte (e.g., B-cell) and may be from any organism including but not limited to humans, mice, rats, rabbits, and monkeys.

[0437] In some embodiments, the regulatory T cell (Treg) is a cell that suppresses an immune response, e.g., to mediate homeostasis and induce immune tolerance. In some embodiments, the Treg cell may include, but may not be limited to, a natural Treg or induced Treg and may be from any organism including but not limited to humans, mice, rats, rabbits, and monkeys.

[0438] In some embodiments, the mammalian cell is a T cell, e.g., a primary T cell. In some embodiments, provided herein is a method of modifying the genome of a mammalian T cell (e.g., a primary T cell), the method comprising contacting the cell with: (a) a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence (e.g., encoding a CAR).

[0439] In some embodiments, provided herein is a method of modifying the genome of a mammalian T cell (e.g., a primary T cell), the method comprising contacting the cell with: (a) a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence having no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid differences thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence (e.g., encoding a CAR).

[0440] In some embodiments, the method is performed ex vivo or in vitro. Page 127 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0441] Additionally provided herein is a reaction mixture comprising a gene modifying system disclosed herein (e.g., comprising a heterologous object sequence encoding a chimeric antigen receptor (CAR)) and a mammalian cell (e.g., a T cell, e.g., a primary T cell).

[0442] In some embodiments, the reaction mixture comprises a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide and a mammalian T cell (e.g., a primary T cell). In some embodiments, the reaction mixture additionally includes a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence.

[0443] In some embodiments, the reaction mixture comprises a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid differences thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide and a mammalian T cell (e.g., a primary T cell). In some embodiments, the reaction mixture additionally includes a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence.

[0444] In some embodiments, the systems, reaction mixtures, and cell populations modified using the methods disclosed herein can be used to treat a subject in need thereof. In some embodiments, the subject has a cancer, e.g., a hematological cancer or a solid tumor. In some embodiments, the subject has an infectious disease. In other embodiments, the subject has an autoimmune or an inflammatory disease. Compositions and Methods for Modifying Mammalian Cells

[0445] In some embodiments, provided herein are methods of modifying mammalian cells and reaction mixtures and systems for the same.

[0446] In some embodiments, provided herein is a method of modifying the genome of a mammalian cell (e.g., a mammalian induced pluripotent stem cell (iPSC) or T cell, e.g., a primary T cell), the method comprising contacting the cell with: Page 128 of 372 12815032v1Attorney Docket No.: 2017469-0043 (a) a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence.

[0447] In some embodiments, provided herein is a method of modifying the genome of a mammalian cell (e.g., a mammalian induced pluripotent stem cell (iPSC) or T cell, e.g., a primary T cell), the method comprising contacting the cell with a system comprising: (a) a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence having no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid differences thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide, and (b) a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence.

[0448] In some embodiments, the method is performed ex vivo or in vitro. In some embodiments, the gene modifying polypeptide and / or template RNA are formulated with an LNP.

[0449] In some embodiments, contacting the cell with the system results in insertion of the heterologous object sequence into at least 1%, 5%, 10%, 15%, 20%, 25%, or 30% of T cells.

[0450] Additionally provided herein is a reaction mixture comprising a gene modifying system disclosed herein (e.g., comprising a heterologous object sequence) and a mammalian cell (e.g., an iPSC or a T cell, e.g., a primary T cell).

[0451] In some embodiments, the reaction mixture comprises a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence at least 70%, 75%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identical thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide and a mammalian T cell (e.g., a primary T cell). In some embodiments, the reaction mixture additionally includes a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence. Page 129 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0452] In some embodiments, the reaction mixture comprises a gene modifying polypeptide comprising an amino acid sequence of Table 1 or a sequence no more than 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 amino acid differences thereto, or a nucleic acid (e.g., DNA or mRNA) encoding the gene modifying polypeptide and a mammalian T cell (e.g., a primary T cell). In some embodiments, the reaction mixture additionally includes a template RNA (or DNA encoding the template RNA) comprising (i) a sequence that binds the polypeptide and (ii) a heterologous object sequence.

[0453] Additionally provided herein is a population of T cells produced according to a method disclosed herein. In some embodiments, the population comprises a plurality of copies of the heterologous object sequence, wherein less than or equal to 70%, 65%, 60%, or 55% of copies of the heterologous object sequence in the population are situated within a gene endogenous to a cell of the population. In some embodiments, less than 10%, 9%, 8%, 7%, 6%, or 5% of copies of the heterologous object sequence in the population are situated within an exon of a gene endogenous to a cell of the population. In some embodiments, less than 70%, 65%, 60%, 55%, or 50% of copies of the heterologous object sequence in the population are situated within an intron of a gene endogenous to a cell of the population. In some embodiments, less than 10%, 9%, 8%, or 7% of copies of the heterologous object sequence in the population are situated upstream of a gene (e.g., within 2 kb of a transcriptional start site (TSS)) endogenous to a cell of the population. In some embodiments, at least 20%, 25%, 30%, 35%, or 40% of copies of the heterologous object sequence in the population are situated within an intergenic region endogenous to a cell of the population. Suitable Indications

[0454] Exemplary suitable diseases and disorders that can be treated by the systems or methods provided herein, for example, those comprising gene modifying systems, include, without limitation, those described in PCT Publication No. WO / 2021 / 178717, incorporated herein by reference in its entirety. In some embodiments, diseases and disorders that can be treated by the systems or methods described herein include cancer or autoimmune diseases. Exemplary heterologous object sequences

[0455] In some embodiments, the systems or methods provided herein comprise a heterologous object sequence, wherein the heterologous object sequence or a reverse complementary sequence Page 130 of 372 12815032v1Attorney Docket No.: 2017469-0043 thereof, encodes a protein (e.g., an antibody) or peptide. In some embodiments, the therapy is one approved by a regulatory agency such as FDA.

[0456] In some embodiments, the heterologous object sequence comprises a sequence as listed in PCT Publication No. WO / 2021 / 178717, incorporated herein by reference in its entirety. In some embodiments, the heterologous object sequence encodes a polypeptide or peptide (e.g., a therapeutic peptide or polypeptide). In some embodiments, the heterologous object sequence comprises a promoter. In some embodiments, the heterologous object sequence encodes a CAR (e.g., as described herein). Administration

[0457] The composition and systems described herein may be used in vitro or in vivo. In some embodiments the system or components of the system are delivered to cells (e.g., mammalian cells, e.g., human cells), e.g., in vitro or in vivo. In some embodiments, the cells are eukaryotic cells, e.g., cells of a multicellular organism, e.g., an animal, e.g., a mammal (e.g., human, swine, bovine) a bird (e.g., poultry, such as chicken, turkey, or duck), or a fish. In some embodiments, the cells are non-human animal cells (e.g., a laboratory animal, a livestock animal, or a companion animal). In some embodiments, the cell is a stem cell (e.g., a hematopoietic stem cell or iPSC), a fibroblast, or a T cell. In some embodiments, the cell is a non-dividing cell, e.g., a non-dividing fibroblast or non-dividing T cell. In some embodiments, the cell is an HSC and p53 is not upregulated or is upregulated by less than 10%, 5%, 2%, or 1%, e.g., as determined according to the method described in Example 30 of PCT Application No. PCT / US2019 / 048607, incorporated herein by reference in its entirety. In some embodiments, the cell is a T cell, e.g., a primary T cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC). The components of the gene modifying system may, in some instances, be delivered in the form of polypeptide, nucleic acid (e.g., DNA, RNA), and combinations thereof.

[0458] For instance, delivery can use any of the following combinations for delivering the retrotransposase (e.g., as DNA encoding the retrotransposase protein, as RNA encoding the retrotransposase protein, or as the protein itself) and the template RNA (e.g., as DNA encoding the RNA, or as RNA): Page 131 of 372 12815032v1Attorney Docket No.: 2017469-0043 1. Retrotransposase DNA + template DNA 2. Retrotransposase RNA + template DNA 3. Retrotransposase DNA + template RNA 4. Retrotransposase RNA + template RNA 5. Retrotransposase protein + template DNA 6. Retrotransposase protein + template RNA 7. Retrotransposase virus + template virus 8. Retrotransposase virus + template DNA 9. Retrotransposase virus + template RNA 10. Retrotransposase DNA + template virus 11. Retrotransposase RNA + template virus 12. Retrotransposase protein + template virus

[0459] As indicated above, in some embodiments, the DNA or RNA that encodes the retrotransposase protein is delivered using a virus, and in some embodiments, the template RNA (or the DNA encoding the template RNA) is delivered using a virus.

[0460] In one embodiments, the system and / or components of the system are delivered as nucleic acid. For example, the gene modifying polypeptide may be delivered in the form of a DNA or RNA encoding the polypeptide and the template RNA may be delivered in the form of RNA or its complementary DNA to be transcribed into RNA. In some embodiments, the system or components of the system are delivered on 1, 2, 3, 4, or more distinct nucleic acid molecules. In some embodiments, the system or components of the system are delivered as a combination of DNA and RNA. In some embodiments, the system or components of the system are delivered as a combination of DNA and protein. In some embodiments, the system or components of the system are delivered as a combination of RNA and protein. In some embodiments, the gene modifying polypeptide is delivered as a protein.

[0461] In some embodiments, the system or components of the system are delivered to cells, e.g. mammalian cells or human cells, using a vector. The vector may be, e.g., a plasmid or a virus. In Page 132 of 372 12815032v1Attorney Docket No.: 2017469-0043 some embodiments, delivery is in vivo, in vitro, ex vivo, or in situ. In some embodiments, the virus is an adeno associated virus (AAV), a lentivirus, an adenovirus. In some embodiments, the virus is selected from a Group I, Group II, Group III, Group IV, Group V, Group VI, or Group VII virus. In some embodiments, the system or components of the system are delivered to cells with a viral-like particle or a virosome. In some embodiments, the delivery uses more than one virus, viral-like particle, or virosome.

[0462] In some embodiments, nucleic acid (e.g., encoding a polypeptide, or a template DNA, or both) delivered to cells is covalently closed linear DNA, or so-called “doggybone” DNA. During its lifecycle, the bacteriophage N15 employs protelomerase to convert its genome from circular plasmid DNA to a linear plasmid DNA (Ravin et al. J Mol Biol 2001). This process has been adapted for the production of covalently closed linear DNA in vitro (see, for example, WO2010086626A1). In some embodiments, a protelomerase is contacted with a DNA containing one or more protelomerase recognition sites, wherein protelomerase results in a cut at the one or more sites and subsequent ligation of the complementary strands of DNA, resulting in the covalent linkage between the complementary strands. In some embodiments, nucleic acid (e.g., encoding a transposase, or a template DNA, or both) is first generated as circular plasmid DNA containing a single protelomerase recognition site that is then contacted with protelomerase to yield a covalently closed linear DNA. In some embodiments, nucleic acid (e.g., encoding a transposase, or a template DNA, or both) flanked by protelomerase recognition sites on plasmid or linear DNA is contacted with protelomerase to generate a covalently closed linear DNA containing only the DNA contained between the protelomerase recognition sites. In some embodiments, the approach of flanking the desired nucleic acid sequence by protelomerase recognition sites results in covalently closed circular DNA lacking plasmid elements used for bacterial cloning and maintenance. In some embodiments, the plasmid or linear DNA containing the nucleic acid and one or more protelomerase recognition sites is optionally amplified prior to the protelomerase reaction, e.g., by rolling circle amplification or PCR.

[0463] In one embodiment, the compositions and systems described herein can be formulated in liposomes or other similar vesicles. Liposomes are spherical vesicle structures composed of a uni- or multilamellar lipid bilayer surrounding internal aqueous compartments and a relatively impermeable outer lipophilic phospholipid bilayer. Liposomes may be anionic, neutral, or Page 133 of 372 12815032v1Attorney Docket No.: 2017469-0043 cationic. Liposomes are biocompatible, nontoxic, can deliver both hydrophilic and lipophilic drug molecules, protect their cargo from degradation by plasma enzymes, and transport their load across biological membranes and the blood brain barrier (BBB) (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for review).

[0464] Vesicles can be made from several different types of lipids; however, phospholipids are most commonly used to generate liposomes as drug carriers. Methods for preparation of multilamellar vesicle lipids are known in the art (see for example U.S. Pat. No.6,693,086, the teachings of which relating to multilamellar vesicle lipid preparation are incorporated herein by reference). Although vesicle formation can be spontaneous when a lipid film is mixed with an aqueous solution, it can also be expedited by applying force in the form of shaking by using a homogenizer, sonicator, or an extrusion apparatus (see, e.g., Spuch and Navarro, Journal of Drug Delivery, vol.2011, Article ID 469679, 12 pages, 2011. doi:10.1155 / 2011 / 469679 for review). Extruded lipids can be prepared by extruding through filters of decreasing size, as described in Templeton et al., Nature Biotech, 15:647-652, 1997, the teachings of which relating to extruded lipid preparation are incorporated herein by reference.

[0465] Lipid nanoparticles are another example of a carrier that provides a biocompatible and biodegradable delivery system for the pharmaceutical compositions described herein. Nanostructured lipid carriers (NLCs) are modified solid lipid nanoparticles (SLNs) that retain the characteristics of the SLN, improve drug stability and loading capacity, and prevent drug leakage. Polymer nanoparticles (PNPs) are an important component of drug delivery. These nanoparticles can effectively direct drug delivery to specific targets and improve drug stability and controlled drug release. Lipid–polymer nanoparticles (PLNs), a new type of carrier that combines liposomes and polymers, may also be employed. These nanoparticles possess the complementary advantages of PNPs and liposomes. A PLN is composed of a core–shell structure; the polymer core provides a stable structure, and the phospholipid shell offers good biocompatibility. As such, the two components increase the drug encapsulation efficiency rate, facilitate surface modification, and prevent leakage of water-soluble drugs. For a review, see, e.g., Li et al.2017, Nanomaterials 7, 122; doi:10.3390 / nano7060122. Page 134 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0466] Exosomes can also be used as drug delivery vehicles for the compositions and systems described herein. For a review, see Ha et al. July 2016. Acta Pharmaceutica Sinica B. Volume 6, Issue 4, Pages 287-296; doi.org / 10.1016 / j.apsb.2016.02.001.

[0467] Fusosomes interact and fuse with target cells, and thus can be used as delivery vehicles for a variety of molecules. They generally consist of a bilayer of amphipathic lipids enclosing a lumen or cavity and a fusogen that interacts with the amphipathic lipid bilayer. The fusogen component has been shown to be engineerable in order to confer target cell specificity for the fusion and payload delivery, allowing the creation of delivery vehicles with programmable cell specificity (see, for example, the relating to fusosome design, preparation, and usage in PCT Publication No. WO / 2020014209, incorporated herein by reference in its entirety).

[0468] A gene modifying system can be introduced into cells, tissues, and multicellular organisms. In some embodiments, the system or components of the system are delivered to the cells via mechanical means or physical means.

[0469] Formulation of protein therapeutics is described in Meyer (Ed.), Therapeutic Protein Drug Products: Practical Approaches to formulation in the Laboratory, Manufacturing, and the Clinic, Woodhead Publishing Series (2012). Tissue Specific Activity / Administration

[0470] In some embodiments, a system, template RNA, or polypeptide described herein is administered to or is active in (e.g., is more active in) a target tissue, e.g., a first tissue. In some embodiments, the system, template RNA, or polypeptide is not administered to or is less active in (e.g., not active in) a non-target tissue. In some embodiments, a system, template RNA, or polypeptide described herein is useful for modifying DNA in a target tissue, e.g., a first tissue, (e.g., and not modifying DNA in a non-target tissue).

[0471] In some embodiments, a system comprises (a) a polypeptide described herein or a nucleic acid encoding the same, (b) a template nucleic acid (e.g., template RNA) described herein, and (c) one or more first tissue-specific expression-control sequences specific to the target tissue, wherein the one or more first tissue-specific expression-control sequences specific to the target tissue are in operative association with (a), (b), or (a) and (b), wherein, when associated with (a), (a) comprises a nucleic acid encoding the polypeptide. Page 135 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0472] In some embodiments, the heterologous object sequence is in operative association with a first promoter.

[0473] In some embodiments, the one or more first tissue-specific expression-control sequences comprises a tissue specific promoter.

[0474] In some embodiments, the tissue-specific promoter comprises a first promoter in operative association with: i. the heterologous object sequence, ii. a nucleic acid encoding the transposase, or iii. (i) and (ii).

[0475] In some embodiments, a system comprises a tissue-specific promoter, and the system further comprises one or more tissue-specific microRNA recognition sequences, wherein: i. the tissue specific promoter is in operative association with: I. the heterologous object sequence, II. a nucleic acid encoding the transposase, or III. (I) and (II); and / or ii. the one or more tissue-specific microRNA recognition sequences are in operative association with: I. the heterologous object sequence, II. a nucleic acid encoding the transposase, or III. (I) and (II).

[0476] In some embodiments, a gene modifying system described herein is delivered to a tissue or cell from the cerebrum, cerebellum, adrenal gland, ovary, pancreas, parathyroid gland, hypophysis, testis, thyroid gland, breast, spleen, tonsil, thymus, lymph node, bone marrow, lung, cardiac muscle, esophagus, stomach, small intestine, colon, liver, salivary gland, kidney, prostate, blood, or other cell or tissue type. In some embodiments, a gene modifying system described herein is used to treat a disease, such as a cancer, inflammatory disease, infectious disease, genetic defect, or other disease. A cancer can be cancer of the cerebrum, cerebellum, adrenal gland, ovary, pancreas, parathyroid gland, hypophysis, testis, thyroid gland, breast, spleen, tonsil, thymus, lymph node, bone marrow, lung, cardiac muscle, esophagus, stomach, small intestine, colon, liver, salivary gland, kidney, prostate, blood, or other cell or tissue type, and can include multiple cancers.

[0477] In some embodiments, a gene modifying system described herein described herein is administered by enteral administration (e.g. oral, rectal, gastrointestinal, sublingual, sublabial, or buccal administration). In some embodiments, a gene modifying system described herein is administered by parenteral administration (e.g., intravenous, intramuscular, subcutaneous, intradermal, epidural, intracerebral, intracerebroventricular, epicutaneous, nasal, intra-arterial, intra-articular, intracavernous, intraocular, intraosseous infusion, intraperitoneal, intrathecal, Page 136 of 372 12815032v1Attorney Docket No.: 2017469-0043 intrauterine, intravaginal, intravesical, perivascular, or transmucosal administration). In some embodiments, a gene modifying system described herein is administered by topical administration (e.g., transdermal administration).

[0478] In some embodiments, a gene modifying system as described herein can be used to modify an animal cell, plant cell, or fungal cell. In some embodiments, a gene modifying system as described herein can be used to modify a mammalian cell (e.g., a human cell). In some embodiments, a gene modifying system as described herein can be used to modify a cell from a livestock animal (e.g., a cow, horse, sheep, goat, pig, llama, alpaca, camel, yak, chicken, duck, goose, or ostrich). In some embodiments, a gene modifying system as described herein can be used as a laboratory tool or a research tool, or used in a laboratory method or research method, e.g., to modify an animal cell, e.g., a mammalian cell (e.g., a human cell), a plant cell, or a fungal cell.

[0479] In some embodiments, a gene modifying system as described herein can be used to express a protein, template, or heterologous object sequence (e.g., in an animal cell, e.g., a mammalian cell (e.g., a human cell), a plant cell, or a fungal cell). In some embodiments, a gene modifying system as described herein can be used to express a protein, template, or heterologous object sequence under the control of an inducible promoter (e.g., a small molecule inducible promoter). In some embodiments, a gene modifying system or payload thereof is designed for tunable control, e.g., by the use of an inducible promoter. For example, a promoter, e.g., Tet, driving a gene of interest may be silent at integration, but may, in some instances, activated upon exposure to a small molecule inducer, e.g., doxycycline. In some embodiments, the tunable expression allows post-treatment control of a gene (e.g., a therapeutic gene), e.g., permitting a small molecule-dependent dosing effect. In embodiments, the small molecule-dependent dosing effect comprises altering levels of the gene product temporally and / or spatially, e.g., by local administration. In some embodiments, a promoter used in a system described herein may be inducible, e.g., responsive to an endogenous molecule of the host and / or an exogenous small molecule administered thereto.

[0480] In some embodiments, a nucleic acid component of a system provided herein comprises a sequence (e.g., encoding the polypeptide or comprising a heterologous object sequence) is flanked by untranslated regions (UTRs) that modify protein expression levels. Various 5’ and 3’ Page 137 of 372 12815032v1Attorney Docket No.: 2017469-0043 UTRs can affect protein expression. For example, in some embodiments, the coding sequence may be preceded by a 5’ UTR that modifies RNA stability or protein translation. In some embodiments, the sequence may be followed by a 3’ UTR that modifies RNA stability or translation. In some embodiments, the sequence may be preceded by a 5’ UTR and followed by a 3’ UTR that modify RNA stability or translation. In some embodiments, the 5’ and / or 3’ UTR may be selected from the 5’ and 3’ UTRs of complement factor 3 (C3) (cactcctccccatcctctccctctgtccctctgtccctctgaccctgcactgtcccagcacc) or orosomucoid 1 (ORM1) (caggacacagccttggatcaggacagagacttgggggccatcctgcccctccaacccgacatgtgtacctcagctttttccctcacttgcat caataaagcttctgtgtttggaacagctaa) (Asrani et al. RNA Biology 2018). In certain embodiments, the 5’ UTR is the 5’ UTR from C3 and the 3’ UTR is the 3’ UTR from ORM1.

[0481] In certain embodiments, a 5’ UTR and 3’ UTR for protein expression, e.g., mRNA (or DNA encoding the RNA) for a gene modifying polypeptide or heterologous object sequence, comprise optimized expression sequences. In some embodiments, the 5’ UTR comprises GGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC and / or the 3’ UTR comprising UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCC AGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA, e.g., as described in Richner et al. Cell 168(6): P1114-1125 (2017), the sequences of which are incorporated herein by reference.

[0482] In some embodiments, a 5’ and / or 3’ UTR may be selected to enhance protein expression. In some embodiments, a 5’ and / or 3’ UTR may be selected to modify protein expression such that overproduction inhibition is minimized. In some embodiments, UTRs are around a coding sequence, e.g., outside the coding sequence and in other embodiments proximal to the coding sequence. In some embodiments, additional regulatory elements (e.g., miRNA binding sites, cis- regulatory sites) are included in the UTRs.

[0483] In some embodiments, an open reading frame (ORF) of a gene modifying system, e.g., an ORF of an mRNA (or DNA encoding an mRNA) encoding a gene modifying polypeptide or one or more ORFs of an mRNA (or DNA encoding an mRNA) of a heterologous object sequence, is flanked by a 5’ and / or 3’ untranslated region (UTR) that enhances the expression thereof. In some embodiments, the 5’ UTR of an mRNA component (or transcript produced from a DNA Page 138 of 372 12815032v1Attorney Docket No.: 2017469-0043 component) of the system comprises the sequence 5’- GGGAAAUAAGAGAGAAAAGAAGAGUAAGAAGAAAUAUAAGAGCCACC-3’. In some embodiments, the 3’ UTR of an mRNA component (or transcript produced from a DNA component) of the system comprises the sequence 5’- UGAUAAUAGGCUGGAGCCUCGGUGGCCAUGCUUCUUGCCCCUUGGGCCUCCCCCC AGCCCCUCCUCCCCUUCCUGCACCCGUACCCCCGUGGUCUUUGAAUAAAGUCUGA- 3’. This combination of 5’ UTR and 3’ UTR has been shown to result in desirable expression of an operably linked ORF by Richner et al. Cell 168(6): P1114-1125 (2017), the teachings and sequences of which are incorporated herein by reference. In some embodiments, a system described herein comprises a DNA encoding a transcript, wherein the DNA comprises the corresponding 5’ UTR and 3’ UTR sequences, with T substituting for U in the above-listed sequence). In some embodiments, a DNA vector used to produce an RNA component of the system further comprises a promoter upstream of the 5’ UTR for initiating in vitro transcription, e.g., a T7, T3, or SP6 promoter. The 5’ UTR above begins with GGG, which is a suitable start for optimizing transcription using T7 RNA polymerase. For tuning transcription levels and altering the transcription start site nucleotides to fit alternative 5’ UTRs, the teachings of Davidson et al. Pac Symp Biocomput 433-443 (2010) describe T7 promoter variants, and the methods of discovery thereof, that fulfill both of these traits. Lipid Nanoparticles and Conjugates Thereof

[0484] The methods and systems provided herein may employ any suitable carrier of delivery modality, including, in certain embodiments, lipid nanoparticles (LNPs). Lipid nanoparticles, in some embodiments, comprise one or more ionic lipids, one or more conjugated lipids, one or more sterols, and optionally, one or more targeting molecules, or combinations of the foregoing.

[0485] In one aspect, the disclosure provides an LNP (conjugate) comprising an ionizable lipid as described herein (e.g., in Table 5), wherein the LNP can deliver a payload, such as a therapeutic agent (e.g., a gene modifying system, such as a retrotransposon gene modifying system and / or a heterologous gene modifying system, as described herein) to an immune cell (e.g., a T cell). In another aspect, the LNP (conjugate) comprises a targeting moiety that binds to a protein (e.g., a protein receptor) on an immune cell (e.g., a T cell), as described herein. In some embodiments, an LNP (conjugate) comprises both an ionizable lipid and a targeting Page 139 of 372 12815032v1Attorney Docket No.: 2017469-0043 moiety. In some embodiments, the LNP (conjugate) delivers greater than 90% of the payload to T cells. In some embodiments, the LNP (conjugate) delivers from about 90% to about 100% of the payload to T cells.

[0486] In one aspect, the disclosure provides targeted LNPs (conjugates) comprising a targeting moiety and a lipid nanoparticle (LNP) encapsulating a payload (e.g., a therapeutic agent, as described herein, such as a gene modifying polypeptide or a gene modifying system), wherein the targeting moiety binds to a protein (e.g., protein receptor) on an immune cell (e.g., T cell). In some embodiments, the targeting moiety is an antibody or antigen binding fragment thereof. In some instances, the targeting moiety is an antibody, a Fab fragment, a scFv, a DARPIN, a VHH domain antibody, a FN3 domain, a nanobody, a single domain antibody or a Centyrin. In other embodiments, the targeting moiety is a folate moiety, an antibiotic mimetic, a polynucleotide (such as a DNA or RNA apatamer), a carbohydrate, a vitamin or a N-Acetylgalactosamine (GalNac). In some embodiments, the payload (e.g., a therapeutic agent, as described herein, such as a gene modifying polypeptide or a gene modifying system) is capable of modifying one or more genes of the target immune cell (e.g., T cell).

[0487] The conjugates described herein may be used to target and modify immune cells. In some embodiments, the conjugates may be used to modify T cells. In some embodiments, T-cells may include any subpopulation of T-cells, e.g., CD4+, CD8+, gamma-delta, naïve T cells, stem cell memory T cells, central memory T cells, or a mixture of subpopulations. In some embodiments, the conjugates may be used to deliver or modify a sequence encoding a T-cell receptor (TCR) in a T cell. In some embodiments, the conjugates may be used to deliver at least one sequence encoding a chimeric antigen receptor (CAR) to T-cells. For instance, in specific embodiments, the conjugates can be used to deliver an RNA encoding a CAR to T-cells. A. Targeting moieties

[0488] In some embodiments, the LNP comprises a targeting moiety. In some embodiments, the targeting moiety is a T-cell targeting moiety, for example, an antibody, Fab fragment or ScFv that binds to a T-cell antigen selected from the group consisting of CD2, CD3, CD4, CD5, CD7, CD8, CD28, CD137, CD45, T-cell receptor (TCR)β,TCR-α, TCR-α / β, TCR-γ / δ, PD1, CTLA4, TIM3, LAG3, CD18, IL-2 receptor, CD11a, TLR2, TLR4, TLR5, IL-7 receptor, or IL-15 receptor. Page 140 of 372 12815032v1Attorney Docket No.: 2017469-0043

[0489] In certain embodiments, the targeting moiety is a T-cell targeting moiety, for example, an antibody, Fab fragment or ScFv that binds to a T-cell antigen selected from the group consisting of CD2, CD3, CD4, CD5, CD7, CD8, CD28, CD137, CD45, T-cell receptor (TCR)β,TCR-α, TCR-α / β, TCR-γ / δ, PD1, CTLA4, TIM3, LAG3, CD18, IL-2 receptor, CD11a, TLR2, TLR4, TLR5, IL-7 receptor, and IL-15 receptor. In some embodiments the CD80 targeting moiety is a CD80 extracellular domain (ECD).

[0490] In some embodiments, the targeted LNP (conjugate) comprises a targeting moiety that targets a receptor on the surface of the T cell selected from CD2, CD3, CD4, CD5, CD6, CD7, CD8, and CD28. In some embodiments, the targeting moiety targets a CD2 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD3 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD4 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD5 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD7 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD8 receptor on the surface of the T cell. In some embodiments, the targeting moiety targets a CD28 receptor on the surface of the T cell.

[0491] In some embodiments, the targeted LNP (conjugate) comprises a targeting moiety that targets CD3 on the surface of the T cell, wherein the targeting moiety is an antibody, Fab fragment or scFv selected from SP34, teclistamab, mosunetuzumab, odronextamab, tebentafusp, tepilizumab, muromonab and visilizumabm, or an antigen-binding portion thereof. In certain embodiments, the targeting moiety is SP34 or an antigen-binding portion thereof. In other embodiments, the targeting moiety is teclistamab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is mosunetuzumab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is odronextamab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is tebentafusp or an antigen-binding portion thereof. In other embodiments, the targeting moiety is muromonab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is visilizumab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is tepilizumab or an antigen-binding portion thereof.. In other embodiments, the targeting moiety is Plamotamab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is HPN536 or an antigen-binding portion thereof. In Page 141 of 372 12815032v1Attorney Docket No.: 2017469-0043 other embodiments, the targeting moiety is Pasotuxizumab or an antigen-binding portion thereof. In other embodiments, the targeting moiety is Flotetuzumab or an antigen-binding portion thereof.

[0492] In some embodiments, the targeted LNP (conjugate) comprises a plurality of targeting moieties conjugated to the LNP, wherein the plurality of targeting moieties bind to at least one targeting moiety on a T cell.

[0493] In some embodiments, the plurality of targeting moieties bind to two or more T-cell antigens selected from the group consisting of CD2, CD3, CD4, CD5, CD7, CD8, CD28, CD80, CD137, CD45, T-cell receptor (TCR)-β,TCR-α, TCR-α / β, TCR-γ / δ, PD1, CTLA4, TIM3, LAG3, CD18, IL-2 receptor, CD11a, TLR2, TLR4, TLR5, IL-7 receptor, and IL-15 receptor. In some embodiments, the targeted LNP (conjugate) comprises a targeting moiety that targets a receptor on the surface of the T cell selected from CD2, CD3, CD4, CD5, CD6, CD7, and CD28. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD2. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD4. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD5. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD6. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD7. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD28. In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD3. In some embodiments, the one targeting moiety that binds to CD3 comprises a different anti-CD3 antibody, Fab fragment, or scFv than the other targeting moiety that binds to CD3.

[0494] In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD2. In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD5. In some Page 142 of 372 12815032v1Attorney Docket No.: 2017469-0043 embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD7. In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD28. In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD5 and the other targeting moiety binds to CD28. In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein one targeting moiety binds to CD7 and the other targeting moiety binds to CD28. In some embodiments the CD28 targeting moiety is a CD80 extracellular domain (ECD). In some embodiments, a targeted LNP comprises two targeting moieties, wherein one targeting moiety binds to CD3 and the other targeting moiety binds to CD3. In some embodiments, the one targeting moiety that binds to CD3 comprises a different anti-CD3 antibody, Fab fragment, or scFv than the other targeting moiety that binds to CD3.

[0495] In some embodiments, the targeted LNP (conjugate) comprises two targeting moieties, wherein each targeting moiety binds to the same target (e.g., receptor) on the T cell. For instance, in some embodiments, both targeting moieties of the conjugate bind to CD3. In some such embodiments, one of the targets is SP34 or an antigen-binding portion thereof and the other is teclistamab or an antigen-binding portion thereof. In other such embodiments, one of the targets is SP34 or an antigen-binding portion thereof and the other is visilizumab or an antigen- binding portion thereof. In other such embodiments, one of the targets is SP34 or an antigen- binding portion thereof and the other is tepilizumab or an antigen-binding portion thereof. In other such embodiments, one of the targets is visilizumab or an antigen-binding portion thereof and the other is tepilizumab or an antigen-binding portion thereof. In other such embodiments, one of the targets is visilizumab or an antigen-binding portion thereof and the other is teclistamab or an antigen-binding portion thereof. In some embodiments, both targeting moieties of the conjugate bind to CD3. In some embodiments, both targeting moieties of the conjugate bind to CD5. In some embodiments, both targeting moieties of the conjugate bind to CD7.

[0496] In certain embodiments, the targeting moiety binds to a CD4+ and / or CD8+ T cell. In other embodiments, the targeting moiety binds to a natural killer (NK) cell. In other embodiments, the targeting moiety binds to a hematopoietic stem cell. In other embodiments, Page 143 of 372 12815032v1Attorney Docket No.: 2017469-0043 the targeting moiety binds to a lymphoid progenitor cell. In other embodiments, the targeting moiety binds to a myeloid cell. In other embodiments, the targeting moiety binds to a macrophage. CD2 Targeting Moieties

[0497] In some embodiments, the target molecule is CD2. In some embodiments, the target cell is CD2+. The glycoprotein CD2 is a costimulatory receptor expressed mainly on T cells, NK cells, thymocytes, and dendritic cells that binds to lymphocyte-associated antigen 3 (LFA3; also known as CD58) which is expressed on the surface of B cells, T cells, monocytes, granulocytes, thymic epithelial cells. CD2 also binds to CD48, albeit with a relatively lower affinity. CD2 has an important role in the formation and organization of the immunological synapse that is formed between T cells and antigen-presenting cells upon cell-cell conjugation and associated intracellular signaling. CD2 expression is upregulated on memory T cells as well as activated T cells and plays an important role in activation of memory T cells. See, e.g., Binder et al. (2020) Front. Immunol.11:1090, hereby incorporated by reference in its entirety.

[0498] In some embodiments, the CD2 targeting moiety includes an antibody or antigen-binding fragment thereof that binds to CD2. In some embodiments, the CD2 targeting moiety is an antibody or antigen-binding fragment thereof (e.g., a Fab, Fab’, F(ab’)2, Fv fragment, scFv, DARPIN, VHH domain, FN3 domain, nanobody, single domain antibody, or Centyrin). In other embodiments, the CD2 targeting moiety includes a ligand, a folate moiety, an antibiotic mimetic, a polynucleotide (such as a DNA or RNA apatamer), a carbohydrate, a vitamin, a cytokine, or a chemokine. In some embodiments, the CD2 targeting moiety is an anti-CD2 antibody or antigen binding fragment thereof. In some embodiments, the CD2 targeting moiety is an IgA, IgG, IgE, or IgM antibody. In some embodiments, the CD2 targeting moiety is a bispecific or multi- specific antibody or fragment thereof. In some embodiments, the CD2 targeting moiety is a humanized antibody or antigen-binding fragment thereof.

[0499] Exemplary anti-CD2 binders, antibodies, or antigen-binding fragments thereof include Siplizumab (i.e., MEDI-507 or TCD601, ITB-Med LLC), BTI-322 (Lo-CD2a), Alefacept (i.e., a chimeric fusion protein consisting of the CD2-binding portion of human LFA3-Fc, Biogen) CB.219 (e.g., BioXCell), UMCD2 (e.g., Santa Cruz Biotechnology), TS1 / 8, RPA-2.10, TS1 / 18, TS1 / 18.1.1, TS2 / 18, AB75, and ZR100, as well as anti-CD2 antibodies or antigen-binding Page 144 of 372 12815032v1Attorney Docket No.: 2017469-0043 fragments thereof disclosed in any of: US 5,730,979; US 5,928,643; US 5,951,983; US 6,764,681; US 7,858,095; US 6,162,432; US 11,732,042; US 12,037,378; US20210032308; US20030068320; US20230365687; WO1999058147; WO2014025198; WO2024180185; WO2023126445; WO2024079046; WO2024079046; etc., each hereby incorporated by reference in its entirety.

[0500] In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:269 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:270. In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:280 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:281. In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:291 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:292. In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:302 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:292. SEQ ID NOs:269, 270, 280, 281, 291, 292, and 302 are shown in Table 2N.

[0501] In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:269, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:270. In some embodiments, the CD2 targeting moiety comprises a CDR-H1 comprising an amino acid sequence SYWVN (SEQ ID NO:271), a CDR-H2 comprising an amino acid sequence RIDPYDSETHYNQKFTD (SEQ ID NO:272), a CDR-H3 comprising an amino acid sequence SPRDSSTNLAD (SEQ ID NO:273), a CDR-L1 comprising an amino acid sequence RASQSISDYLH (SEQ ID NO:274), a CDR-L2 comprising an amino acid sequence YASQSIS (SEQ ID NO:275), and a CDR-L3 comprising an amino acid sequence QNGHSFPLT (SEQ ID NO:276). In some embodiments, the CD2 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% Page 145 of 372 12815032v1Attorney Docket No.: 2017469-0043 (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:266 or 267, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:268.

[0502] In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:280, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:281. In some embodiments, the CD2 targeting moiety comprises a CDR-H1 comprising an amino acid sequence RYWIH (SEQ ID NO:282), a CDR-H2 comprising an amino acid sequence NIDPSDSETHYNQKFKD (SEQ ID NO:283), a CDR-H3 comprising an amino acid sequence EDLYYAMEY (SEQ ID NO:284), a CDR-L1 comprising an amino acid sequence KSSQSVLYSSNQKNYLA (SEQ ID NO:285), a CDR-L2 comprising an amino acid sequence WASTRES (SEQ ID NO:147), and a CDR-L3 comprising an amino acid sequence HQYLSSHT (SEQ ID NO:287). In some embodiments, the CD2 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:277 or 278, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:279.

[0503] In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:291, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid Page 146 of 372 12815032v1Attorney Docket No.: 2017469-0043 sequence of SEQ ID NO:292. In some embodiments, the CD2 targeting moiety comprises a CDR-H1 comprising an amino acid sequence EYYMY (SEQ ID NO:293), a CDR-H2 comprising an amino acid sequence RIDPEDGSIDYVEKFKK (SEQ ID NO:294), a CDR-H3 comprising an amino acid sequence GKFNYRFAY (SEQ ID NO:295), a CDR-L1 comprising an amino acid sequence RSSQSLLHSSGNTYLN (SEQ ID NO:296), a CDR-L2 comprising an amino acid sequence LVSKLES (SEQ ID NO:297), and a CDR-L3 comprising an amino acid sequence MQFTHYPYT (SEQ ID NO:298). In some embodiments, the CD2 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:288 or 289, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:290.

[0504] In some embodiments, the CD2 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:302, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:292. In some embodiments, the CD2 targeting moiety comprises a CDR-H1 comprising an amino acid sequence EYYMY (SEQ ID NO:293), a CDR-H2 comprising an amino acid sequence RIDPEDGSIDYVEKFKK (SEQ ID NO:294), a CDR-H3 comprising an amino acid sequence GKFNYRFAY (SEQ ID NO:295), a CDR-L1 comprising an amino acid sequence RSSQSLLHSSGNTYLN (SEQ ID NO:296), a CDR-L2 comprising an amino acid sequence LVSKLES (SEQ ID NO:297), and a CDR-L3 comprising an amino acid sequence MQFTHYPYT (SEQ ID NO:298). In some embodiments, the CD2 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:299 or 300, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at Page 147 of 372 12815032v1Attorney Docket No.: 2017469-0043 least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:290. CD3 Targeting Moieties

[0505] In some embodiments, the target molecule is CD3. In some embodiments, the target cell is CD3+. CD3 is a multimeric protein complex made up of four polypeptide chains (CD3-epsilon (ε), CD3-gamma (γ), CD3-delta (δ), and CD3-zeta (ζ)) to form a CD3γε–CD3δε–CD3ζζ signaling hexamer that associates with the T cell receptor (TCR). The CD3 / TCR complex is critical for T cells to recognize foreign antigens and activate T-cell adaptive immunity. CD3 is expressed by all T cells and is a defining marker of the T lymphocyte lineage. See, e.g., Dong et al. (2019) Nature 573:546-552, hereby incorporated by reference in its entirety.

[0506] In some embodiments, the CD3 targeting moiety includes an antibody or antigen-binding fragment thereof that binds to CD3. In some embodiments, the CD3 targeting moiety is an antibody or antigen-binding fragment thereof (e.g., a Fab, Fab’, F(ab’)2, Fv fragment, scFv, DARPIN, VHH domain, FN3 domain, nanobody, single domain antibody, or Centyrin). In other embodiments, the CD3 targeting moiety includes a ligand, a folate moiety, an antibiotic mimetic, a polynucleotide (such as a DNA or RNA apatamer), a carbohydrate, a vitamin, a cytokine, or a chemokine. In some embodiments, the CD3 targeting moiety is an anti-CD3 antibody or antigen binding fragment thereof. In some embodiments, the CD3 targeting moiety is an IgA, IgG, IgE, or IgM antibody. In some embodiments, the CD3 targeting moiety is a bispecific or multi- specific antibody or fragment thereof. In some embodiments, the CD3 targeting moiety is a humanized antibody or antigen-binding fragment thereof.

[0507] Exemplary anti-CD3 binders, antibodies, or antigen-binding fragments thereof include SP34 mouse monoclonal antibody (see, for example, Pressano, S. The EMBO J.4:337-344, 1985; Alarcon, B. EMBO J.10:903-912, 1991; Salmeron A. et al., J. Immunol.147:3047-52, 1991; Yoshino N. et al., Exp. Anim 49:97-110, 2000; Conrad M L. et al., Cytometry 71A:925-33, 2007; Yang et al., J. Immunol.137:1097-1100: 1986; US 8,846,042; US 11,013,800; and US 10,870,701), Cris-7 monoclonal antibody (Reinherz, E. L. et al. (eds.), Leukocyte typing II, Springer Verlag, New York, (1986)), BC3 monoclonal antibody (Anasetti et al. (1990) J. Exp. Med.172:1691), OKT3 (Ortho multicenter Transplant Study Group (1985) N. Engl. J. Med. 313:337) and derivatives thereof such as OKT3 ala-ala (Herold et al. (2003) J. Clin. Invest. Page 148 of 372 12815032v1Attorney Docket No.: 2017469-0043 11:409), visilizumab (Carpenter et al. (2002) Blood 99:2712), mosunetuzumab, odronextamab, tebentafusp, teplizumab, teclistamab, muromonab, plamotamab, HPN536, pasotuxizumab, flotetuzumab, and 145-2C11 monoclonal antibody (Hirsch et al. (1988) J. Immunol.140: 3766). Further CD3 binding molecules contemplated herein include UCHT-1 (Beverley, P C and Callard, R. E. (1981) Eur. J. Immunol.11: 329-334) and CD3 binding molecules described in WO2004 / 106380; WO2010 / 037838; WO2008 / 119567; WO2007 / 042261; WO2010 / 0150918; the contents of each of which are incorporated herein by reference in their entirety.

[0508] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:141 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:142. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:152 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:153. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:163 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:164. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:174 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:175. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:185 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:186. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:196 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:197. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:207 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:208. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:218 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:219. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:228 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:229. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of Page 149 of 372 12815032v1Attorney Docket No.: 2017469-0043 SEQ ID NO:238 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:239. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:248 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:249. In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:258 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:259. SEQ ID NOs:141-142, 152-153, 163-163, 174-175, 185-186, 196-197, 207-208, 218-219, 228-229, 238-239, 248-249, and 258-259 are shown in tables below.

[0509] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:141, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:142. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence NYYIH (SEQ ID NO:143), a CDR-H2 comprising an amino acid sequence WIYPGDGNTKYNEKFKG (SEQ ID NO:144), a CDR-H3 comprising an amino acid sequence DSYSNYYFDY (SEQ ID NO:145), a CDR-L1 comprising an amino acid sequence KSSQSLLNSRTRKNYLA (SEQ ID NO:146), a CDR-L2 comprising an amino acid sequence WASTRES (SEQ ID NO:147), and a CDR-L3 comprising an amino acid sequence TQSFILRT (SEQ ID NO:148). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:138 or 139, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:140. In some embodiments, the CD3 targeting moiety is mosunetuzumab or an antigen-binding fragment thereof.

[0510] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least Page 150 of 372 12815032v1Attorney Docket No.: 2017469-0043 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:152, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:153. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence DYTMH (SEQ ID NO:154), a CDR-H2 comprising an amino acid sequence GISWNSGSIGYADSVKG (SEQ ID NO:155), a CDR-H3 comprising an amino acid sequence DNSGYGHYYYGMDV (SEQ ID NO:156), a CDR-L1 comprising an amino acid sequence RASQSVSSNLA (SEQ ID NO:157), a CDR-L2 comprising an amino acid sequence GASTRAT (SEQ ID NO:158), and a CDR-L3 comprising an amino acid sequence QHYINWPLT (SEQ ID NO:159). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:149 or 150, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:151. In some embodiments, the CD3 targeting moiety is odronextamab or an antigen-binding fragment thereof.

[0511] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:163, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:164. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence GYTMN (SEQ ID NO:165), a CDR-H2 comprising an amino acid sequence LINPYKGVSTYNQKFKD (SEQ ID NO:166), a CDR-H3 comprising an amino acid sequence SGYYGDSDWYFDV (SEQ ID NO:167), a CDR-L1 comprising an amino acid sequence RASQDIRNYLN (SEQ ID NO:168), a CDR-L2 comprising an amino acid sequence YTSRLES (SEQ ID NO:169), and a CDR-L3 comprising an amino acid sequence QQGNTLPWT (SEQ ID NO:170). In some embodiments, the CD3 targeting moiety is Page 151 of 372 12815032v1Attorney Docket No.: 2017469-0043 a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:160 or 161, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:162. In some embodiments, the CD3 targeting moiety is tebentafusp or an antigen-binding fragment thereof.

[0512] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:174, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:175. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence RYTMH (SEQ ID NO:176), a CDR-H2 comprising an amino acid sequence YINPSRGYTNYNQKVKD (SEQ ID NO:177), a CDR-H3 comprising an amino acid sequence YYDDHYCLDY (SEQ ID NO:178), a CDR-L1 comprising an amino acid sequence SASSSVSYMN (SEQ ID NO:179), a CDR-L2 comprising an amino acid sequence DTSKLAS (SEQ ID NO:180), and a CDR-L3 comprising an amino acid sequence QQWSSNPFT (SEQ ID NO:181). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:171 or 172, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:173. In some embodiments, the CD3 targeting moiety is teplizumab or an antigen-binding fragment thereof.

[0513] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to Page 152 of 372 12815032v1Attorney Docket No.: 2017469-0043 the amino acid sequence of SEQ ID NO:185, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:186. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence NTYAMN (SEQ ID NO:187), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYAASVKG (SEQ ID NO:188), a CDR- H3 comprising an amino acid sequence HGNFGNSYVSWFAY (SEQ ID NO:189), a CDR-L1 comprising an amino acid sequence RSSTGAVTTSNYAN (SEQ ID NO:190), a CDR-L2 comprising an amino acid sequence GTNKRAP (SEQ ID NO:191), and a CDR-L3 comprising an amino acid sequence ALWYSNLWV (SEQ ID NO:192). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:182 or 183, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:184. In some embodiments, the CD3 targeting moiety is teclistamab or an antigen-binding fragment thereof.

[0514] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:196, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:197. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence SYTMH (SEQ ID NO:198), a CDR-H2 comprising an amino acid sequence YINPRSGYTHYNQKLKD (SEQ ID NO:199), a CDR-H3 comprising an amino acid sequence SAYYDYDGFAY (SEQ ID NO:200), a CDR-L1 comprising an amino acid sequence SASSSVSYMN (SEQ ID NO:179), a CDR-L2 comprising an amino acid sequence DTSKLAS (SEQ ID NO:180), and a CDR-L3 comprising an amino acid sequence QQWSSNPPT (SEQ ID NO:203). In some embodiments, the CD3 targeting moiety is Page 153 of 372 12815032v1Attorney Docket No.: 2017469-0043 a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:193 or 194, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:195. In some embodiments, the CD3 targeting moiety is visilizumab or an antigen-binding fragment thereof.

[0515] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:207, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:208. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence RYTMH (SEQ ID NO:176), a CDR-H2 comprising an amino acid sequence YINPSRGYTNYNQKFKD (SEQ ID NO:210), a CDR-H3 comprising an amino acid sequence YYDDHYCLDY (SEQ ID NO:178), a CDR-L1 comprising an amino acid sequence SASSSVSYMN (SEQ ID NO:179), a CDR-L2 comprising an amino acid sequence DTSKLAS (SEQ ID NO:180), and a CDR-L3 comprising an amino acid sequence QQWSSNPFT (SEQ ID NO:181). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:204 or 205, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:206. In some embodiments, the CD3 targeting moiety is muromonab or an antigen-binding fragment thereof.

[0516] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to Page 154 of 372 12815032v1Attorney Docket No.: 2017469-0043 the amino acid sequence of SEQ ID NO:218, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:219. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence KYAMN (SEQ ID NO:220), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYADSVKD (SEQ ID NO:221), a CDR- H3 comprising an amino acid sequence HGNFGNSYISYWAY (SEQ ID NO:222), a CDR-L1 comprising an amino acid sequence GSSTGAVTSGNYPN (SEQ ID NO:223), a CDR-L2 comprising an amino acid sequence GTKFLAP (SEQ ID NO:224), and a CDR-L3 comprising an amino acid sequence VLWYSNRWV (SEQ ID NO:225). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:215 or 216, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:217. In some embodiments, the CD3 targeting moiety is SP34 or an antigen-binding fragment thereof.

[0517] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:228, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:229. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence TYAMN (SEQ ID NO:230), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYADSVKG (SEQ ID NO:231), a CDR- H3 comprising an amino acid sequence HGNFGDSYVSWFAY (SEQ ID NO:232), a CDR-L1 comprising an amino acid sequence GSSTGAVTTSNYAN (SEQ ID NO:233), a CDR-L2 comprising an amino acid sequence GTNKRAP (SEQ ID NO:191), and a CDR-L3 comprising an amino acid sequence ALWYSNHWV (SEQ ID NO:235). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid Page 155 of 372 12815032v1Attorney Docket No.: 2017469-0043 sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:226, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:227. In some embodiments, the CD3 targeting moiety is plamotamab or an antigen-binding fragment thereof.

[0518] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:238, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:239. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence KYAIN (SEQ ID NO:240), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYADQVKD (SEQ ID NO:241), a CDR-H3 comprising an amino acid sequence HANFGNSYISYWAY (SEQ ID NO:242), a CDR-L1 comprising an amino acid sequence ASSTGAVTSGNYPN (SEQ ID NO:243), a CDR-L2 comprising an amino acid sequence GTKFLVP (SEQ ID NO:244), and a CDR-L3 comprising an amino acid sequence TLWYSNRWV (SEQ ID NO:245). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:236, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:237. In some embodiments, the CD3 targeting moiety is HPN536 or an antigen-binding fragment thereof.

[0519] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least Page 156 of 372 12815032v1Attorney Docket No.: 2017469-0043 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:218, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:249. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence KYAMN (SEQ ID NO:220), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYADSVKD (SEQ ID NO:221), a CDR- H3 comprising an amino acid sequence HGNFGNSYISYWAY (SEQ ID NO:222), a CDR-L1 comprising an amino acid sequence GSSTGAVTSGNYPN (SEQ ID NO:223), a CDR-L2 comprising an amino acid sequence GTKFLAP (SEQ ID NO:224), and a CDR-L3 comprising an amino acid sequence VLWYSNRWV (SEQ ID NO:225). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:246, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:247. In some embodiments, the CD3 targeting moiety is pasotuxizumab or an antigen-binding fragment thereof.

[0520] In some embodiments, the CD3 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:258, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:259. In some embodiments, the CD3 targeting moiety comprises a CDR-H1 comprising an amino acid sequence TYAMN (SEQ ID NO:230), a CDR-H2 comprising an amino acid sequence RIRSKYNNYATYYADSVKD (SEQ ID NO:221), a CDR- H3 comprising an amino acid sequence HGNFGNSYVSWFAY (SEQ ID NO:189), a CDR-L1 comprising an amino acid sequence RSSTGAVTTSNYAN (SEQ ID NO:190), a CDR-L2 comprising an amino acid sequence GTNKRAP (SEQ ID NO:191), and a CDR-L3 comprising Page 157 of 372 12815032v1Attorney Docket No.: 2017469-0043 an amino acid sequence ALWYSNLWV (SEQ ID NO:192). In some embodiments, the CD3 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:256, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:257. In some embodiments, the CD3 targeting moiety is flotetuzumab or an antigen-binding fragment thereof. CD5 Targeting Moieties

[0521] In some embodiments, the target molecule is CD5. In some embodiments, the target cell is CD5+. CD5 is a type-I transmembrane glycoprotein with an extracellular region composed of three scavenger receptor cysteine-rich (SRCR) domains. Several CD5 ligands have been reported such as CD72, the IgV(H) frame-work region and several polypeptides (gp40-80, gp150) whose identity remains undetermined. CD5 regulates T cell functions and development, including negative regulation of TCR signaling. CD5 is an activation marker of T cells, wherein the expression of CD5 increases according to the magnitude of the signal delivered by the TCR. Consequently, CD5 expression reflects the heterogeneity of the signal strength associated with each individual TCR within a polyclonal T cell population. See, e.g., Voisinne et al. (2018) Front. Immunol.9:2900, hereby incorporated by reference in its entirety.

[0522] In some embodiments, the CD5 targeting moiety includes an antibody or antigen-binding fragment thereof that binds to CD5. In some embodiments, the CD5 targeting moiety is an antibody or antigen-binding fragment thereof (e.g., a Fab, Fab’, F(ab’)2, Fv fragment, scFv, DARPIN, VHH domain, FN3 domain, nanobody, single domain antibody, or Centyrin). In other embodiments, the CD5 targeting moiety includes a ligand, a folate moiety, an antibiotic mimetic, a polynucleotide (such as a DNA or RNA apatamer), a carbohydrate, a vitamin, a cytokine, or a chemokine. In some embodiments, the CD5 targeting moiety is an anti-CD5 antibody or antigen binding fragment thereof. In some embodiments, the CD5 targeting moiety is an IgA, IgG, IgE, or IgM antibody. In some embodiments, the CD5 targeting moiety is a bispecific or multi- Page 158 of 372 12815032v1Attorney Docket No.: 2017469-0043 specific antibody or fragment thereof. In some embodiments, the CD5 targeting moiety is a humanized antibody or antigen-binding fragment thereof.

[0523] Exemplary anti-CD5 binders, antibodies, or antigen-binding fragments thereof include AFM 16 (e.g., Affimed Therapeutics); AFM 17 (e.g., Affimed Therapeutics), RM354, L17F12, CRIS-1, UCHT2, RM314, SP19, and CD5-5D7, as well as anti-CD5 antibodies or antigen- binding fragments thereof disclosed in any of: US 10,786,549; US20110250203; Dai et al. (2021) Mol Ther.29(9)2707-2722; etc., each hereby incorporated by reference in its entirety.

[0524] In some embodiments, the CD5 targeting moiety comprises a heavy chain variable region comprising the amino acid sequence of SEQ ID NO:357 and a light chain variable region comprising the amino acid sequence of SEQ ID NO:358. SEQ ID NOs:357 and 358 are shown in Table 20, with complementary determining regions (CDRs) marked in bold.

[0525] In some embodiments, the CD5 targeting moiety comprises a heavy chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:357, and / or a light chain variable region comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:358. In some embodiments, the CD5 targeting moiety comprises a CDR-H1 comprising an amino acid sequence TSGMGVG (SEQ ID NO:359), a CDR-H2 comprising an amino acid sequence HIWWDDDVYYNPSLKS (SEQ ID NO:360), a CDR-H3 comprising an amino acid sequence RRATGTGFDY (SEQ ID NO:361), a CDR-L1 comprising an amino acid sequence QASQDVGTAVA (SEQ ID NO:362), a CDR-L2 comprising an amino acid sequence WTSTRHT (SEQ ID NO:363), and a CDR-L3 comprising an amino acid sequence HQYNSYNT (SEQ ID NO:364). In some embodiments, the CD5 targeting moiety is a Fab fragment comprising a first polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98% or at least 99%) sequence identity to the amino acid sequence of SEQ ID NO:354 or 355, and a second polypeptide comprising an amino acid sequence having at least 90% (e.g., at least 92%, at least 93%, at least 94%...

Claims

Attorney Docket No.: 2017469-0043 CLAIMS What is claimed is:

1. A nucleic acid molecule comprising a modified RTE-13’ UTR that comprises a nucleotide difference (e.g., substitution, insertion, or deletion) at one or more of: Region A, Region B, Region D, or Region E, as numbered according to SEQ ID NO:

2.

2. The nucleic acid molecule of claim 1, wherein the modified RTE-13’ UTR comprises 5, 6, 7, or 8 of, or all of: Region A, Region B, Region C, Region D, Region E, Region F, Region G, or Region H.

3. A nucleic acid molecule comprising a modified RTE-13’ UTR that comprises a nucleotide difference (e.g., substitution or deletion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or all of positions A4, C6, A21, U29, C30, G31, A32, C47, C48, A49, C50, or U51, as numbered according to SEQ ID NO:

2.

4. The nucleic acid molecule of any of the preceding claims, which has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, 98%, or 99% identity to SEQ ID NO:

2.

5. The nucleic acid molecule of any of the preceding claims, which has 1, 2, 3, or all of substitutions: A4U, C6A, A21U, or A32G, as numbered according to SEQ ID NO:

2.

6. The nucleic acid molecule of any of the preceding claims, which has a deletion of position 29, 30, 31, or 32 as numbered according to SEQ ID NO:

2.

7. The nucleic acid of any of the preceding claims, wherein the modified RTE-13’ UTR has a length of 25-30, 30-35, 35-40, 40-45, 45-50, or 50-51 nucleotides.

8. The nucleic acid molecule of any of the preceding claims, which has an A21U substitution as numbered according to SEQ ID NO:

2.

9. The nucleic acid of any of the preceding claims, which has an C6A substitution as numbered according to SEQ ID NO:

2. Page 359 of 372 12815032v1Attorney Docket No.: 2017469-0043 10. The nucleic acid of any one of the preceding claims, which has an A32G substitution as numbered according to SEQ ID NO:

2.

11. The nucleic acid molecule of nucleic acid of any one of the preceding claims, which comprises a deletion of position 31, as numbered according to SEQ ID NO:

2.

12. The nucleic acid of any one of the preceding claims, which comprises a deletion of position 32, as numbered according to SEQ ID NO:

2.

13. The nucleic acid molecule of any of the preceding claims, which has an A32G substitution and a deletion of position 30 as numbered according to SEQ ID NO:

2.

14. The nucleic acid molecule of any of the preceding claims, which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

15. The nucleic acid molecule of any of the preceding claims, which has an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

16. The nucleic acid molecule of any of the preceding claims, which has a C6A substitution, an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

17. The nucleic acid molecule of any of the preceding claims, which has a C6A substitution and an A32G substitution as numbered according to SEQ ID NO:

2.

18. The nucleic acid molecule of any of the preceding claims, which has an A21U substitution and an A32G substitution as numbered according to SEQ ID NO:

2.

19. The nucleic acid molecule of any of the preceding claims, which has a C6A substitution and an A21U substitution as numbered according to SEQ ID NO:

2.

20. A nucleic acid molecule comprising a modified RTE-15’ UTR that comprises a nucleotide difference (e.g., substitution, deletion, or insertion) at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO:

1. Page 360 of 372 12815032v1Attorney Docket No.: 2017469-0043 21. The nucleic acid molecule of claim 20, which has a sequence having at least 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 98% identity to SEQ ID NO:

1.

22. The nucleic acid molecule of any of claims 20 or 21, which comprises a nucleotide substitution at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or all of positions U138, U180, A208, U264, A271, G307, G358, A398, U399, A405, U407, G469, U488, G570, or U658, as numbered according to SEQ ID NO:

1.

23. The nucleic acid of any of claims 20-22, wherein the modified RTE-15’ UTR has a length of 400-450, 450-500, 500-550, 550-600, 600-650, 650-700, or 700-710 nucleotides, e.g., 700-710 nucleotides.

24. The nucleic acid of any one of claims 20-23, which has an A405G substitution as numbered according to SEQ ID NO:

1.

25. The nucleic acid of any one of claims 20-24, which has a U180C substitution and a U399C substitution as numbered according to SEQ ID NO:

1.

26. The nucleic acid of any one of claims 20-25, which has a U180C substitution, an A398G substitution, and an A405G substitution as numbered according to SEQ ID NO:

1.

27. The nucleic acid of any one of claims 20-26, which has a U138C substitution, a U180C substitution, an A405G substitution, and a U488C substitution as numbered according to SEQ ID NO:

1.

28. The nucleic acid of any one of claims 20-27, which has a U180C substitution, an A398G substitution, an A405G substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO:

1.

29. The nucleic acid of any one of claims 20-28, which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO:

1. Page 361 of 372 12815032v1Attorney Docket No.: 2017469-0043 30. The nucleic acid of any one of claims 20-29, which has a U138C substitution, a U180C substitution, an A405G substitution, a U488C substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO:

1.

31. The nucleic acid of any one of claims 20-30, which has a U138C substitution, a U180C substitution, a U264G substitution, a U399C substitution, an A405G substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO:

1.

32. The nucleic acid of any one of claims 20-30, which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G358A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO:

1.

33. The nucleic acid molecule of any of the preceding claims, which binds an RTE-1 polypeptide of Table 1 or Table 22.

34. The nucleic acid molecule of any preceding claims, wherein the nucleic acid comprises one or more chemically modified nucleotides.

35. A template RNA comprising, from 5’ to 3’: a) optionally, a RTE-15’ UTR; b) a heterologous object sequence; and c) a nucleic acid molecule of any of claims 1-19, 33, or 34.

36. The template RNA of claim 35, wherein the template RNA comprises an RTE-15’ UTR, and wherein the RTE-15’ UTR comprises a nucleic acid sequence according to SEQ ID NO:

1.

37. The template RNA of any of claims 35 or 36, wherein the template RNA comprises an RTE-15’ UTR, and wherein the RTE-15’ UTR comprises a modified RTE-15’ UTR.

38. A template RNA comprising, from 5’ to 3’: a) a nucleic acid molecule of any of claims 20-34; Page 362 of 372 12815032v1Attorney Docket No.: 2017469-0043 b) a heterologous object sequence; and c) optionally, an RTE-13’ UTR.

39. The template RNA of claim 38, wherein the template RNA comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a nucleic acid sequence according to SEQ ID NO:

2.

40. The template RNA of any of claims 38 or 39, wherein the template RNA comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a modified RTE-13’ UTR.

41. The template RNA of any of claims 38-40, wherein the template comprises an RTE-13’ UTR, and wherein the RTE-13’ UTR comprises a modified RTE-13’ UTR according to any of claims 1-19, 33, or 34.

42. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence having an A21U substitution as numbered according to SEQ ID NO:

2.

43. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence having a deletion of position 32 as numbered according to SEQ ID NO:

2.

44. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and Page 363 of 372 12815032v1Attorney Docket No.: 2017469-0043 c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

45. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

46. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution, an A21U substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

47. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and an A32G substitution as numbered according to SEQ ID NO:

2.

48. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and Page 364 of 372 12815032v1Attorney Docket No.: 2017469-0043 c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution and an A32G substitution as numbered according to SEQ ID NO:

2.

49. A template RNA comprising, from 5’ to 3’: a) an RTE-15’ UTR comprising a nucleic acid sequence according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and an A21U substitution as numbered according to SEQ ID NO:

2.

50. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

51. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO:

2.

52. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and Page 365 of 372 12815032v1Attorney Docket No.: 2017469-0043 c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO:

2.

53. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

54. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution and a U399C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

55. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution, an A398G substitution, and an A405G substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

56. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, an A405G substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; Page 366 of 372 12815032v1Attorney Docket No.: 2017469-0043 b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

57. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U180C substitution, an A398G substitution, an A405G substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

58. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

59. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

60. A template RNA comprising, from 5’ to 3’: Page 367 of 372 12815032v1Attorney Docket No.: 2017469-0043 a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

61. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO:

2.

62. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U264G substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO:

2.

63. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, an A405G substitution, a U488C substitution, and an insertion of an A following U658 as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and Page 368 of 372 12815032v1Attorney Docket No.: 2017469-0043 c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

64. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, a U180C substitution, a U264G substitution, a U399C substitution, an A405G substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

65. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) an RTE-13’ UTR comprising a nucleic acid sequence according to SEQ ID NO:

2.

66. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has an A21U substitution as numbered according to SEQ ID NO:

2.

67. A template RNA comprising, from 5’ to 3’: Page 369 of 372 12815032v1Attorney Docket No.: 2017469-0043 a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a deletion of position 32 as numbered according to SEQ ID NO:

2.

68. A template RNA comprising, from 5’ to 3’: a) a modified RTE-15’ UTR comprising a nucleic acid sequence which has a U138C substitution, an A208G substitution, a U264G substitution, an A271G substitution, a G307U substitution, a G538A substitution, a U399C substitution, a G469A substitution, and a U488C substitution as numbered according to SEQ ID NO: 1; b) a heterologous object sequence; and c) a modified RTE-13’ UTR comprising a nucleic acid sequence which has a C6A substitution and a deletion of position 32 as numbered according to SEQ ID NO:

2.

69. The template RNA of any of claims 35-68, which promotes integration of a heterologous object sequence into a target nucleic acid in an assay comprising: a) providing a nucleic acid encoding an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392; and b) contacting the nucleic acid of a) and the template RNA with a cell comprising the target nucleic acid, under conditions that allow for integration of the heterologous object sequence into the target nucleic acid.

70. A gene modifying system comprising: a template RNA of any of claims 35-68; and Page 370 of 372 12815032v1Attorney Docket No.: 2017469-0043 a gene modifying polypeptide, or a nucleic acid encoding the gene modifying polypeptide, the gene modifying polypeptide comprising an RTE-1 polypeptide having an amino acid sequence of SEQ ID NO: 403, SEQ ID NO: 402, or SEQ ID NO: 392, or an amino acid sequence having at least 80%, 90%, 95%, 97%, 98%, 99% identity to SEQ ID NO: 403, SEQ ID NO: 402, SEQ ID NO:

392.

71. A pharmaceutical composition comprising the nucleic acid of any one of claims 1-34 or the gene modifying system of claim 70, and a pharmaceutically acceptable excipient or carrier.

72. The pharmaceutical composition of claim 71, wherein the pharmaceutically acceptable excipient or carrier is selected from the group consisting of a plasmid vector, a viral vector, a vesicle, and a lipid nanoparticle (LNP).

73. A host cell (e.g., a mammalian cell, e.g., a human cell) comprising the gene modifying system, template RNA, or nucleic acid of any of claims 1-70.

74. A method of making the cell of claim 73, the method comprising providing the cell and contacting the cell with the gene modifying system of claim 70. 75.A method of making the nucleic acid or template RNA of any one of claims 1-69, the method comprising synthesizing the template RNA in vitro (e.g., by in vitro transcription or solid state synthesis).

76. A method of modifying the genome of a mammalian cell, comprising contacting the cell with a gene modifying system of claim 70, thereby modifying the genome of the mammalian cell.

77. A cell made by the method of claim 76.

78. A method for treating a subject having cancer, the method comprising administering to the subject the gene modifying system of claim 70 or DNA encoding the same, or the pharmaceutical composition of any one of claims 71-72, thereby treating the subject having the cancer. Page 371 of 372 12815032v1

Citation Information

Patent Citations

  • Compositions and methods for modulating a genome in t cells, induced pluripotent stem cells, and respiratory epithelial cells

    WO2023212724A2